[Infra] Promote Internal Staging to main - #26731
Conversation
…page The edit page was calling POST /mcp-rest/test/tools/list (the temp-session endpoint that requires inline credentials) on mount. Since fetchTools deliberately omits credentials from the request body, any server with auth_type api_key/bearer_token/basic/authorization would 422. Switch to GET /mcp-rest/tools/list?server_id=... which looks up stored credentials on the backend — no inline creds needed for saved servers.
Previously, the "Store Prompts in Spend Logs" and "Maximum Spend Logs Retention Period" settings were surfaced via a gear-icon modal on the Logs page. The gear was visible to every authenticated user even though the backend endpoints (/config/update, /config/list) require PROXY_ADMIN — so non-admins could open the modal but the request would 403 on load and save, giving a confusing UX. Move the controls into a new "Logging Settings" tab under Admin Settings, which is already gated to admins at the sidebar. Remove the gear button and the onOpenSettings prop chain (ConfigInfoMessage → LogDetailContent → LogDetailsDrawer). ConfigInfoMessage now points users to "Admin Settings → Logging Settings" inline.
…erun CircleCI's 'Rerun failed tests' feature passes test identifiers from the JUnit XML classname attribute (dot notation, e.g. 'tests.local_testing.test_router') via stdin. pytest receives these paths and collects 0 items, causing the rerun to exit 123 with no tests run. Add an awk preprocessor before xargs that detects dot-notation module paths and converts them to file paths (tests/local_testing/test_router.py). File paths already containing '.py' are passed through unchanged. Applied to all three jobs using the 'circleci tests run' + 'xargs pytest' pattern: local_testing_part1, local_testing_part2, and the router test job.
Made-with: Cursor
Made-with: Cursor
Made-with: Cursor
…test Pytest tests inside a class produce JUnit XML classnames like 'tests.local_testing.test_file_types.TestFileConsts' (module + class). The previous awk preprocessor would convert this to 'tests/local_testing/test_file_types/TestFileConsts.py', which doesn't exist, causing pytest to collect 0 items on rerun. Strip a trailing '.<UppercaseSegment>' before the dot-to-slash conversion. Module path segments are lowercase (test files start with 'test_'), and the class name is the only segment beginning with an uppercase letter, so this is unambiguous. Verified affected files in tests/local_testing/: test_file_types.py (TestFileConsts), test_gcs_cache_unit_tests.py, test_disk_cache_unit_tests.py, test_docker_no_network_on_deploy.py, test_sagemaker_nova_integration.py, test_cache_preset_key.py.
Made-with: Cursor
Pass externalTools/externalIsLoading/externalError/externalCanFetch from the edit page so MCPToolConfiguration consumes the parent's GET fetch instead of firing its own POST /test/tools/list via useTestMCPConnection. Eliminates the spurious POST that caused the user-visible "Unable to load tools" error for api_key/bearer_token/basic/authorization servers.
…itellm_fix-edit-page-tools-fetch-422
…26441) * fix(redis): cache GCP IAM token to prevent async event loop blocking ## Problem GCPIAMCredentialProvider.get_credentials() calls _generate_gcp_iam_access_token on every Redis connection establishment. This function performs synchronous HTTP and gRPC calls (google-auth + google-cloud-iam) which block Python's asyncio event loop while running. Under concurrent load (e.g. connection pool warm-up, parallel health checks), multiple connections are established simultaneously, each triggering an independent blocking IAM token refresh. These refreshes serialise behind each other inside the single-threaded event loop, causing individual Redis spans to take 20-25 seconds instead of milliseconds. Observed in production via Datadog APM: a single INCRBYFLOAT Redis span took 25.6 seconds (90% of a 28.4s trace), with GCP metadata + GenerateAccessToken gRPC calls visible inside the span. This cascaded into aiohttp SocketTimeoutError on upstream LLM API calls — not because the upstream was slow, but because the event loop was frozen and the 30-second sock_read timer fired on a connection that was never given CPU time. ## Fix Add a module-level token cache (dict keyed by service account, value is (token, expiry_monotonic)). _get_cached_gcp_iam_token() returns the cached token on cache hit (no I/O), and refreshes only when expired using double-checked locking so only one thread performs the network round-trip. GCP IAM tokens are valid for 1 hour; the cache TTL is set to 55 minutes (_GCP_IAM_TOKEN_TTL_SECONDS = 3300) to refresh safely before expiry. The cache is shared across all GCPIAMCredentialProvider instances for the same service account, so N concurrent Redis connections on the same pod share a single token and avoid N concurrent blocking refreshes. get_credentials_async() already used asyncio.to_thread (non-blocking), and is updated to call _get_cached_gcp_iam_token so it also benefits from caching. ## Tests - Updated existing test that expected a fresh token on every call to reflect the new caching behaviour. - Added tests for: cache hit (no redundant I/O), cache expiry and refresh, and cache sharing across multiple provider instances. - Added autouse fixture to clear the module-level cache between tests. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com> * refactor(redis): remove unused Optional import from _redis_credential_provider.py * refactor(redis): improve documentation for GCPIAMCredentialProvider class Updated the docstring for the GCPIAMCredentialProvider class to clarify its purpose and the caching mechanism for GCP IAM tokens. The changes enhance readability and maintainability by providing a more concise explanation of the token caching strategy and its benefits for Redis authentication. * refactor(redis): improve documentation for GCPIAMCredentialProvider class Updated the docstring for the GCPIAMCredentialProvider class to clarify its purpose and the caching mechanism for GCP IAM tokens. The changes enhance readability and maintainability by providing a more concise explanation of the token caching strategy and its benefits for Redis authentication. --------- Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
…5855) Bedrock enforces non-increasing TTL ordering across cache_control blocks (tools → system → messages). The tool cache_control TTL was being unconditionally dropped to the default 5m, while system blocks preserved the user-specified TTL for Claude 4.5+ models. This mismatch caused "a ttl='1h' block must not come after a ttl='5m' block" errors when users set ttl='1h' on both tools and system. Converse path: add_cache_point_tool_block() now accepts a model param and preserves TTL for Claude 4.5+, matching _get_cache_point_block(). Invoke path: _remove_ttl_from_cache_control() now also processes tools (was only processing system and messages). Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…onses (#20270) (#26262) * fix(proxy): invoke post-call guardrails on pass-through endpoint responses (#20270) Wire post_call_success_hook into non-streaming pass-through response path, gated on explicit guardrail config (opt-in only, no backwards-compat break). - Call post_call_success_hook after reading non-streaming response body - Build enriched hook_data with guardrails metadata and litellm_logging_obj at call site (avoids mutation of _parsed_body which is shared by logging) - Handle ModifyResponseException with provider-agnostic error envelope, post_call_failure_hook, and defensive try/except - Strip stale content-length when guardrail modifies response body - Move ModifyResponseException to litellm.exceptions to break cyclic import; re-export from custom_guardrail for backwards compat - Add call_type fallback in UnifiedLLMGuardrails for pass-through endpoints using CallTypes.pass_through.value enum * test: add unit tests for pass-through post-call guardrails 5 tests covering the post-call guardrail invocation on pass-through endpoints: - post_call_success_hook fires when guardrails configured - post_call_success_hook skipped when no guardrails (backwards compat) - ModifyResponseException returns 200 with provider-agnostic error - UnifiedLLMGuardrails resolves call_type from logging_obj for pass-through - ModifyResponseException re-export from custom_guardrail stays in sync
…n fallback path (#25888)
…#26122) tool_calls on assistant messages were translated to OllamaToolCall format but never copied into the outgoing OllamaChatCompletionMessage, so Ollama received {role: assistant, content: ''} with no tool_calls. The model then had no record of having made a tool call, causing it to re-issue the identical call on every turn (infinite loop). Similarly, tool_call_id on role:tool messages was silently dropped. Ollama uses this field to resolve the tool name from conversation history. Also add tool_call_id to OllamaChatCompletionMessage TypedDict. Fixes #26094
litellm oss branch
* Use auth key name if there are no app id in in headers or in extra_data * use key alias instead of key name * Fix * last priority key alias * Fix * Add tests * [Feat] Day-0 support for GPT-5.5 and GPT-5.5 Pro (#26449) * feat(openai): day-0 support for GPT-5.5 and GPT-5.5 Pro Add pricing + capability entries for the new GPT-5.5 family launched by OpenAI on 2026-04-24: - gpt-5.5 / gpt-5.5-2026-04-23 (chat): $5/$30/$0.50 per 1M input/output/cached input - gpt-5.5-pro / gpt-5.5-pro-2026-04-23 (responses-only): $60/$360/$6 per 1M input/output/cached input Other fees (long-context >272k, flex, batches, priority, cache discounts) follow the same ratios as GPT-5.4, with context window retained at 1.05M input / 128K output. No transformation / classifier code changes are required: OpenAIGPT5Config.is_model_gpt_5_4_plus_model() already matches 5.5+ via numeric version parsing, and model registration is driven from the JSON. The existing responses-API bridge for tools + reasoning_effort (litellm/main.py:970) already covers gpt-5.5-pro. Tests: - GPT5_MODELS regression list now covers gpt-5.5-pro and dated variants - New test_generic_cost_per_token_gpt55_pro cost-calc test - Updated test_generic_cost_per_token_gpt55 for long-context fields * fix(openai): mirror reasoning_effort flags onto gpt-5.5 dated variants gpt-5.5-2026-04-23 and gpt-5.5-pro-2026-04-23 were missing the supports_none_reasoning_effort, supports_xhigh_reasoning_effort, and supports_minimal_reasoning_effort flags that their non-dated counterparts define. Reasoning-effort routing in OpenAIGPT5Config is fully capability-driven from these JSON flags — since an absent flag is treated as False for opt-in levels (xhigh), users pinning to a dated snapshot would silently lose xhigh support and diverge from the base alias on logprobs + flexible temperature handling. Copy the flags onto both dated variants so every dated snapshot inherits the base model's reasoning-effort capability profile. Adds a parametrized regression test that asserts supports_{none,minimal,xhigh}_reasoning_effort parity between each dated variant and its non-dated counterpart, preventing future drift when new snapshots are added. * [Feat] Add azure/gpt-5.5 + azure/gpt-5.5-pro entries (+ dated variants) (#26361) * feat(azure): add azure/gpt-5.5 + azure/gpt-5.5-pro entries (+ dated variants) Azure variants of OpenAI's GPT-5.5 family. Microsoft has not yet shipped GPT-5.5 on Azure OpenAI (latest GA on the Foundry models page is GPT-5.4 as of 2026-04-24), but adding the entries day-0 mirrors the established precedent for azure/gpt-5.4* (which were in the cost map before the Azure rollout) so cost tracking and capability flags work the moment customers deploy. Schema follows the existing azure/gpt-5.4* shape: - Same base/long-context pricing as openai/gpt-5.5*: $5/$30 chat, $60/$360 pro per 1M, with priority tier 2x base - Azure variants drop the flex/batches keys (Azure has no flex tier) but keep priority pricing, matching gpt-5.4* precedent - mode=chat for the thinking model, mode=responses for pro reasoning_effort capability flags mirror the OpenAI variants exactly since Azure proxies the same API contract: minimal rejection on both chat and pro, low/none rejection on pro. Once #26456 (which sets supports_low_reasoning_effort + minimal=false on openai/gpt-5.5*) lands, OpenAI and Azure flag profiles align. Tests pin entry presence + pricing for all four Azure variants and verify the live-API-derived reasoning_effort flags. * test: register supports_low_reasoning_effort in cost-map JSON schema azure/gpt-5.5-pro and azure/gpt-5.5-pro-2026-04-23 added in this branch carry supports_low_reasoning_effort=false. The strict 'additionalProperties: false' schema in test_aaamodel_prices_and_context_window_json_is_valid rejected the new key. Register it alongside the other supports_*_reasoning_effort entries. Note: the runtime side of this flag (code that reads it) lands in #26456. Until that PR merges the flag is inert for both Azure and OpenAI pro entries, but having the schema accept it lets cost-map tests pass on either merge order. * Use sanitize deep copy style to replace deepcopy usage * Added test checking error is not happening anymore * Added warning log when json copy failed * Reduce to one change * Fix spaces --------- Co-authored-by: Ido Lavi <ido@noma.security> Co-authored-by: yuneng-jiang <yuneng@berri.ai> Co-authored-by: Mateo Wang <277851410+mateo-berri@users.noreply.github.com> Co-authored-by: TomAlon <tom@noma.security>
…tch-422 fix(ui): use stored-credentials endpoint for tools fetch on MCP edit page
… triage Adds a CLI flag (`--timeout_worker_healthcheck`, env `TIMEOUT_WORKER_HEALTHCHECK`) that forwards to uvicorn's `timeout_worker_healthcheck` Config kwarg (added in uvicorn 0.37.0). Lets operators raise the supervisor's worker-ping timeout above the default 5s when triaging workers being killed and respawned under load. The helper introspects `uvicorn.Config.__init__` and only sets the kwarg if supported, otherwise prints a warning - so the existing uvicorn>=0.32.1,<1.0.0 floor pin is unaffected. Gunicorn and Hypercorn paths are unchanged (the uvicorn supervisor isn't running there); the value is also not passed to the helper at all on those paths so the "uvicorn too old" warning never fires spuriously.
…itellm_fix-logging-settings-admin-only
…lthcheck-flag feat(proxy): add --timeout_worker_healthcheck flag for uvicorn worker triage
fix(ci): support CircleCI rerun failed tests for local_testing jobs
…backs Switch the spend-logs save flow from mutateAsync + try/catch to mutate + callbacks. Errors now surface through a single onError path (no more double toast on failure), and the delete-then-update sequencing runs through onSettled instead of awaited promises. handleFormSubmit is no longer async. Tighten the corresponding test to assert exactly one error toast fires.
Previously, useStoreRequestInSpendLogs and useDeleteProxyConfigField did not refresh the proxyConfig cache on success, so the Logging Settings form continued to render the pre-save values until React Query refetched on its own. Wire both hooks to invalidate proxyConfigKeys on success so any active observer (currently the Logging Settings page) repulls fresh data. Export proxyConfigKeys for cross-hook reuse.
…dmin-only fix(ui): move 'Store Prompts in Spend Logs' toggle to Admin Settings
convert_anyof_null_to_nullable was stripping the items field from array
branches inside anyOf when a sibling null branch was present, leaving
{"type": "array"} without items. Vertex requires items whenever
type == "array" (even inside anyOf) and rejects the call with
INVALID_ARGUMENT.
Leave the (possibly empty) items in place so the downstream process_items
step can convert {} to {"type": "object"}, which is what Vertex wants.
Also:
- Update test_build_vertex_schema expected output, which was codifying
the broken shape.
- Convert test_gemini_tool_calling_not_working to a hermetic mock test
that asserts the request body sent to Vertex includes items inside
the callbacks anyOf array branch. The previous form made a real
network call and was flaky in CI.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…round-trip (#26653) * fix(caching): preserve prompt_tokens_details through embedding cache round-trip The embedding caching layer was dropping prompt_tokens_details (including image_count) because CachedEmbedding had no field for usage metadata and the cache retrieval code reconstructed Usage without it. This caused inconsistent responses where the first call returned image_count but cached responses did not, breaking cost tracking for multimodal embeddings. Add prompt_tokens_details to CachedEmbedding, persist per-item details during cache storage, aggregate them on retrieval, and merge them in combine_usage() for partial cache hits. * style: apply Black formatting to caching files * fix(caching): address Greptile review — cyclic import, guarded construction, nested dict merge Move PromptTokensDetailsWrapper to inline import to resolve CodeQL cyclic import warning. Guard PromptTokensDetailsWrapper construction with try/except to handle unexpected cached keys. Add recursive dict merging in _merge_prompt_tokens_details for nested fields like cache_creation_token_details.
* Add retry settings for generic API logger Made-with: Cursor * Refine generic API retry behavior Made-with: Cursor
* fix(logging): backfill streaming hidden response cost Made-with: Cursor * fix(logging): avoid mutating streaming hidden params Backfill calculated streaming response cost into logging payload copies so OTEL spans expose hidden_params.response_cost without mutating the response object. Made-with: Cursor * fix black formatting Apply the repo-pinned Black 24.10.0 formatting expected by CI. Made-with: Cursor * fix(types): allow numeric hidden response cost Allow standard logging hidden params to carry numeric response_cost values, matching LiteLLM's calculated cost payloads. Made-with: Cursor * refactor(logging): simplify hidden response cost backfill Clean up metadata initialization and reuse the raw response cost when deciding whether to backfill hidden params. Made-with: Cursor
Cache provider config lookups for Vertex Anthropic messages so repeated requests reuse the same config object and preserve credential cache state. Add a regression test to catch any future loss of config reuse. Made-with: Cursor
Companion to the prior commit. process_items only converted empty
`items: {}` to `{"type": "object"}`. But anyOf branches like
`{"type": "array"}` (no items field at all) were untouched, so after
convert_anyof_null_to_nullable stripped the null branch and added
nullable, the array branch was sent to Vertex as
`{"type": "array", "nullable": true}` — which Vertex rejects with
INVALID_ARGUMENT (`any_of[0].items: missing field`).
Make process_items synthesize `items: {"type": "object"}` for any
`type == "array"` schema where items is missing or empty.
Also:
- Convert test_gemini_tool_calling_working_demo to a hermetic mock
test asserting items is present on the array branch in the sent
body. Was previously a real-network call to Vertex and was the
test the user reported still failing in CI.
- Add unit test test_build_vertex_schema_array_branch_missing_items_in_anyof
covering the missing-items shape directly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
fix(vertex): preserve items on array branches in anyOf with null + de-flake test
AWS Bedrock has reached end-of-life for `claude-3-7-sonnet-20250219-v1:0`, returning 404s with "This model version has reached the end of its life." Update test references to `claude-sonnet-4-5-20250929-v1:0` (same capability surface: thinking, tools, prompt caching, PDF input, vision, computer use). The bedrock/invoke pass-through tests stay on Sonnet 3.5 since Sonnet 4.5 is converse-only on Bedrock.
Claude 3.5 Sonnet v2 reached EOL on Bedrock 2026-03-01, returning the same 404 EOL error as 3.7 Sonnet. Sonnet 4.5 supports both InvokeModel and Converse APIs on Bedrock, so use the same model for both routes.
…-model fix(tests): replace deprecated Bedrock Claude 3.7 Sonnet model ID
[Fix] Cache LiteLLM_Config param reads in DualCache and batch
Co-authored-by: Michael Riad Zaky <michaelr@Mac.localdomain>
[Infra] Version Bump
|
Michael Riad Zaky seems not to be a GitHub user. You need a GitHub account to be able to sign the CLA. If you have already a GitHub account, please add the email address used for this commit to your account. You have signed the CLA already but the status is still pending? Let us recheck it. |
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
Greptile SummaryThis infrastructure promotion merges internal staging into Confidence Score: 4/5Safe to merge; all findings are P2 style/reliability issues with no blocking correctness bugs. No P0/P1 issues found. Three P2 findings: silent KeyError drop in XecGuard logging_only mode, inconsistent HTTP 200 for guardrail blocks on pass-through endpoints, and a narrow multi-pod cache-miss staleness window in the new config cache layer. litellm/proxy/guardrails/guardrail_hooks/xecguard/xecguard.py (logging_only reliability), litellm/proxy/pass_through_endpoints/pass_through_endpoints.py (status code inconsistency)
|
| Filename | Overview |
|---|---|
| litellm/proxy/guardrails/guardrail_hooks/xecguard/xecguard.py | New XecGuard guardrail integration; raises HTTPException(400) to block requests but silently drops log entries in logging_only mode when standard_logging_object is absent from kwargs |
| litellm/proxy/pass_through_endpoints/pass_through_endpoints.py | Adds post-call guardrail execution and ModifyResponseException handling; returns HTTP 200 for blocked responses, which is inconsistent with the 400 returned by apply_guardrail |
| litellm/proxy/utils.py | Adds DualCache-backed config param caching with prefetch and invalidation; cache-miss sentinels stored in-memory may suppress fresh reads for up to TTL seconds in multi-pod scenarios |
| litellm/_redis_credential_provider.py | Adds module-level token cache with double-checked locking (threading.Lock) for GCP IAM tokens, preventing thundering-herd on pool warm-up; pattern is correct |
| litellm/llms/predibase/chat/transformation.py | Migrates Predibase response parsing, URL building, and request transformation out of handler.py into the standard BaseConfig pattern; logic is preserved faithfully |
| litellm/proxy/common_utils/expired_ui_session_key_cleanup_manager.py | New background job that deletes expired UI session keys; uses PodLockManager correctly to coordinate across pods |
| litellm/exceptions.py | Moves ModifyResponseException from integrations/custom_guardrail.py to the canonical exceptions module; no logic changes |
| litellm/proxy/proxy_server.py | Wires config cache into startup, switches all LiteLLM_Config reads to get_config_param/prefetch_config_params, and adds expired UI session key cleanup job |
| litellm/caching/caching_handler.py | Aggregates prompt_tokens_details from cached embedding responses; logic correctly handles numeric fields |
Sequence Diagram
sequenceDiagram
participant Client
participant PassThrough as pass_through_endpoints
participant GuardrailHook as XecGuardGuardrail
participant XecGuardAPI as XecGuard API
participant LLM as Upstream LLM
Client->>PassThrough: POST /passthrough/{path}
PassThrough->>LLM: Forward request (pre-call guardrails checked)
LLM-->>PassThrough: Raw response
PassThrough->>PassThrough: post_call_success_hook (if guardrails_to_run)
PassThrough->>GuardrailHook: apply_guardrail(input_type=response)
GuardrailHook->>XecGuardAPI: POST /xecguard/v1/scan
XecGuardAPI-->>GuardrailHook: decision SAFE or UNSAFE
alt UNSAFE raises HTTPException(400)
GuardrailHook-->>PassThrough: HTTPException caught as ModifyResponseException
PassThrough-->>Client: HTTP 200 + error body
else SAFE
GuardrailHook-->>PassThrough: modified or original response_body
PassThrough-->>Client: HTTP 200 + response content
end
Comments Outside Diff (1)
-
litellm/proxy/utils.py, line 2520-2537 (link)Config cache misses not invalidated by
prefetch_config_paramswhen a param is later writtenprefetch_config_paramsstores_CONFIG_CACHE_MISSfor any param not yet in the database. If that param is created (viainsert_data) shortly after the prefetch, the cache entry is invalidated correctly. However, if the write happens on a different pod (before the cache TTL of 60 s expires), the local pod's in-memory_CONFIG_CACHE_MISSentry will suppress the read from Redis for the full TTL duration, silently ignoring the new config.
Reviews (1): Last reviewed commit: "Merge pull request #26728 from BerriAI/y..." | Re-trigger Greptile
| "guardrail_name": "xecguard", | ||
| "guardrail_response": scan_result, | ||
| "guardrail_status": guardrail_status, | ||
| "masked_entity_count": None, | ||
| "start_time": start_time.timestamp(), | ||
| } | ||
|
|
||
| except Exception as exc: | ||
| verbose_proxy_logger.debug( | ||
| "XecGuard logging_only swallowed exception: %s", | ||
| str(exc), | ||
| ) | ||
| return kwargs, result | ||
|
|
||
| def logging_hook( |
There was a problem hiding this comment.
Silent KeyError in
async_logging_hook drops guardrail log entries
kwargs["standard_logging_object"] is accessed without first checking that the key exists. When this path runs in logging_only mode, if standard_logging_object is absent from kwargs (e.g. on older SDK callers or non-standard invocation paths), the KeyError is swallowed by the broad except Exception and the scan result is never written to the downstream observability pipeline — making logging_only a silent no-op in those cases.
| try: | ||
| await proxy_logging_obj.post_call_failure_hook( | ||
| user_api_key_dict=user_api_key_dict, | ||
| original_exception=e, | ||
| request_data=e.request_data, | ||
| ) | ||
| except Exception: | ||
| verbose_proxy_logger.warning( | ||
| "pass_through_endpoint: post_call_failure_hook raised during guardrail block", | ||
| exc_info=True, | ||
| ) | ||
| error_body = { | ||
| "error": { | ||
| "message": e.message or "Response blocked by guardrail", | ||
| "type": "content_filter", |
There was a problem hiding this comment.
HTTP 200 returned for guardrail-blocked pass-through responses
When a guardrail raises ModifyResponseException, the handler returns a 200 OK with a structured {"error": {...}} body. Clients that rely on HTTP status codes to detect blocked or filtered responses (e.g. the Anthropic SDK, httpx retry logic, or upstream load-balancers) will not recognize this as an error. Returning 400 or 403 (as apply_guardrail does via HTTPException) would make the signal consistent with the rest of the guardrail stack.
| ## INIT PROXY REDIS USAGE CLIENT ## | ||
| redis_usage_cache = litellm.cache.cache | ||
| spend_counter_cache.redis_cache = redis_usage_cache | ||
| litellm_config_cache.redis_cache = redis_usage_cache |
| @@ -2929,8 +2933,13 @@ def _init_cache( | |||
| ## INIT PROXY REDIS USAGE CLIENT ## | |||
| @@ -2929,8 +2933,13 @@ def _init_cache( | |||
| hash_password, | ||
| hash_token, | ||
| invalidate_config_param, | ||
| litellm_config_cache, |
| self.guardrail_name = guardrail_name | ||
| self.detection_info = detection_info or {} | ||
| super().__init__(message) | ||
| from litellm.exceptions import ModifyResponseException as ModifyResponseException |
|
|
||
| if TYPE_CHECKING: | ||
| from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj | ||
| from litellm.types.utils import PromptTokensDetailsWrapper |
| custom_llm_provider: Optional field - custom LLM provider for the endpoint | ||
| guardrails_config: Optional field - guardrails configuration for passthrough endpoint | ||
| """ | ||
| from litellm.exceptions import ModifyResponseException |
| ) | ||
| from litellm.types.utils import LiteLLMLoggingBaseClass | ||
| from litellm.utils import Choices, CustomStreamWrapper, Message, ModelResponse, Usage | ||
| from litellm.utils import CustomStreamWrapper, ModelResponse |
| def add_cache_point_tool_block( | ||
| tool: dict, model: Optional[str] = None | ||
| ) -> Optional[BedrockToolBlock]: | ||
| from litellm.llms.bedrock.common_utils import is_claude_4_5_on_bedrock |
[Infra] Promote Internal Staging to main
Relevant issues
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
tests/test_litellm/directory, Adding at least 1 test is a hard requirement - see detailsmake test-unit@greptileaiand received a Confidence Score of at least 4/5 before requesting a maintainer reviewDelays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
CI (LiteLLM team)
Branch creation CI run
Link:
CI run for the last commit
Link:
Merge / cherry-pick CI run
Links:
Screenshots / Proof of Fix
Type
🚄 Infrastructure
Changes