Repository navigation
feat(guardrails): honor litellm_params.timeout in every HTTP guardrail - #43134
Conversation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
|
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
…ler post stubs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…eout, drop akto from timeout coverage Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… cast Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…eout_all_providers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> # Conflicts: # litellm/proxy/guardrails/guardrail_initializers.py
…only rubrik moderation calls Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit eadb7ca. Configure here.
…ject_key_prefix * upstream/main: (62 commits) fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786) feat(lens): investigate sampled traces and retain batch results (BerriAI#43942) fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962) fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916) chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898) chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949) feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134) test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044) fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973) feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949) feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872) feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393) test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958) feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063) fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082) chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961) fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956) fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914) fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770) chore(cost-map): sync openrouter prices from the models API (BerriAI#43950) ...
TLDR
Problem this solves:
litellm_params.timeoutexists but most guardrails ignore itHow it solves it:
CustomGuardrailstorestimeoutonly when it is configuredtimeout=self.timeouton its outbound calltimeoutis unsetIntentional product change: a guardrail that already had
timeoutin its config now actually stops waiting at that value, so a stale low value starts failing checks that used to finishUser Flow
Before: a proxy admin sets
timeout: 2on anopenai_moderationguardrail, and requests still hang while the moderation endpoint is downtimeout: 2under the guardrail'slitellm_paramsand restarts the proxy"guardrails": ["moderation-stalled-timeout-2"]After: the same request fails fast at the configured limit
timeout: 2under the guardrail'slitellm_paramsand restarts the proxy"guardrails": ["moderation-stalled-timeout-2"]litellm.Timeout: Connection timed out. Timeout passed=2.0Admin UI Playground at http://localhost:4000/ui, model
gpt-5.4-mini, guardrailmoderation-stalled-timeout-2, prompt "Say hi in one word". The Playground client retries the error, so the times shown are the sum of its attempts (see Caveats)Before, error shown after 91.50s (three 30s attempts):
After, 408 timeout shown after 17.71s (each attempt cut at 2.0s):
Linear ticket
Resolves LIT-8333
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Both legs ran at the same time against one stalled sink, real OpenAI
gpt-5.4-mini, realomni-moderationfor the control cases, and a real Postgres per proxy. Base is the merge base in a worktree on port 4001, head is the PR tip on port 4002Sink (
sink.py, every POST sleeps 30s and then returns{}):Proxy config (same file for both legs):
Each case ran
curl -s -w "%{http_code} %{time_total}" localhost:<port>/v1/chat/completions -H "Authorization: Bearer $MASTER_KEY" -d '{"model":"gpt-5.4-mini","messages":[{"role":"user","content":"Say hi in one word"}],"guardrails":["<case>"]}'Before (f8f0576)
moderation-stalled-timeout-2
HTTP 500 in 30.04s{"error":{"message":"3 validation errors for OpenAIModerationResponse\nid\n Field required ..., the configured 2s is ignoredmoderation-real-timeout-5
HTTP 200 in 0.63s, model replied "Hi"moderation-real-unset
HTTP 200 in 1.12s, model replied "Hi"cisco-stalled-unset
HTTP 503 in 10.03sCisco AI Defense guardrail unavailable, its own 10s defaultpillar-stalled-unset
HTTP 200 in 5.45s, its own 5s default, then fails open to the modelPlayground
After (75e7d7c)
moderation-stalled-timeout-2
HTTP 408 in 2.03s{"error":{"message":"litellm.Timeout: Connection timed out. Timeout passed=2.0, time taken=2.002 seconds","code":"408"}}moderation-real-timeout-5
HTTP 200 in 0.74s, model replied "Hi", a healthy vendor is unaffected by the limitmoderation-real-unset
HTTP 200 in 0.63s, model replied "Hi"cisco-stalled-unset
HTTP 503 in 10.03sCisco AI Defense guardrail unavailable, default kept whentimeoutis unsetpillar-stalled-unset
HTTP 200 in 5.47s, default kept whentimeoutis unsetPlayground
Test notes:
tests/integration/observability/test_guardrail_timeout_all_providers.pyboots a real proxy on Postgres with 45 guardrail cases pointed at one sink that sleeps 20s, each withtimeout: 1, and asserts each round trip ends under 8s. All probes go out concurrently, so the file runs in about 26s instead of 72s. model_armor gets a synthetic service account whose token URL is the same sink. A control guardrail with notimeoutmust still wait for the sink's 3s reply. It fails 34 of 45 on the merge base and passes 45 of 45 at the tip. Mutation checks: makingCustomGuardrailassigntimeoutunconditionally failstest_unset_timeout_keeps_default_assigned_before_super_init, and passingtimeout=self.timeoutto Rubrik's log flush failstest_flush_queue_does_not_inherit_guardrail_timeoutType
🆕 New Feature
✅ Test
Caveats (if any)
Medium
timeout: 2timeoutor setmax_retries=0on the clienttimeoutthrough a live guardrail update drops cisco, pillar, singulr, straiker and panw to the 600s client default until restart. Same as mainLow
guardrail_timeoutparam and agent_365 keeps its ownrequest_timeout, not wired hereFinal Attestation
ran /live-pr-risk and found no regressions/backward incompatible risks
REVIEWER MUST KNOW BEFORE APPROVING
timeoutin its config now stops its vendor call at that valueHTTP 500 in 30.04sagainst a 30s stall withtimeout: 2HTTP 408 in 2.03swithlitellm.Timeout: Connection timed out. Timeout passed=2.0timeoutnow fails checks that used to finishtimeoutwhen set and keep the constant when unsettimeoutsaidtimeouttimeoutwhen set, unchanged when unset. Covered by unit tests, not a live HiddenLayer runtimeout, its batch log shipping does nottimeoutbehave the same as mainmoderation-real-unsetHTTP 200 in 1.12sLink to Devin session: https://app.devin.ai/sessions/ddf36d5cf5434de19370ceb10b5249b2
Open in Devin Desktop: https://app.devin.ai/desktop/session/ddf36d5cf5434de19370ceb10b5249b2?variant=devin
Requested by: @kerry-berri
Note
Medium Risk
Wide guardrail surface area: configs with
timeoutnow fail fast instead of waiting on the shared client default, which can change production pass/fail behavior for stale low values.Overview
Adds a shared
timeoutonCustomGuardrail(from guardraillitellm_params) and threads it through proxy guardrail initializers so outbound vendor HTTP calls passtimeout=self.timeout.When
timeoutis unset, integrations keep their existing per-provider defaults (Compresr, Lasso, Cisco, etc.) by assigningself.timeoutaftersuper().__init__. Presidio maps numeric timeouts to aiohttp; HiddenLayer bounds SaaS JWT fetch/refresh; MCP JWT signer applies it to JWKS/OIDC/introspection; Rubrik moderation is bounded but log batch flush does not inherit it.Adds a broad integration test (real proxy + slow sink) and unit tests for constructor/update behavior and mock fixes for
timeoutkwargs.Reviewed by Cursor Bugbot for commit eadb7ca. Bugbot is set up for automated code reviews on this repo. Configure here.