Skip to content

feat(guardrails): honor litellm_params.timeout in every HTTP guardrail - #43134

Merged
kerry-berri merged 9 commits into
mainfrom
litellm_guardrail_timeout_all_providers
Oct 1, 2026
Merged

kerry-berri merged 9 commits into
mainfrom
litellm_guardrail_timeout_all_providers

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • litellm_params.timeout exists but most guardrails ignore it
  • Guardrail vendor calls fall back to the shared client's 600s default
  • A hung moderation or guardrail vendor stalls every guarded request for minutes

How it solves it:

  • CustomGuardrail stores timeout only when it is configured
  • Every HTTP guardrail passes timeout=self.timeout on its outbound call
  • HiddenLayer SaaS JWT fetch and refresh use the same timeout
  • Providers with their own default keep it when timeout is unset
  • Rubrik log shipping keeps its own timeout, only its moderation call is bounded
  • One integration test boots a real proxy with 45 guardrails against a hanging sink

Intentional product change: a guardrail that already had timeout in its config now actually stops waiting at that value, so a stale low value starts failing checks that used to finish

User Flow

Before: a proxy admin sets timeout: 2 on an openai_moderation guardrail, and requests still hang while the moderation endpoint is down

  1. Admin adds timeout: 2 under the guardrail's litellm_params and restarts the proxy
  2. A user sends POST https://litellm-domain/v1/chat/completions with "guardrails": ["moderation-stalled-timeout-2"]
  3. The moderation endpoint never answers, and the user waits 30s (up to 600s) before getting a 500

After: the same request fails fast at the configured limit

  1. Admin adds timeout: 2 under the guardrail's litellm_params and restarts the proxy
  2. A user sends POST https://litellm-domain/v1/chat/completions with "guardrails": ["moderation-stalled-timeout-2"]
  3. After about 2s the user gets a 408 litellm.Timeout: Connection timed out. Timeout passed=2.0

Admin UI Playground at http://localhost:4000/ui, model gpt-5.4-mini, guardrail moderation-stalled-timeout-2, prompt "Say hi in one word". The Playground client retries the error, so the times shown are the sum of its attempts (see Caveats)

Before, error shown after 91.50s (three 30s attempts):

Before: Playground validation error after the 30s stall

After, 408 timeout shown after 17.71s (each attempt cut at 2.0s):

After: Playground timeout error

Linear ticket

Resolves LIT-8333

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Both legs ran at the same time against one stalled sink, real OpenAI gpt-5.4-mini, real omni-moderation for the control cases, and a real Postgres per proxy. Base is the merge base in a worktree on port 4001, head is the PR tip on port 4002

Sink (sink.py, every POST sleeps 30s and then returns {}):

import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

class H(BaseHTTPRequestHandler):
    protocol_version = "HTTP/1.1"
    def do_POST(self):
        self.rfile.read(int(self.headers.get("content-length", "0")))
        time.sleep(30)
        body = b"{}"
        self.send_response(200); self.send_header("content-type", "application/json")
        self.send_header("content-length", str(len(body))); self.end_headers(); self.wfile.write(body)
    def log_message(self, *a): pass

class S(ThreadingHTTPServer):
    daemon_threads = True
    allow_reuse_address = True

S(("127.0.0.1", 9911), H).serve_forever()

Proxy config (same file for both legs):

model_list:
  - model_name: gpt-5.4-mini
    litellm_params:
      model: openai/gpt-5.4-mini
      api_key: os.environ/OPENAI_API_KEY
guardrails:
  - guardrail_name: moderation-stalled-timeout-2
    litellm_params: {guardrail: openai_moderation, mode: pre_call, api_base: http://127.0.0.1:9911/v1, api_key: sk-synthetic, timeout: 2}
  - guardrail_name: moderation-real-timeout-5
    litellm_params: {guardrail: openai_moderation, mode: pre_call, api_key: os.environ/OPENAI_API_KEY, timeout: 5}
  - guardrail_name: moderation-real-unset
    litellm_params: {guardrail: openai_moderation, mode: pre_call, api_key: os.environ/OPENAI_API_KEY}
  - guardrail_name: cisco-stalled-unset
    litellm_params: {guardrail: cisco_ai_defense, mode: pre_call, api_base: http://127.0.0.1:9911, api_key: sk-synthetic}
  - guardrail_name: pillar-stalled-unset
    litellm_params: {guardrail: pillar, mode: pre_call, api_base: http://127.0.0.1:9911, api_key: sk-synthetic}

Each case ran curl -s -w "%{http_code} %{time_total}" localhost:<port>/v1/chat/completions -H "Authorization: Bearer $MASTER_KEY" -d '{"model":"gpt-5.4-mini","messages":[{"role":"user","content":"Say hi in one word"}],"guardrails":["<case>"]}'

Before (f8f0576)

moderation-stalled-timeout-2

  1. Ran the curl above on port 4001
  2. HTTP 500 in 30.04s {"error":{"message":"3 validation errors for OpenAIModerationResponse\nid\n Field required ..., the configured 2s is ignored

moderation-real-timeout-5

  1. Same curl with this guardrail
  2. HTTP 200 in 0.63s, model replied "Hi"

moderation-real-unset

  1. Same curl with this guardrail
  2. HTTP 200 in 1.12s, model replied "Hi"

cisco-stalled-unset

  1. Same curl with this guardrail
  2. HTTP 503 in 10.03s Cisco AI Defense guardrail unavailable, its own 10s default

pillar-stalled-unset

  1. Same curl with this guardrail
  2. HTTP 200 in 5.45s, its own 5s default, then fails open to the model

Playground

  1. Steps and screenshot in User Flow, error shown after 91.50s

After (75e7d7c)

moderation-stalled-timeout-2

  1. Ran the curl above on port 4002
  2. HTTP 408 in 2.03s {"error":{"message":"litellm.Timeout: Connection timed out. Timeout passed=2.0, time taken=2.002 seconds","code":"408"}}

moderation-real-timeout-5

  1. Same curl with this guardrail
  2. HTTP 200 in 0.74s, model replied "Hi", a healthy vendor is unaffected by the limit

moderation-real-unset

  1. Same curl with this guardrail
  2. HTTP 200 in 0.63s, model replied "Hi"

cisco-stalled-unset

  1. Same curl with this guardrail
  2. HTTP 503 in 10.03s Cisco AI Defense guardrail unavailable, default kept when timeout is unset

pillar-stalled-unset

  1. Same curl with this guardrail
  2. HTTP 200 in 5.47s, default kept when timeout is unset

Playground

  1. Steps and screenshot in User Flow, 408 shown after 17.71s, each server-side attempt cut at 2.0s

Test notes: tests/integration/observability/test_guardrail_timeout_all_providers.py boots a real proxy on Postgres with 45 guardrail cases pointed at one sink that sleeps 20s, each with timeout: 1, and asserts each round trip ends under 8s. All probes go out concurrently, so the file runs in about 26s instead of 72s. model_armor gets a synthetic service account whose token URL is the same sink. A control guardrail with no timeout must still wait for the sink's 3s reply. It fails 34 of 45 on the merge base and passes 45 of 45 at the tip. Mutation checks: making CustomGuardrail assign timeout unconditionally fails test_unset_timeout_keeps_default_assigned_before_super_init, and passing timeout=self.timeout to Rubrik's log flush fails test_flush_queue_does_not_inherit_guardrail_timeout

Type

🆕 New Feature
✅ Test

Caveats (if any)

Medium

  • Clients that retry a 408 multiply the wait
    • The OpenAI SDK retries twice, the Playground showed the error at 17.71s with timeout: 2
    • Lower timeout or set max_retries=0 on the client
  • Clearing timeout through a live guardrail update drops cisco, pillar, singulr, straiker and panw to the 600s client default until restart. Same as main
  • microsoft_purview passes the timeout but has no integration case, since its login and Graph URLs are fixed and cannot point at a sink

Low

  • grayswan and akto keep their own guardrail_timeout param and agent_365 keeps its own request_timeout, not wired here
  • The timeout bounds each HTTP call, so a guardrail making several calls can run longer in total
  • Rubrik log-flush behavior is covered by a unit test, not a live run
  • Docs live in the litellm-docs repo and are not part of this PR
  • CI: documentation, code-quality, misc, proxy-endpoints and proxy-behavior also fail on main at ed4caeb, and rust-wheel fails on main at 6fd9334 on a stub mismatch. This PR touches none of those areas

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

ran /live-pr-risk and found no regressions/backward incompatible risks

REVIEWER MUST KNOW BEFORE APPROVING

  • A guardrail with timeout in its config now stops its vendor call at that value
    • Before: HTTP 500 in 30.04s against a 30s stall with timeout: 2
    • After: HTTP 408 in 2.03s with litellm.Timeout: Connection timed out. Timeout passed=2.0
    • A config carrying a stale low timeout now fails checks that used to finish
    • Approved by kerry-berri, who asked for a guardrail timeout and picked this PR's approach in the linked Devin session
  • Providers that had a fixed timeout constant (compresr, lasso, promptguard, typesafe, xecguard and similar) now use the configured timeout when set and keep the constant when unset
    • Before: the constant applied whatever timeout said
    • After: the configured value wins, unset keeps the constant (cisco 503 in 10.03s and pillar 200 in 5.47s on both legs)
    • Approved by kerry-berri in the linked Devin session
  • HiddenLayer SaaS token fetch at startup and token refresh use the configured timeout
    • Before: fixed 30s on the token request regardless of config
    • After: bounded by timeout when set, unchanged when unset. Covered by unit tests, not a live HiddenLayer run
    • Approved by kerry-berri in the linked Devin session
  • Rubrik moderation calls use the configured timeout, its batch log shipping does not
    • Before: neither call used it
    • After: moderation bounded, log flush keeps the client default (unit test, not live)
    • Approved by kerry-berri, who chose this PR after the log-shipping fix was proposed in the linked Devin session
  • Guardrails with no timeout behave the same as main
    • Before: moderation-real-unset HTTP 200 in 1.12s
    • After: HTTP 200 in 0.63s, same reply, no limit applied
    • No sign-off needed, unchanged

Link to Devin session: https://app.devin.ai/sessions/ddf36d5cf5434de19370ceb10b5249b2
Open in Devin Desktop: https://app.devin.ai/desktop/session/ddf36d5cf5434de19370ceb10b5249b2?variant=devin
Requested by: @kerry-berri


Note

Medium Risk
Wide guardrail surface area: configs with timeout now fail fast instead of waiting on the shared client default, which can change production pass/fail behavior for stale low values.

Overview
Adds a shared timeout on CustomGuardrail (from guardrail litellm_params) and threads it through proxy guardrail initializers so outbound vendor HTTP calls pass timeout=self.timeout.

When timeout is unset, integrations keep their existing per-provider defaults (Compresr, Lasso, Cisco, etc.) by assigning self.timeout after super().__init__. Presidio maps numeric timeouts to aiohttp; HiddenLayer bounds SaaS JWT fetch/refresh; MCP JWT signer applies it to JWKS/OIDC/introspection; Rubrik moderation is bounded but log batch flush does not inherit it.

Adds a broad integration test (real proxy + slow sink) and unit tests for constructor/update behavior and mock fixes for timeout kwargs.

Reviewed by Cursor Bugbot for commit eadb7ca. Bugbot is set up for automated code reviews on this repo. Configure here.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 25, 2026 04:26
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Adds timeout parameter to guardrail HTTP calls.

The PR appears safe to merge; no outstanding finding or new actionable issue was identified.

Summary

The PR forwards configured guardrail timeouts to outbound HTTP calls while retaining provider defaults when no timeout is set. The latest change tightens sink-call attribution in the integration test so sibling guardrails cannot satisfy each other’s call assertions.

Reviews (10) · Last reviewed commit: "test(guardrails): match sink calls to th..."

Comment thread litellm/proxy/guardrails/guardrail_hooks/hiddenlayer/__init__.py

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@codecov

codecov Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 87.50000% with 4 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...rdrails/guardrail_hooks/hiddenlayer/hiddenlayer.py 50.00% 2 Missing ⚠️
...s/guardrail_hooks/mcp_jwt_signer/mcp_jwt_signer.py 85.71% 1 Missing ⚠️
...ardrail_hooks/zscaler_ai_guard/zscaler_ai_guard.py 0.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_guardrail_timeout_all_providers (eadb7ca) with main (6997223)

Open in CodSpeed

…ler post stubs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

…eout, drop akto from timeout coverage

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

… cast

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread tests/integration/observability/test_guardrail_timeout_all_providers.py Outdated
kerry and others added 2 commits September 30, 2026 23:38
…eout_all_providers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

# Conflicts:
#	litellm/proxy/guardrails/guardrail_initializers.py
…only rubrik moderation calls

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit eadb7ca. Configure here.

@kerry-berri
kerry-berri merged commit a308a8e into main Oct 1, 2026
98 of 104 checks passed
jan-sauer-reef added a commit to jan-sauer-reef/litellm that referenced this pull request Oct 1, 2026
…ject_key_prefix

* upstream/main: (62 commits)
  fix(guardrails): scan Responses API input in Azure Prompt Shield (BerriAI#43786)
  feat(lens): investigate sampled traces and retain batch results (BerriAI#43942)
  fix(proxy): restore pre-config-wins handling of pass-through endpoints (BerriAI#43962)
  fix(cost-map): raise baseten DeepSeek-V4.1-Flash max output to 262144 (BerriAI#43916)
  chore(cost-map): add deprecation date for anthropic claude-sonnet-4-5 (BerriAI#43898)
  chore(cost-map): add fireworks inkling priority prices from the prices api (BerriAI#43949)
  feat(guardrails): honor litellm_params.timeout in every HTTP guardrail (BerriAI#43134)
  test(e2e): typed per-test metadata for the e2e suite (BerriAI#42044)
  fix(caching): write the response-cache SET to Redis at once instead of on the post-call batch (BerriAI#43973)
  feat(ui): filter tags by name and description on the Tag Management page (BerriAI#42949)
  feat(providers): add Cortecs as an OpenAI-compatible provider (BerriAI#43872)
  feat(e2e): record each e2e test's steps, starting with ProxyClient (BerriAI#42393)
  test(ci): repair stale tests and move retired OpenAI text-completion fixtures (BerriAI#43958)
  feat(proxy): record in spend logs whether a request used a client-forwarded Anthropic OAuth token (BerriAI#43063)
  fix(azure_storage): keep the DataLakeServiceClient alive until its TTL elapses (BerriAI#43082)
  chore(deps): bump gitpython and tornado, extend diskcache osv ignore to Nov 1 (BerriAI#43961)
  fix(guardrails): treat an unknown straiker api_version as unset instead of skipping the guardrail (BerriAI#43956)
  fix(azure_storage): name Data Lake objects without base64 padding or slashes (BerriAI#43914)
  fix(grayswan): send request conversation and tool calls to post-call monitor (BerriAI#43770)
  chore(cost-map): sync openrouter prices from the models API (BerriAI#43950)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants