Repository navigation
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
With fire_and_forget enabled the guardrail POST runs as a detached task and apply_guardrail returns the inputs unchanged right away, so the guardrail becomes observe-only. A dispatched call is recorded as success with a response saying the verdict was not read. At most fire_and_forget_max_inflight calls (default 100) are in flight at once per guardrail and worker. Extra calls are dropped, counted with a rate-limited warning, and recorded as guardrail_status not_run. The background POST has a fixed 30 second timeout and starts from an empty context, so it does not keep request-scoped context vars alive Failures in the background call are logged and never raised. A failure while building the payload before dispatch is logged, passes the request through regardless of fail_on_error, and is recorded as not_run with its own reason. A non-bool fire_and_forget or a non-int fire_and_forget_max_inflight is rejected at startup. Enabling it also forces streaming_end_of_stream_only so a stream sends one call
6169d6b to
73904c9
Compare
Parse fire_and_forget and fire_and_forget_max_inflight the way pydantic parses config values, so "true" and "5" work. A value that cannot be read is ignored with a warning: fire_and_forget falls back to false, so the guardrail keeps enforcing, and fire_and_forget_max_inflight falls back to 100 Force streaming_transform_mode to block_only under fire_and_forget. With incremental_diff the unified guardrail buffered the stream and dispatched one call per chunk, so a stream now sends a single end-of-stream call while the chunks reach the client live The background POST now honors the guardrail's timeout and falls back to 30 seconds when it is unset. It is built from the same URL, headers and payload as the awaited call. A payload that cannot be serialized is recorded as not_run before anything is dispatched Tests that build GenericGuardrailAPI move to a mirror test_generic_guardrail_api.py, the dispatcher and parser tests stay in test_background_dispatch.py, and a shared conftest.py captures proxy warnings
|
@greptileai please re-review the latest commit, which hardens config parsing, forces block_only streaming and honors the guardrail timeout for background calls |
|
…free BackgroundDispatcher.dispatch now takes a prepare callable and calls it only after the in-flight cap check passes. A call dropped by fire_and_forget_max_inflight no longer serializes its payload on the request path, which matters most for large or image-heavy requests while the endpoint is slow A payload that cannot be serialized is still recorded as not_run before anything is dispatched
…ails With fire_and_forget the response audit ran only after a stream finished, so a stream the client hung up on, or one that failed upstream, never reached the audit endpoint Guardrails can now set streaming_observe_only. UnifiedLLMGuardrails then forwards every chunk untouched and, when the stream closes for any reason, runs one check over a copy of what reached the client. The check is shielded from cancellation, and a failure in it is logged and never breaks the stream. GenericGuardrailAPI sets the flag from fire_and_forget instead of forcing end-of-stream and block_only, which the new path makes unnecessary. Other guardrails are unchanged
|
@greptileai please re-review: abandoned streams are now observed, dropped calls skip serialization, and the PR body has a fresh live proof |
TLDR
Problem this solves:
How it solves it:
fire_and_forgetoption sends the guardrail call in the backgroundfire_and_forget_max_inflightcaps background calls and drops the extra onesRelevant issues
Split out of #37055. Docs: BerriAI/litellm-docs#2047
User Flow
Before: an admin who only wants to audit traffic makes every request wait on the audit endpoint, and its answer can still block
generic_guardrail_apiguardrail at an audit endpoint that takes 5 seconds to answerAfter: the same request goes straight to the model while the audit endpoint still receives every call
fire_and_forget: trueto the guardrail, or sets it in the Add Guardrail form at https://litellm-domain/ui/?page=guardrails, and the boot log warns it is observe-onlyTo reproduce the screenshots, open http://localhost:4000/ui/?page=guardrails, click Add New Guardrail, then Add Provider Guardrail, type any name, pick Generic Guardrail API as the provider, click Next and scroll to the end of the form. For the logs, run case 3 below, open http://localhost:4000/ui/?page=logs and click each of the three requests
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
successandnot_run, returning empty inputs, removing the observe-only warning, not forwarding either option from the config or the max_inflight to the dispatcher, letting an unreadable value turn fire_and_forget on, rejecting "true", raising on or ignoring max_inflight, acceptingtrueas a max_inflight, an off-by-one cap, never freeing a slot, letting a background failure escape, running the task in the request's context, warning on every drop, and building the payload before the cap checkuv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or morebudget-ratchetfails because main dropped the LIT002 budget in chore(lint): remove the LIT002 mutable-construction rule #43971 after this branch's merge base, andosv-scanflags abracesadvisory in the dashboard's package-lock.json. This PR touches neither budgets nor dependenciesmake checkpasses locally withBASE_REF=upstream/main. A fork whosemainlags upstream fails the budget gates against that stale base, so setBASE_REFwhen checking a forkcodecov/patchpasses, and Codecov reports every modified line as covered@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
All runs are against a live proxy on
localhost:4000started withpython litellm/proxy/proxy_cli.py --config proof_config.yaml --detailed_debug --use_v2_migration_resolver, backed by a real Postgres, and every completion is a real, billed call to Bedrockus.anthropic.claude-haiku-4-5-20251001-v1:0.$PROOF_KEYis a virtual key fromPOST /key/generatewith access tobedrock-haiku-4-5. Non-streaming requests setmetadata.include_guardrail_response, so the response lists the guardrail entries recorded for that call. Each request picks its guardrail with"guardrails": [...]At the merge base
fire_and_forgetandfire_and_forget_max_inflightdo not exist, so the proxy ignores them and awaits every guardrail call. The audit endpoint is a small FastAPI app on127.0.0.1:8787that waits 5 seconds, logs what it received, and answers BLOCKED for everything.audit-downpoints at a port where nothing listens, andaudit-postonly checks responses. Nothing else is stubbedThe Admin UI screenshots for the same runs are in User Flow above
Before (6f123b7)
Case 1: the audit endpoint takes 5 s and answers BLOCKED
Case 2: the audit endpoint still receives every call with its api_key
Case 3: fire_and_forget_max_inflight=1 with three requests at once
Case 4: the audit endpoint is down
Case 5: streaming, time to first byte and total time
Case 6: streaming with streaming_transform_mode: incremental_diff
Case 7: the client hangs up mid-stream
Case 8: startup warning
After (de597a5)
Case 1: the audit endpoint takes 5 s and answers BLOCKED
Case 2: the audit endpoint still receives every call with its api_key
Case 3: fire_and_forget_max_inflight=1 with three requests at once
Case 4: the audit endpoint is down
Case 5: streaming, time to first byte and total time
Case 6: streaming with streaming_transform_mode: incremental_diff
Case 7: the client hangs up mid-stream
Case 8: startup warning
Type
🆕 New Feature
Caveats (if any)
Severe
fire_and_forgeton, the guardrail never blocks or rewrites anythingfail_on_errorandunreachable_fallbackcan no longer fail the requestMedium
fire_and_forget_max_inflightLow
streaming_*settingsstreaming_observe_onlypath to the shared streaming hooktimeout, or 30 seconds when it is unsetfire_and_forgetis ignored with a warning and keeps enforcingfire_and_forget_max_inflightbehaves differently by config pathlitellm_call_id, input type and guardrail nametest_generic_guardrail_api.py, so the second to land merges both by handFinal Attestation