Repository navigation
feat(guardrails): add Airia guardrail integration - #40784
airia-open-source wants to merge 3 commits into
Conversation
226f0a4 to
612ca2c
Compare
|
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 3 · PR risk: 0/10 |
46d65e4 to
7babb5b
Compare
|
@greptileai please re-review at 7babb5b: fixes the AttributeError on a non-dict body, the rewrite length-mismatch, replaces the flag-only streaming test with one that drives the real streaming iterator and checks client-visible output, and switches the HTTP client to constructor-based dependency injection. The tool-call streaming timing gap and the tool-call-only response skip are analyzed in detail in the PR description under Caveats; no fix is available from this guardrail alone without regressing streamed redaction. |
7babb5b to
899e719
Compare
|
@yassin-berriai @yuneng-berri @yucheng-berri Can you help us review this PR, to add Airia as a Guardrail integration. We've addressed the previous review feedback and verified that all tests pass locally. Our Docs PR is linked here: BerriAI/litellm-docs#1427 |
|
@greptileai review latest head |
|
bugbot run |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Want reviews to match your repository better? Bugbot Learning can learn team-specific rules from PR activity. A team admin can enable Learning in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 899e719. Configure here.
| """ | ||
| if rewrite is None: | ||
| return True | ||
| return isinstance(rewrite, list) and len(rewrite) == len(original or ()) |
There was a problem hiding this comment.
Intervened empty rewrite skips fail-closed
Low Severity
An INTERVENED verdict can be accepted without rewriting any field that was actually sent. _rewritten only fail-closes when both texts and structured_messages are missing, while _is_applicable_rewrite treats an empty list as a valid match for an omitted field. Post-call never sends structured_messages, so a body like structured_messages: [] with texts omitted returns the original payload unchanged instead of blocking.
Reviewed by Cursor Bugbot for commit 899e719. Configure here.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Co-authored-by: Airia Open Source <opensource@airia.com> Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Hi @airia-open-source, can you send me the relevant longstanding credentials e.g api key for airia guardrail through email so that I can verify this PR e2e? Also we require them for maintaining the integration going forward. |
Adds `airia` as a built-in guardrail provider. The hook sends the prompt under `pre_call` and the model's response under `post_call` to the Airia AI Gateway, which answers allow, block, or redacted content; the only configuration is `api_base`, `api_key`, and an optional `timeout`. `during_call` is not offered: it runs concurrently with the model call, so a block could land after the prompt has already reached the provider. A BLOCKED verdict, or any action this version does not recognise, raises with `blocked_content=True`; a transport error, a non-2xx, or a response body that is not a JSON object raises with `blocked_content=False`, so callers can tell "could not evaluate" from "evaluated and blocked" while both still fail closed. On GUARDRAIL_INTERVENED a copy of the inputs is returned with every rewritten field substituted; an intervention carrying no applicable rewrite, a rewrite of the wrong shape, or a rewrite whose item count does not match what was sent (which would misalign the positional write-back downstream) blocks instead of letting the original through. Streamed responses are moderated whole and then emitted redacted: the hook opts into the unified hook's `incremental_diff` mode (the default `block_only` drops rewrites) with `streaming_end_of_stream_only`, so a redaction can never span transform rounds and underflow. Covered by a test that drives the real streaming iterator hook end to end and checks the assembled client output, not just the guardrail's own config attributes. The HTTP client is dependency-injected (an `async_handler` constructor parameter, defaulting to a real client) rather than mutated post-construction, so tests inject a fake client instead of monkeypatching an instance method. Known limitation, not fixable from this guardrail alone: LiteLLM's unified guardrail framework yields a tool-call chunk to the client immediately during `incremental_diff` streaming, inspecting the assembled tool calls against the guardrail's verdict only once the stream ends (`_run_incremental_transform_stream` in `unified_guardrail.py`). A blocked tool call can therefore reach an executing client before the terminal error arrives. Switching to `block_only` with `streaming_buffer_until_moderated` does not close this without a regression: that flag has no effect under `incremental_diff` (the mode branches away before it is read), and `block_only` itself replays the original, unredacted chunks on release, silently dropping any REDACT verdict on streamed text. Relatedly, LiteLLM's OpenAI chat completions guardrail translation skips `post_call` guardrails entirely for a response with no text content, so a tool-call-only response bypasses this guardrail (and every other post-call guardrail) even outside streaming. Both are shared-framework behavior, not specific to Airia; fixing either is out of scope for an isolated single-guardrail contribution.
… not sent Bugbot finding on BerriAI#40784. An INTERVENED verdict carrying `structured_messages: []` with `texts` omitted passed `_rewritten` unchanged: an empty list has the same length as an omitted field, so it counted as an applied rewrite while the `texts` that actually carried the content went through untouched. post_call sends only `texts`, so this was reachable. A rewrite for a field the request did not send is now a mismatch, same as a wrong length, and the verdict blocks. Regression test covers both the empty and the non-empty case. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
899e719 to
4747b6a
Compare
|
Rebased onto
Locally on this head: 34/34 in @greptileai please re-review at 4747b6a. |
…write cannot reach Veria finding on BerriAI#40784. incremental_diff, the mode that carries a rewrite into the stream, exists only for OpenAI chat completions; UnifiedLLMGuardrails sends every other streamed route (/v1/responses, /v1/messages) to block_only, which streams the original text live and drops rewrites. So with post_call on, a streamed Responses or Anthropic call could receive text that Airia then blocked or redacted. Two changes, both confined to this guardrail: - streaming_buffer_until_moderated is set, so on the block_only fallback the framework holds every chunk until the end-of-stream verdict instead of releasing it live. The flag is never read on the incremental_diff path, so streamed chat completions are unchanged. - A REDACT verdict on a streamed response whose route cannot carry a rewrite is enforced as a block, since block_only would otherwise replay the original. The route check mirrors _resolve_transform_call_type so the two cannot disagree; an unresolvable route fails closed. Request-side rewrites and non-streamed calls are untouched. The end-to-end test drives the real UnifiedLLMGuardrails iterator over an Anthropic /v1/messages SSE stream and asserts not one original chunk is released for either verdict. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Streamed chat completions deliver redactions at end of stream; streamed Responses and Anthropic Messages calls cannot carry a rewrite, so the proxy holds the response and enforces a Redact outcome as a Block (BerriAI/litellm#40784 review). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
@greptileai please re-review at 74daaf4. Since 4747b6a: Veria's non-chat streaming finding is fixed (buffered block_only fallback, redact enforced as block where the stream cannot carry a rewrite), with an end-to-end Anthropic-stream test; Veria passes on this head. 46/46 in test_airia.py, 220/220 in the shared guardrail suites, ruff strict and basedpyright gates pass locally. The 20 workflow runs still await maintainer approval. |


TLDR
Problem this solves:
How it solves it:
airiaas a built-in guardrail provider, configured withapi_base+api_keypre_call), responses after it (post_call)incremental_diff+streaming_end_of_stream_only)User Flow
Before: a proxy admin who wants Airia guardrails cannot reference the provider by name
guardrail: airiaunderguardrails:inconfig.yamland start the proxySkipping guardrail 'airia': invalid configuration ... Unsupported guardrail: airiahttp://litellm-domain/v1/chat/completionsreaches the model uncheckedAfter: the same config enforces the Airia policy on every call
guardrail: airiaunderguardrails:inconfig.yamland start the proxy/v1/chat/completionswith"my aws key is AKIAIOSFODNN7EXAMPLE"returns400—Guardrail raised an exception, Guardrail: airia, Message: <the message configured on their Airia guardrail>. The model is never called"Contact me at ada@example.com"returns200; the model receivedContact me at [EmailAddress1]200unchangedRelevant issues
—
Pre-Submission checklist
uv run pytest tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py -v→ 21 passedScreenshots / Proof of Fix
Setup shared by both runs.
config.yaml:The guardrail side is a live Airia deployment, not a mock. The model is a local echo endpoint because that is the only way to show what the provider actually received; nothing in the guardrail path is stubbed.
Before (a426dc4)
litellm --config config.yaml --port 4096After (899e719, unit-tested; see Caveats for what changed since the run below)
(rebased onto the current base after the run below; the tree is byte-identical to the commit it was captured on)
Blocked — secret in the prompt
HTTP 400{"error": {"message": "Guardrail raised an exception, Guardrail: airia, Message: Secret blocked by A51-317 demo policy.", "type": "None", "param": "None", "code": "400"}}Redacted — PII in the prompt
HTTP 200,choices[0].message.content = "echo: Contact me at [EmailAddress1]"['Contact me at [EmailAddress1]']— the address never reached itAllowed — benign prompt
HTTP 200,choices[0].message.content = "echo: Explain recursion in one sentence."Redacted — PII in the model's response (
post_call)The echo endpoint answers the fixed prompt
please share the demo contact(no PII on the way in) withSure, reach me at ada@example.com or call 415-555-0199.HTTP 200,choices[0].message.content = "echo: Sure, reach me at [EmailAddress1] or call 415-555-0199."Redacted — PII in a streamed response (
post_call,stream: true)HTTP 200; the stream is held until the response is complete and Airia has answered, then one delta carriesecho: Sure, reach me at [EmailAddress1] or call 415-555-0199.grep -c 'ada@example.com'over the raw SSE bytes:0Type
🆕 New Feature
Caveats (if any)
Medium
_run_incremental_transform_streaminunified_guardrail.pyyields a passthrough chunk for each tool-call delta as it arrives and only inspects the assembled tool calls against the guardrail's verdict once the stream ends, so a blocked tool call can be executed by the client before the terminal error lands. This is shared framework behavior, not specific to Airia: every guardrail usingincremental_diffstreaming has the same exposure. I looked for a way to close it from this guardrail alone and could not find one that does not regress the redaction path:streaming_buffer_until_moderated(which withholds every chunk until the full response is moderated) has no effect underincremental_diff— the mode branches away and returns before that flag is ever read — and switching toblock_onlyto make it apply would silently drop REDACT verdicts on streamed text, replaying the original unredacted content on release. Closing this needs a change to the shared streaming framework (buffering tool-call chunks specifically, independent of text transform mode), which is out of scope for an isolated single-guardrail PROpenAIChatCompletionsHandler.process_output_responseskips everypost_callguardrail entirely when a response has no text content (if not self._has_text_content(response): ... return response), so a tool-call-only response with no text bypasses Airia's response-side check regardless of streaming. Also shared framework behavior; also out of scope hereLow
Streamed text arrives in one delta at end of stream rather than incrementally: the trade for redactions that can never underflow. Blocks still terminate the stream immediately
The HTTP client is dependency-injected (
async_handlerconstructor param) rather than mutated after construction, so tests build a fake client instead of monkeypatching an instance methodduring_callis deliberately not offered; it runs concurrently with the model call, so a block could land after the prompt reached the providerThe proof's model is a local echo endpoint; the guardrail path is a live Airia deployment. No provider key is available to this machine for a paid-model run
If Airia is unreachable the request is refused (
blocked_content=False), not passed through — same fail-closed stance as the other guardrail providersDocs page will follow in a separate PR to
litellm-docsBerriAI/litellm retired
litellm_internal_stagingas the trunk while this PR was open (main is now the sole default branch); rebased onto currentmainwith no conflicts, and re-verified against the strict-ruleset rules that were added upstream in the same window (test return-type annotations, bannedtyping.cast,**kwargstyping). No behavior change, no budget-JSON editsFinal Attestation
Note
Medium Risk
New code sits on the proxy request path and can block or rewrite all traffic when enabled; fail-closed behavior on gateway errors denies requests rather than bypassing policy.
Overview
Adds
airiaas a first-class LiteLLM proxy guardrail so admins can setguardrail: airiain config instead of shipping a custom callback module.AiriaGuardrailPOSTs request/response content to the Airia Gateway (/v1/guardrails/litellm) onpre_callandpost_callonly. Gateway verdicts map to allow (NONE), block (BLOCKED+ optional reason), or redact (GUARDRAIL_INTERVENEDwith rewrittentexts/structured_messages). Streaming usesincremental_diffwithstreaming_end_of_stream_onlyso redacted text is applied before the client sees the full stream.Safety-oriented behavior is new in this hook: malformed or non-object JSON, unknown actions, and rewrites with wrong list lengths fail closed as blocks (or evaluation errors with
blocked_content=Falsewhen the gateway is unreachable). Credentials and timeout resolve from config orAIRIA_*env vars.Registration follows the existing pattern (
initialize_guardrail, initializer/class registries,AiriaGuardrailConfigModelfor the UI). A largetest_airia.pysuite covers payload forwarding, intervention edge cases, fail-closed paths,init_guardrails_v2wiring, and end-to-end streaming redaction viaUnifiedLLMGuardrails.Reviewed by Cursor Bugbot for commit 899e719. Bugbot is set up for automated code reviews on this repo. Configure here.