Repository navigation
feat(guardrails): add Airia guardrail integration - #39418
rchawla-airia wants to merge 1 commit into
Conversation
4a36824 to
f697903
Compare
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
0f23c1f to
549827b
Compare
Greptile SummaryThe PR adds Airia as a built-in pre-call and post-call guardrail, including fail-closed remote evaluation, response rewriting, configuration registration, and whole-response streamed redaction.
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains.
|
| Filename | Overview |
|---|---|
| litellm/proxy/guardrails/guardrail_hooks/airia/airia.py | Implements Airia evaluation, fail-closed decision handling, content rewrites, and end-of-stream moderation. |
| litellm/proxy/guardrails/guardrail_hooks/airia/init.py | Registers and initializes the Airia callback with configured credentials, modes, and timeout. |
| litellm/types/proxy/guardrails/guardrail_hooks/airia.py | Defines the proxy configuration schema for the Airia gateway URL, API key, and timeout. |
| tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py | Adds mocked coverage of Airia initialization, request payloads, decisions, rewrites, failures, and streaming configuration. |
Reviews (3): Last reviewed commit: "feat(guardrails): add Airia guardrail in..." | Re-trigger Greptile
PR overviewThis pull request adds Airia guardrail integration to LiteLLM, including moderation of unified streaming responses and assembled tool calls. One issue remains after another was addressed. In incremental streaming mode, tool-call chunks can reach an auto-executing client before Airia returns its moderation verdict, allowing a prohibited call to run under that specific configuration. Buffering tool calls until approval or disabling incremental delivery for them would close the remaining gap. Open issues (1)
Fixed/addressed: 1 · PR risk: 4/10 |
549827b to
f2f7a0b
Compare
|
@greptileai please re-review: f2f7a0b addresses the streaming rewrite mode, timeout passthrough, comment conventions, and fail-closed handling of unapplied rewrites |
|
|
||
| self.api_base = (api_base or os.getenv("AIRIA_GATEWAY_URL", "")).rstrip("/") | ||
| self.api_key = api_key or os.getenv("AIRIA_API_KEY") | ||
| self.streaming_transform_mode: Final[Literal["block_only", "incremental_diff"]] = "incremental_diff" |
There was a problem hiding this comment.
Low: Streamed tool calls are released before moderation
Enabling incremental_diff sends tool-call chunks to the client immediately; the unified streaming hook only submits the assembled tool call for its block decision at end of stream. An attacker can prompt the model to emit a prohibited tool call that reaches an auto-executing client before Airia's block verdict. Buffer tool-call chunks until the complete call passes Airia, or disable this streaming mode when tool calls are present.
There was a problem hiding this comment.
Framework behavior of incremental mode, now called out under Caveats; buffering tool-call chunks belongs in the unified hook as a follow-up
Adds `airia` as a built-in guardrail provider. Airia evaluates the prompt under `pre_call` and the model's response under `post_call`, returning allow, block, or redacted content; the tenant is resolved from the API key, so the only configuration is `api_base` and `api_key`. `during_call` is not offered: it runs concurrently with the model call, so a block could land after the prompt has already reached the provider. A BLOCKED verdict, or any action this version does not recognise, raises with `blocked_content=True`; a transport error or non-2xx raises with `blocked_content=False`, so callers can tell "could not evaluate" from "evaluated and blocked" while both still fail closed. On GUARDRAIL_INTERVENED a copy of the inputs is returned with every rewritten field substituted; an intervention carrying no applicable rewrite, or a rewrite of the wrong shape, blocks instead of letting the original through. Streamed responses are moderated whole and then emitted redacted: the hook opts into the unified hook's `incremental_diff` mode (the default `block_only` drops rewrites) with `streaming_end_of_stream_only`, so a redaction can never span transform rounds and underflow.
f2f7a0b to
5e60018
Compare
|
@greptileai please re-review: 5e60018 moderates streamed responses whole via streaming_end_of_stream_only, so rewrites cannot underflow; tool-call passthrough is documented under Caveats |
|
Closing this draft for now while we settle which contributor account submits it; the change will be resubmitted unchanged once that is sorted |
|
Resubmitted from the Airia open source account as #40784 with the same change; this closed draft is superseded |
TLDR
Problem this solves:
How it solves it:
airiaas a built-in guardrail provider, configured withapi_base+api_keypre_call), responses after it (post_call)incremental_diff+streaming_end_of_stream_only)User Flow
Before: a proxy admin who wants Airia guardrails cannot reference the provider by name
guardrail: airiaunderguardrails:inconfig.yamland start the proxySkipping guardrail 'airia': invalid configuration ... Unsupported guardrail: airiahttp://litellm-domain/v1/chat/completionsreaches the model uncheckedAfter: the same config enforces the Airia policy on every call
guardrail: airiaunderguardrails:inconfig.yamland start the proxy/v1/chat/completionswith"my aws key is AKIAIOSFODNN7EXAMPLE"returns400—Guardrail raised an exception, Guardrail: airia, Message: <the message configured on their Airia guardrail>. The model is never called"Contact me at ada@example.com"returns200; the model receivedContact me at [EmailAddress1]200unchangedRelevant issues
—
Pre-Submission checklist
uv run pytest tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py -v→ 21 passedScreenshots / Proof of Fix
Setup shared by both runs.
config.yaml:The guardrail side is a live Airia deployment, not a mock. The model is a local echo endpoint because that is the only way to show what the provider actually received; nothing in the guardrail path is stubbed.
Before (b4f5b6a)
litellm --config config.yaml --port 4096After (5e60018)
Blocked — secret in the prompt
HTTP 400{"error": {"message": "Guardrail raised an exception, Guardrail: airia, Message: Secret blocked by A51-317 demo policy.", "type": "None", "param": "None", "code": "400"}}Redacted — PII in the prompt
HTTP 200,choices[0].message.content = "echo: Contact me at [EmailAddress1]"['Contact me at [EmailAddress1]']— the address never reached itAllowed — benign prompt
HTTP 200,choices[0].message.content = "echo: Explain recursion in one sentence."Redacted — PII in the model's response (
post_call)The echo endpoint answers the fixed prompt
please share the demo contact(no PII on the way in) withSure, reach me at ada@example.com or call 415-555-0199.HTTP 200,choices[0].message.content = "echo: Sure, reach me at [EmailAddress1] or call 415-555-0199."Redacted — PII in a streamed response (
post_call,stream: true)HTTP 200; the stream is held until the response is complete and Airia has answered, then one delta carriesecho: Sure, reach me at [EmailAddress1] or call 415-555-0199.grep -c 'ada@example.com'over the raw SSE bytes:0Type
🆕 New Feature
Caveats (if any)
Medium
incremental_diffguardrail. A follow-up in the unified hook (buffer tool-call chunks until moderated) would close it for all providersLow
Streamed text arrives in one delta at end of stream rather than incrementally: the trade for redactions that can never underflow. Blocks still terminate the stream immediately
during_callis deliberately not offered; it runs concurrently with the model call, so a block could land after the prompt reached the providerThe proof's model is a local echo endpoint; the guardrail path is a live Airia deployment. No provider key is available to this machine for a paid-model run
If Airia is unreachable the request is refused (
blocked_content=False), not passed through — same fail-closed stance as the other guardrail providersDocs page will follow in a separate PR to
litellm-docsFinal Attestation