Skip to content

feat(guardrails): add Airia guardrail integration - #40784

Open
airia-open-source wants to merge 3 commits into
BerriAI:mainfrom
airia-open-source:airia_guardrail
Open

airia-open-source wants to merge 3 commits into
BerriAI:mainfrom
airia-open-source:airia_guardrail

Conversation

@airia-open-source

@airia-open-source airia-open-source commented Sep 11, 2026 •

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Airia customers on LiteLLM have no way to enforce Airia guardrails on proxy traffic
  • Today they must copy a Python file onto the proxy host and reference it by dotted path

How it solves it:

  • Adds airia as a built-in guardrail provider, configured with api_base + api_key
  • Prompts are checked before the model call (pre_call), responses after it (post_call)
  • Airia returns allow / block / redact; redactions are applied before the text moves on
  • Streamed responses are moderated whole, then delivered redacted (incremental_diff + streaming_end_of_stream_only)
  • A Gateway response that is not a JSON object, or a rewrite whose item count does not match what was sent, fails closed instead of erroring or misaligning the positional write-back

User Flow

Before: a proxy admin who wants Airia guardrails cannot reference the provider by name

  1. They add guardrail: airia under guardrails: in config.yaml and start the proxy
  2. Startup logs Skipping guardrail 'airia': invalid configuration ... Unsupported guardrail: airia
  3. Every POST to http://litellm-domain/v1/chat/completions reaches the model unchecked

After: the same config enforces the Airia policy on every call

  1. They add guardrail: airia under guardrails: in config.yaml and start the proxy
  2. The proxy starts with the guardrail active
  3. POST /v1/chat/completions with "my aws key is AKIAIOSFODNN7EXAMPLE" returns 400 — Guardrail raised an exception, Guardrail: airia, Message: <the message configured on their Airia guardrail>. The model is never called
  4. POST with "Contact me at ada@example.com" returns 200; the model received Contact me at [EmailAddress1]
  5. POST with a benign prompt returns 200 unchanged

Relevant issues

—

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally: uv run pytest tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py -v → 21 passed
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • Greptile Confidence Score of at least 4/5

Screenshots / Proof of Fix

Setup shared by both runs. config.yaml:

model_list:
  - model_name: gpt-4o-mini
    litellm_params:
      model: openai/harness-echo          # local echo endpoint that records exactly what it received,
      api_base: http://127.0.0.1:8931     # so the redacted text the MODEL saw is observable
      api_key: not-a-real-key
guardrails:
  - guardrail_name: "airia"
    litellm_params:
      guardrail: airia
      mode: ["pre_call", "post_call"]
      default_on: true
      api_base: os.environ/AIRIA_GATEWAY_URL   # a running Airia AI Gateway; policy: PII → redact, Secrets → block
      api_key: os.environ/AIRIA_API_KEY
general_settings:
  master_key: sk-1234

The guardrail side is a live Airia deployment, not a mock. The model is a local echo endpoint because that is the only way to show what the provider actually received; nothing in the guardrail path is stubbed.

Before (a426dc4)

  1. litellm --config config.yaml --port 4096
  2. Output:
    LiteLLM Proxy:ERROR: init_guardrails.py:37 - Skipping guardrail 'airia': invalid configuration, proxy is starting WITHOUT this guardrail: Unsupported guardrail: airia
    

After (899e719, unit-tested; see Caveats for what changed since the run below)

(rebased onto the current base after the run below; the tree is byte-identical to the commit it was captured on)

Blocked — secret in the prompt

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"my aws key is AKIAIOSFODNN7EXAMPLE"}]}'
  2. HTTP 400
    {"error": {"message": "Guardrail raised an exception, Guardrail: airia, Message: Secret blocked by A51-317 demo policy.", "type": "None", "param": "None", "code": "400"}}
  3. Echo endpoint call log: 0 calls for this request — the model was never invoked

Redacted — PII in the prompt

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Contact me at ada@example.com"}]}'
  2. HTTP 200, choices[0].message.content = "echo: Contact me at [EmailAddress1]"
  3. Echo endpoint call log shows the model received ['Contact me at [EmailAddress1]'] — the address never reached it

Allowed — benign prompt

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain recursion in one sentence."}]}'
  2. HTTP 200, choices[0].message.content = "echo: Explain recursion in one sentence."
  3. Echo endpoint call log shows the model received the prompt unchanged

Redacted — PII in the model's response (post_call)

The echo endpoint answers the fixed prompt please share the demo contact (no PII on the way in) with Sure, reach me at ada@example.com or call 415-555-0199.

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"please share the demo contact"}]}'
  2. HTTP 200, choices[0].message.content = "echo: Sure, reach me at [EmailAddress1] or call 415-555-0199."

Redacted — PII in a streamed response (post_call, stream: true)

  1. curl -sN http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","stream":true,"messages":[{"role":"user","content":"please share the demo contact"}]}'
  2. HTTP 200; the stream is held until the response is complete and Airia has answered, then one delta carries echo: Sure, reach me at [EmailAddress1] or call 415-555-0199.
  3. grep -c 'ada@example.com' over the raw SSE bytes: 0

Type

🆕 New Feature

Caveats (if any)

Medium

  • Streamed tool-call chunks reach the client before the end-of-stream verdict. _run_incremental_transform_stream in unified_guardrail.py yields a passthrough chunk for each tool-call delta as it arrives and only inspects the assembled tool calls against the guardrail's verdict once the stream ends, so a blocked tool call can be executed by the client before the terminal error lands. This is shared framework behavior, not specific to Airia: every guardrail using incremental_diff streaming has the same exposure. I looked for a way to close it from this guardrail alone and could not find one that does not regress the redaction path: streaming_buffer_until_moderated (which withholds every chunk until the full response is moderated) has no effect under incremental_diff — the mode branches away and returns before that flag is ever read — and switching to block_only to make it apply would silently drop REDACT verdicts on streamed text, replaying the original unredacted content on release. Closing this needs a change to the shared streaming framework (buffering tool-call chunks specifically, independent of text transform mode), which is out of scope for an isolated single-guardrail PR
  • Separately, and outside streaming too: OpenAIChatCompletionsHandler.process_output_response skips every post_call guardrail entirely when a response has no text content (if not self._has_text_content(response): ... return response), so a tool-call-only response with no text bypasses Airia's response-side check regardless of streaming. Also shared framework behavior; also out of scope here

Low

  • Streamed text arrives in one delta at end of stream rather than incrementally: the trade for redactions that can never underflow. Blocks still terminate the stream immediately

  • The HTTP client is dependency-injected (async_handler constructor param) rather than mutated after construction, so tests build a fake client instead of monkeypatching an instance method

  • during_call is deliberately not offered; it runs concurrently with the model call, so a block could land after the prompt reached the provider

  • The proof's model is a local echo endpoint; the guardrail path is a live Airia deployment. No provider key is available to this machine for a paid-model run

  • If Airia is unreachable the request is refused (blocked_content=False), not passed through — same fail-closed stance as the other guardrail providers

  • Docs page will follow in a separate PR to litellm-docs

  • BerriAI/litellm retired litellm_internal_staging as the trunk while this PR was open (main is now the sole default branch); rebased onto current main with no conflicts, and re-verified against the strict-ruleset rules that were added upstream in the same window (test return-type annotations, banned typing.cast, **kwargs typing). No behavior change, no budget-JSON edits

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Note

Medium Risk
New code sits on the proxy request path and can block or rewrite all traffic when enabled; fail-closed behavior on gateway errors denies requests rather than bypassing policy.

Overview
Adds airia as a first-class LiteLLM proxy guardrail so admins can set guardrail: airia in config instead of shipping a custom callback module.

AiriaGuardrail POSTs request/response content to the Airia Gateway (/v1/guardrails/litellm) on pre_call and post_call only. Gateway verdicts map to allow (NONE), block (BLOCKED + optional reason), or redact (GUARDRAIL_INTERVENED with rewritten texts / structured_messages). Streaming uses incremental_diff with streaming_end_of_stream_only so redacted text is applied before the client sees the full stream.

Safety-oriented behavior is new in this hook: malformed or non-object JSON, unknown actions, and rewrites with wrong list lengths fail closed as blocks (or evaluation errors with blocked_content=False when the gateway is unreachable). Credentials and timeout resolve from config or AIRIA_* env vars.

Registration follows the existing pattern (initialize_guardrail, initializer/class registries, AiriaGuardrailConfigModel for the UI). A large test_airia.py suite covers payload forwarding, intervention edge cases, fail-closed paths, init_guardrails_v2 wiring, and end-to-end streaming redaction via UnifiedLLMGuardrails.

Reviewed by Cursor Bugbot for commit 899e719. Bugbot is set up for automated code reviews on this repo. Configure here.

@CLAassistant

CLAassistant commented Sep 11, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@greptile-apps

greptile-apps Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Adds a new guardrail integration for content filtering.

The PR appears safe to merge, with the previously reported issues resolved, conceded, or manually closed and no new actionable defects found

Summary

Adds Airia as a built-in pre-call and post-call guardrail integration, including fail-closed response validation and redaction support

  • Registers Airia in guardrail initialization and configuration models
  • Sends normalized request and response content through LiteLLM's shared HTTP client
  • Applies validated positional rewrites and blocks malformed or unsupported interventions
  • Buffers non-chat streams until moderation and blocks redaction verdicts where the stream cannot carry rewritten content
  • Adds behavioral coverage for initialization, verdict handling, streaming chat redaction, and Anthropic stream blocking

Reviews (5) · Last reviewed commit: "fix(guardrails): fail closed on streamed..." · Reviewed by Greptile

Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py
Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py
Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py Outdated
Comment thread tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py Outdated
Comment thread tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py Outdated
Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py
Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py
@veria-ai

veria-ai Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 3 · PR risk: 0/10

@airia-open-source

Copy link
Copy Markdown
Author

@greptileai please re-review at 7babb5b: fixes the AttributeError on a non-dict body, the rewrite length-mismatch, replaces the flag-only streaming test with one that drives the real streaming iterator and checks client-visible output, and switches the HTTP client to constructor-based dependency injection. The tool-call streaming timing gap and the tool-call-only response skip are analyzed in detail in the PR description under Caveats; no fix is available from this guardrail alone without regressing streamed redaction.

@devin-ai-integration
devin-ai-integration Bot changed the base branch from litellm_internal_staging to main September 23, 2026 14:35
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 23, 2026 14:35
@airia-open-source

Copy link
Copy Markdown
Author

@yassin-berriai @yuneng-berri @yucheng-berri Can you help us review this PR, to add Airia as a Guardrail integration. We've addressed the previous review feedback and verified that all tests pass locally.

Our Docs PR is linked here: BerriAI/litellm-docs#1427

@yucheng-berri

Copy link
Copy Markdown
Contributor

@greptileai review latest head

@yucheng-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want reviews to match your repository better? Bugbot Learning can learn team-specific rules from PR activity. A team admin can enable Learning in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 899e719. Configure here.

"""
if rewrite is None:
return True
return isinstance(rewrite, list) and len(rewrite) == len(original or ())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intervened empty rewrite skips fail-closed

Low Severity

An INTERVENED verdict can be accepted without rewriting any field that was actually sent. _rewritten only fail-closes when both texts and structured_messages are missing, while _is_applicable_rewrite treats an empty list as a valid match for an omitted field. Post-call never sends structured_messages, so a body like structured_messages: [] with texts omitted returns the original payload unchanged instead of blocking.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 899e719. Configure here.

@codspeed

codspeed Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing airia-open-source:airia_guardrail (899e719) with main (6b642f3)

Open in CodSpeed

@codecov

codecov Bot commented Sep 28, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

devin-ai-integration Bot added a commit that referenced this pull request Sep 29, 2026
Co-authored-by: Airia Open Source <opensource@airia.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@yucheng-berri

Copy link
Copy Markdown
Contributor

Hi @airia-open-source, can you send me the relevant longstanding credentials e.g api key for airia guardrail through email so that I can verify this PR e2e? Also we require them for maintaining the integration going forward.

airia-open-source and others added 2 commits October 7, 2026 00:11
Adds `airia` as a built-in guardrail provider. The hook sends the prompt under `pre_call` and the
model's response under `post_call` to the Airia AI Gateway, which answers allow, block, or redacted
content; the only configuration is `api_base`, `api_key`, and an optional `timeout`.

`during_call` is not offered: it runs concurrently with the model call, so a block could land after
the prompt has already reached the provider.

A BLOCKED verdict, or any action this version does not recognise, raises with
`blocked_content=True`; a transport error, a non-2xx, or a response body that is not a JSON object
raises with `blocked_content=False`, so callers can tell "could not evaluate" from "evaluated and
blocked" while both still fail closed. On GUARDRAIL_INTERVENED a copy of the inputs is returned with
every rewritten field substituted; an intervention carrying no applicable rewrite, a rewrite of the
wrong shape, or a rewrite whose item count does not match what was sent (which would misalign the
positional write-back downstream) blocks instead of letting the original through.

Streamed responses are moderated whole and then emitted redacted: the hook opts into the unified
hook's `incremental_diff` mode (the default `block_only` drops rewrites) with
`streaming_end_of_stream_only`, so a redaction can never span transform rounds and underflow. Covered
by a test that drives the real streaming iterator hook end to end and checks the assembled client
output, not just the guardrail's own config attributes.

The HTTP client is dependency-injected (an `async_handler` constructor parameter, defaulting to a
real client) rather than mutated post-construction, so tests inject a fake client instead of
monkeypatching an instance method.

Known limitation, not fixable from this guardrail alone: LiteLLM's unified guardrail framework
yields a tool-call chunk to the client immediately during `incremental_diff` streaming, inspecting
the assembled tool calls against the guardrail's verdict only once the stream ends
(`_run_incremental_transform_stream` in `unified_guardrail.py`). A blocked tool call can therefore
reach an executing client before the terminal error arrives. Switching to `block_only` with
`streaming_buffer_until_moderated` does not close this without a regression: that flag has no effect
under `incremental_diff` (the mode branches away before it is read), and `block_only` itself replays
the original, unredacted chunks on release, silently dropping any REDACT verdict on streamed text.
Relatedly, LiteLLM's OpenAI chat completions guardrail translation skips `post_call` guardrails
entirely for a response with no text content, so a tool-call-only response bypasses this guardrail
(and every other post-call guardrail) even outside streaming. Both are shared-framework behavior,
not specific to Airia; fixing either is out of scope for an isolated single-guardrail contribution.
… not sent

Bugbot finding on BerriAI#40784. An INTERVENED verdict carrying `structured_messages: []` with
`texts` omitted passed `_rewritten` unchanged: an empty list has the same length as an
omitted field, so it counted as an applied rewrite while the `texts` that actually carried
the content went through untouched. post_call sends only `texts`, so this was reachable.

A rewrite for a field the request did not send is now a mismatch, same as a wrong length,
and the verdict blocks. Regression test covers both the empty and the non-empty case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@airia-open-source

Copy link
Copy Markdown
Author

Rebased onto main at 4747b6a and addressed the open findings:

Locally on this head: 34/34 in test_airia.py, 200/200 in the shared guardrail registry/init/endpoint tests, ruff check (default and strict), ruff format --check, and scripts/type_check_gate.py --base main all pass.

@greptileai please re-review at 4747b6a.

Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py
…write cannot reach

Veria finding on BerriAI#40784. incremental_diff, the mode that carries a rewrite into the stream,
exists only for OpenAI chat completions; UnifiedLLMGuardrails sends every other streamed route
(/v1/responses, /v1/messages) to block_only, which streams the original text live and drops
rewrites. So with post_call on, a streamed Responses or Anthropic call could receive text that
Airia then blocked or redacted.

Two changes, both confined to this guardrail:

- streaming_buffer_until_moderated is set, so on the block_only fallback the framework holds
  every chunk until the end-of-stream verdict instead of releasing it live. The flag is never
  read on the incremental_diff path, so streamed chat completions are unchanged.
- A REDACT verdict on a streamed response whose route cannot carry a rewrite is enforced as a
  block, since block_only would otherwise replay the original. The route check mirrors
  _resolve_transform_call_type so the two cannot disagree; an unresolvable route fails closed.
  Request-side rewrites and non-streamed calls are untouched.

The end-to-end test drives the real UnifiedLLMGuardrails iterator over an Anthropic
/v1/messages SSE stream and asserts not one original chunk is released for either verdict.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
airia-open-source added a commit to airia-open-source/litellm-docs that referenced this pull request Oct 7, 2026
Streamed chat completions deliver redactions at end of stream; streamed Responses and
Anthropic Messages calls cannot carry a rewrite, so the proxy holds the response and
enforces a Redact outcome as a Block (BerriAI/litellm#40784 review).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@airia-open-source

Copy link
Copy Markdown
Author

@greptileai please re-review at 74daaf4. Since 4747b6a: Veria's non-chat streaming finding is fixed (buffered block_only fallback, redact enforced as block where the stream cannot carry a rewrite), with an end-to-end Anthropic-stream test; Veria passes on this head. 46/46 in test_airia.py, 220/220 in the shared guardrail suites, ruff strict and basedpyright gates pass locally. The 20 workflow runs still await maintainer approval.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants