Skip to content

feat(guardrails): add Airia guardrail integration - #39418

Closed
rchawla-airia wants to merge 1 commit into
BerriAI:litellm_internal_stagingfrom
rchawla-airia:feature/airia-guardrail
Closed

rchawla-airia wants to merge 1 commit into
BerriAI:litellm_internal_stagingfrom
rchawla-airia:feature/airia-guardrail

Conversation

@rchawla-airia

@rchawla-airia rchawla-airia commented Sep 2, 2026 •

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Airia customers on LiteLLM have no way to enforce Airia guardrails on proxy traffic
  • Today they must copy a Python file onto the proxy host and reference it by dotted path

How it solves it:

  • Adds airia as a built-in guardrail provider, configured with api_base + api_key
  • Prompts are checked before the model call (pre_call), responses after it (post_call)
  • Airia returns allow / block / redact; redactions are applied before the text moves on
  • Streamed responses are moderated whole, then delivered redacted (incremental_diff + streaming_end_of_stream_only)

User Flow

Before: a proxy admin who wants Airia guardrails cannot reference the provider by name

  1. They add guardrail: airia under guardrails: in config.yaml and start the proxy
  2. Startup logs Skipping guardrail 'airia': invalid configuration ... Unsupported guardrail: airia
  3. Every POST to http://litellm-domain/v1/chat/completions reaches the model unchecked

After: the same config enforces the Airia policy on every call

  1. They add guardrail: airia under guardrails: in config.yaml and start the proxy
  2. The proxy starts with the guardrail active
  3. POST /v1/chat/completions with "my aws key is AKIAIOSFODNN7EXAMPLE" returns 400 — Guardrail raised an exception, Guardrail: airia, Message: <the message configured on their Airia guardrail>. The model is never called
  4. POST with "Contact me at ada@example.com" returns 200; the model received Contact me at [EmailAddress1]
  5. POST with a benign prompt returns 200 unchanged

Relevant issues

—

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally: uv run pytest tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py -v → 21 passed
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • Greptile Confidence Score of at least 4/5

Screenshots / Proof of Fix

Setup shared by both runs. config.yaml:

model_list:
  - model_name: gpt-4o-mini
    litellm_params:
      model: openai/harness-echo          # local echo endpoint that records exactly what it received,
      api_base: http://127.0.0.1:8931     # so the redacted text the MODEL saw is observable
      api_key: not-a-real-key
guardrails:
  - guardrail_name: "airia"
    litellm_params:
      guardrail: airia
      mode: ["pre_call", "post_call"]
      default_on: true
      api_base: os.environ/AIRIA_GATEWAY_URL   # a running Airia AI Gateway; policy: PII → redact, Secrets → block
      api_key: os.environ/AIRIA_API_KEY
general_settings:
  master_key: sk-1234

The guardrail side is a live Airia deployment, not a mock. The model is a local echo endpoint because that is the only way to show what the provider actually received; nothing in the guardrail path is stubbed.

Before (b4f5b6a)

  1. litellm --config config.yaml --port 4096
  2. Output:
    LiteLLM Proxy:ERROR: init_guardrails.py:37 - Skipping guardrail 'airia': invalid configuration, proxy is starting WITHOUT this guardrail: Unsupported guardrail: airia
    

After (5e60018)

Blocked — secret in the prompt

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"my aws key is AKIAIOSFODNN7EXAMPLE"}]}'
  2. HTTP 400
    {"error": {"message": "Guardrail raised an exception, Guardrail: airia, Message: Secret blocked by A51-317 demo policy.", "type": "None", "param": "None", "code": "400"}}
  3. Echo endpoint call log: 0 calls for this request — the model was never invoked

Redacted — PII in the prompt

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Contact me at ada@example.com"}]}'
  2. HTTP 200, choices[0].message.content = "echo: Contact me at [EmailAddress1]"
  3. Echo endpoint call log shows the model received ['Contact me at [EmailAddress1]'] — the address never reached it

Allowed — benign prompt

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain recursion in one sentence."}]}'
  2. HTTP 200, choices[0].message.content = "echo: Explain recursion in one sentence."
  3. Echo endpoint call log shows the model received the prompt unchanged

Redacted — PII in the model's response (post_call)

The echo endpoint answers the fixed prompt please share the demo contact (no PII on the way in) with Sure, reach me at ada@example.com or call 415-555-0199.

  1. curl -s http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"please share the demo contact"}]}'
  2. HTTP 200, choices[0].message.content = "echo: Sure, reach me at [EmailAddress1] or call 415-555-0199."

Redacted — PII in a streamed response (post_call, stream: true)

  1. curl -sN http://127.0.0.1:4097/v1/chat/completions -H 'Content-Type: application/json' -H 'Authorization: Bearer sk-1234' \
      -d '{"model":"gpt-4o-mini","stream":true,"messages":[{"role":"user","content":"please share the demo contact"}]}'
  2. HTTP 200; the stream is held until the response is complete and Airia has answered, then one delta carries echo: Sure, reach me at [EmailAddress1] or call 415-555-0199.
  3. grep -c 'ada@example.com' over the raw SSE bytes: 0

Type

🆕 New Feature

Caveats (if any)

Medium

  • Streamed tool-call chunks are forwarded before the end-of-stream verdict: the unified hook's incremental mode passes them through and inspects the assembled call only at the end. Framework behavior shared by every incremental_diff guardrail. A follow-up in the unified hook (buffer tool-call chunks until moderated) would close it for all providers

Low

  • Streamed text arrives in one delta at end of stream rather than incrementally: the trade for redactions that can never underflow. Blocks still terminate the stream immediately

  • during_call is deliberately not offered; it runs concurrently with the model call, so a block could land after the prompt reached the provider

  • The proof's model is a local echo endpoint; the guardrail path is a live Airia deployment. No provider key is available to this machine for a paid-model run

  • If Airia is unreachable the request is refused (blocked_content=False), not passed through — same fail-closed stance as the other guardrail providers

  • Docs page will follow in a separate PR to litellm-docs

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@rchawla-airia
rchawla-airia force-pushed the feature/airia-guardrail branch from 4a36824 to f697903 Compare September 2, 2026 22:09
@CLAassistant

CLAassistant commented Sep 2, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@rchawla-airia
rchawla-airia changed the base branch from main to litellm_internal_staging September 2, 2026 22:10
@codspeed

codspeed Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing rchawla-airia:feature/airia-guardrail (5e60018) with litellm_internal_staging (11a02b9)

Open in CodSpeed

@codecov

codecov Bot commented Sep 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@rchawla-airia
rchawla-airia force-pushed the feature/airia-guardrail branch 2 times, most recently from 0f23c1f to 549827b Compare September 2, 2026 22:28
@rchawla-airia
rchawla-airia marked this pull request as ready for review September 2, 2026 22:46
@greptile-apps

greptile-apps Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR adds Airia as a built-in pre-call and post-call guardrail, including fail-closed remote evaluation, response rewriting, configuration registration, and whole-response streamed redaction.

  • Registers the Airia integration and its configuration model.
  • Sends normalized request and response content to the Airia gateway and applies block or rewrite decisions.
  • Buffers streamed output until completion before applying and emitting redactions.
  • Adds mocked coverage for configuration, transport failures, decisions, rewrites, and streaming settings.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
litellm/proxy/guardrails/guardrail_hooks/airia/airia.py Implements Airia evaluation, fail-closed decision handling, content rewrites, and end-of-stream moderation.
litellm/proxy/guardrails/guardrail_hooks/airia/init.py Registers and initializes the Airia callback with configured credentials, modes, and timeout.
litellm/types/proxy/guardrails/guardrail_hooks/airia.py Defines the proxy configuration schema for the Airia gateway URL, API key, and timeout.
tests/test_litellm/proxy/guardrails/guardrail_hooks/test_airia.py Adds mocked coverage of Airia initialization, request payloads, decisions, rewrites, failures, and streaming configuration.

Reviews (3): Last reviewed commit: "feat(guardrails): add Airia guardrail in..." | Re-trigger Greptile

Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py
Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/__init__.py
Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py Outdated
Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py Outdated
@veria-ai

veria-ai Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

This pull request adds Airia guardrail integration to LiteLLM, including moderation of unified streaming responses and assembled tool calls.

One issue remains after another was addressed. In incremental streaming mode, tool-call chunks can reach an auto-executing client before Airia returns its moderation verdict, allowing a prohibited call to run under that specific configuration. Buffering tool calls until approval or disabling incremental delivery for them would close the remaining gap.

Open issues (1)

Fixed/addressed: 1 · PR risk: 4/10

@rchawla-airia
rchawla-airia force-pushed the feature/airia-guardrail branch from 549827b to f2f7a0b Compare September 3, 2026 13:08
@rchawla-airia

Copy link
Copy Markdown
Author

@greptileai please re-review: f2f7a0b addresses the streaming rewrite mode, timeout passthrough, comment conventions, and fail-closed handling of unapplied rewrites


self.api_base = (api_base or os.getenv("AIRIA_GATEWAY_URL", "")).rstrip("/")
self.api_key = api_key or os.getenv("AIRIA_API_KEY")
self.streaming_transform_mode: Final[Literal["block_only", "incremental_diff"]] = "incremental_diff"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Low: Streamed tool calls are released before moderation

Enabling incremental_diff sends tool-call chunks to the client immediately; the unified streaming hook only submits the assembled tool call for its block decision at end of stream. An attacker can prompt the model to emit a prohibited tool call that reaches an auto-executing client before Airia's block verdict. Buffer tool-call chunks until the complete call passes Airia, or disable this streaming mode when tool calls are present.

@rchawla-airia rchawla-airia Sep 3, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Framework behavior of incremental mode, now called out under Caveats; buffering tool-call chunks belongs in the unified hook as a follow-up

Comment thread litellm/proxy/guardrails/guardrail_hooks/airia/airia.py
Adds `airia` as a built-in guardrail provider. Airia evaluates the prompt under `pre_call` and the
model's response under `post_call`, returning allow, block, or redacted content; the tenant is
resolved from the API key, so the only configuration is `api_base` and `api_key`.

`during_call` is not offered: it runs concurrently with the model call, so a block could land after
the prompt has already reached the provider.

A BLOCKED verdict, or any action this version does not recognise, raises with
`blocked_content=True`; a transport error or non-2xx raises with `blocked_content=False`, so callers
can tell "could not evaluate" from "evaluated and blocked" while both still fail closed. On
GUARDRAIL_INTERVENED a copy of the inputs is returned with every rewritten field substituted; an
intervention carrying no applicable rewrite, or a rewrite of the wrong shape, blocks instead of
letting the original through.

Streamed responses are moderated whole and then emitted redacted: the hook opts into the unified
hook's `incremental_diff` mode (the default `block_only` drops rewrites) with
`streaming_end_of_stream_only`, so a redaction can never span transform rounds and underflow.
@rchawla-airia
rchawla-airia force-pushed the feature/airia-guardrail branch from f2f7a0b to 5e60018 Compare September 3, 2026 13:22
@rchawla-airia

Copy link
Copy Markdown
Author

@greptileai please re-review: 5e60018 moderates streamed responses whole via streaming_end_of_stream_only, so rewrites cannot underflow; tool-call passthrough is documented under Caveats

@rchawla-airia

Copy link
Copy Markdown
Author

Closing this draft for now while we settle which contributor account submits it; the change will be resubmitted unchanged once that is sorted

@airia-open-source

Copy link
Copy Markdown

Resubmitted from the Airia open source account as #40784 with the same change; this closed draft is superseded

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants