Skip to content

fix(guardrails): scan Responses API input in Azure Text Moderation - #43965

Merged
yucheng-berri merged 2 commits into
mainfrom
litellm_azure_text_moderation_responses_input
Oct 1, 2026
Merged

yucheng-berri merged 2 commits into
mainfrom
litellm_azure_text_moderation_responses_input

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Azure Text Moderation never scans /v1/responses prompts
  • Harmful Responses input reaches the model with a 200
  • An empty messages: [] stub next to input would also skip it

How it solves it:

  • Text Moderation's pre-call hook uses the shared call-type-aware prompt helper from the PR below
  • Responses calls read input, every other call type reads messages

Intentional product change: /v1/responses requests are now scanned by Azure Text Moderation, so harmful text sent there is blocked with a 400 like on chat completions and messages

Stacked on #43786, which fixes the same gap for Prompt Shield only. Split out so that PR stays within its ticket

User Flow

Before: a team with Azure Text Moderation on gets harmful chat and messages prompts blocked, but not Responses

  1. They send POST https://litellm-domain/v1/chat/completions with harmful text and get a 400 from Text Moderation
  2. They send the same text to POST https://litellm-domain/v1/messages and get the same 400
  3. They send the same text as input to POST https://litellm-domain/v1/responses (string, message list, streamed, or next to an empty messages stub)
  4. The request comes back 200 and the model answers, because Azure was never called

After: Responses is moderated the same way

  1. Same chat request, same 400
  2. Same messages request, same 400
  3. Same POST https://litellm-domain/v1/responses requests
  4. Each comes back 400 from Text Moderation and the model is never called

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Two local proxies, base 9bd0602 (the PR below) and head dee9caf (this PR's tip), 2 workers each, own fresh Postgres database each, real Azure AI Content Safety and real OpenAI gpt-5.4-mini, no mocks. A pass-through recorder sits in front of Azure and OpenAI on both sides, forwarding every request to the real API and counting it. Config is the one in #43786 with azure/text_moderations set to default_on: true and the default severity threshold

Payloads, each sent with curl -s localhost:4000<route> -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '<body>', where HARM is I am going to find you and kill you with a knife, then hurt your whole family.

chat             /v1/chat/completions  {"model":"gpt","messages":[{"role":"user","content":"HARM"}]}
messages         /v1/messages          {"model":"gpt","max_tokens":50,"messages":[{"role":"user","content":"HARM"}]}
responses        /v1/responses         {"model":"gpt","input":"HARM"}
responses list   /v1/responses         {"model":"gpt","input":[{"role":"user","content":[{"type":"input_text","text":"HARM"}]}]}
responses stub   /v1/responses         {"model":"gpt","messages":[],"input":"HARM"}
responses stream /v1/responses         {"model":"gpt","input":"HARM","stream":true}
benign           /v1/responses         {"model":"gpt","input":"Say hi in one word"}

The Admin UI case uses the Playground at http://localhost:4000/ui/?page=llm-playground with Endpoint Type /v1/responses and model gpt, then the request's row at http://localhost:4000/ui/?page=logs and its "2 guardrails evaluated" link

Before (9bd0602)

/v1/chat/completions

  1. Send the chat payload: HTTP 400 from text-mod, 0 OpenAI calls

/v1/messages

  1. Send the messages payload: HTTP 400 from text-mod, 0 OpenAI calls

/v1/responses

  1. Send the responses, responses list, responses stub and responses stream payloads: all HTTP 200, the model answers
  2. Recorder: 0 Azure analyze calls and 1 OpenAI call for each

Benign /v1/responses

  1. Send the benign payload: HTTP 200, 0 Azure analyze calls

Admin UI

  1. Send HARM in the Playground: the model answers
  2. Open its log: Success with an LLM call, prompt-shield PASSED with 1 text record and $0.00038000, text-mod PASSED at 0ms because Azure was never asked (the recorder logged a Prompt Shield call and no analyze call)
Playground Request log
Base harmful Base log

After (dee9caf)

/v1/chat/completions

  1. Send the chat payload: HTTP 400 from text-mod, 0 OpenAI calls, unchanged

/v1/messages

  1. Send the messages payload: HTTP 400 from text-mod, 0 OpenAI calls, unchanged

/v1/responses

  1. Send the responses, responses list, responses stub and responses stream payloads: all HTTP 400 with x-litellm-applied-guardrails: text-mod
  2. Recorder: 1 Azure analyze call and 0 OpenAI calls for each

Benign /v1/responses

  1. Send the benign payload: HTTP 200, 1 Azure analyze call, the model answers

Admin UI

  1. Send HARM in the Playground: "400 Azure Content Safety Guardrail: Violence crossed severity 2, Got severity: 4"
  2. Open its log: Failure, prompt-shield PASSED with 1 text record and $0.00038000, text-mod FAILED after a real 477ms Azure call, and the model was never called
  3. Send "Say hi in one word": the model answers "Hi", and its log shows prompt-shield and text-mod both PASSED with a real Azure call each
Playground Blocked request log Benign request log
Head harmful Head log Head benign

Recorder totals over the seven payloads: base made 7 Prompt Shield calls, 2 Azure analyze calls and 5 OpenAI calls. Head made 7 Prompt Shield calls, 7 Azure analyze calls and 1 OpenAI call

Deterministic audit (dee9caf)

Same two tests/integration files as the PR below plus 6 Text Moderation cells, 33 cells total

uv run --no-sync python tests/integration/run.py extensions \
  tests/integration/observability/test_azure_content_safety_audit.py \
  tests/integration/observability/test_azure_content_safety_endpoints.py \
  --seed 4106601 --order-seed 7
leg sha result
base 9bd0602 5 failed, 28 passed, the failures are the Responses string, list, stub, over-threshold and streamed over-threshold Text Moderation cells
head run 1 dee9caf 33 passed, 0 skipped
head run 2 dee9caf 33 passed, 0 skipped, same node list as run 1

Type

🐛 Bug Fix
✅ Test

Caveats (if any)

Low

  • Only the trailing user message block of input is scanned
    • Same rule chat completions already uses for this guardrail
  • The scanned prompt is now logged at debug instead of info, matching Prompt Shield, so it no longer reaches info logs on any endpoint
  • Text Moderation records a status but no usage or cost
    • Existing behavior on every endpoint, unchanged here

REVIEWER MUST KNOW BEFORE APPROVING

  • Responses text over the Text Moderation threshold: 200 and forwarded to the model on base, 400 with no model call on head. Approval pending from @yucheng-berri
  • Azure Content Safety now receives the user text of Responses requests for Text Moderation, adding Azure latency per request. Approval pending from @yucheng-berri
  • Text Moderation's scanned prompt log line moves from info to debug on every endpoint, so operators running info logs stop seeing user prompts there. Approval pending from @yucheng-berri

Link to Devin session: https://app.devin.ai/sessions/5c9051d5c1294188a325003170ceb7df
Open in Devin Desktop: https://app.devin.ai/desktop/session/5c9051d5c1294188a325003170ceb7df?variant=devin


Note

Medium Risk
Changes pre-call content moderation for /v1/responses, adding Azure analyze calls and blocking behavior where requests previously reached the model; chat behavior is unchanged aside from log level.

Overview
Azure Text Moderation now moderates /v1/responses the same way as chat: the pre-call hook uses shared get_user_prompt_from_request so Responses calls read input (string, structured list, streaming) and other routes still read messages, including when messages: [] would have skipped scanning before.

The redundant get_user_prompt wrapper on AzureGuardrailBase is removed. Scanned prompt text is logged at debug instead of info on all endpoints.

Integration and unit tests cover opt-in moderation on Responses (including block above threshold), chat-only vs shadow input, and the Azure wire mock for text:analyze.

Reviewed by Cursor Bugbot for commit 436904a. Bugbot is set up for automated code reviews on this repo. Configure here.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@devin-ai-integration
devin-ai-integration Bot added this pull request to stack #43966 October 1, 2026 01:58
@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Refactors Azure Text Moderation guardrail to scan API responses.

The PR appears safe to merge; no outstanding finding or new actionable issue was identified.

Summary

Azure Text Moderation now extracts Responses API input for pre-call scanning while continuing to use messages for other call types.

  • Moves the scanned-prompt log line from info to debug.
  • Adds unit and integration coverage for Responses input, blocking, and streaming.

Reviews (3) · Last reviewed commit: "fix(guardrails): log Azure Text Moderati..."

Comment thread litellm/proxy/guardrails/guardrail_hooks/azure/text_moderation.py
Comment thread tests/integration/observability/test_azure_content_safety_audit.py
@codspeed

codspeed Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_azure_text_moderation_responses_input (436904a) with main (6d7d183)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (19842da) during the generation of this report, so 6d7d183 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@yucheng-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Base automatically changed from litellm_azure_guardrail_responses_input to main October 1, 2026 05:48
yucheng-berri and others added 2 commits September 30, 2026 22:48
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… streamed Responses blocking

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@yucheng-berri
yucheng-berri force-pushed the litellm_azure_text_moderation_responses_input branch from dee9caf to 436904a Compare October 1, 2026 05:48
@yucheng-berri

Copy link
Copy Markdown
Contributor

@greptileai review latest head

@yucheng-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 436904a. Configure here.

@yucheng-berri
yucheng-berri merged commit 0980f75 into main Oct 1, 2026
103 of 105 checks passed
@yucheng-berri
yucheng-berri deleted the litellm_azure_text_moderation_responses_input branch October 1, 2026 05:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant