Repository navigation
fix(guardrails): scan Responses API input in Azure Text Moderation - #43965
Conversation
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… streamed Responses blocking Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
dee9caf to
436904a
Compare
|
@greptileai review latest head |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 436904a. Configure here.
TLDR
Problem this solves:
/v1/responsespromptsinputreaches the model with a 200messages: []stub next toinputwould also skip itHow it solves it:
input, every other call type readsmessagesIntentional product change:
/v1/responsesrequests are now scanned by Azure Text Moderation, so harmful text sent there is blocked with a 400 like on chat completions and messagesStacked on #43786, which fixes the same gap for Prompt Shield only. Split out so that PR stays within its ticket
User Flow
Before: a team with Azure Text Moderation on gets harmful chat and messages prompts blocked, but not Responses
inputto POST https://litellm-domain/v1/responses (string, message list, streamed, or next to an emptymessagesstub)After: Responses is moderated the same way
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Two local proxies, base 9bd0602 (the PR below) and head dee9caf (this PR's tip), 2 workers each, own fresh Postgres database each, real Azure AI Content Safety and real OpenAI
gpt-5.4-mini, no mocks. A pass-through recorder sits in front of Azure and OpenAI on both sides, forwarding every request to the real API and counting it. Config is the one in #43786 withazure/text_moderationsset todefault_on: trueand the default severity thresholdPayloads, each sent with
curl -s localhost:4000<route> -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'content-type: application/json' -d '<body>', where HARM isI am going to find you and kill you with a knife, then hurt your whole family.The Admin UI case uses the Playground at http://localhost:4000/ui/?page=llm-playground with Endpoint Type
/v1/responsesand modelgpt, then the request's row at http://localhost:4000/ui/?page=logs and its "2 guardrails evaluated" linkBefore (9bd0602)
/v1/chat/completions
HTTP 400fromtext-mod, 0 OpenAI calls/v1/messages
HTTP 400fromtext-mod, 0 OpenAI calls/v1/responses
HTTP 200, the model answersBenign /v1/responses
HTTP 200, 0 Azure analyze callsAdmin UI
After (dee9caf)
/v1/chat/completions
HTTP 400fromtext-mod, 0 OpenAI calls, unchanged/v1/messages
HTTP 400fromtext-mod, 0 OpenAI calls, unchanged/v1/responses
HTTP 400withx-litellm-applied-guardrails: text-modBenign /v1/responses
HTTP 200, 1 Azure analyze call, the model answersAdmin UI
Recorder totals over the seven payloads: base made 7 Prompt Shield calls, 2 Azure analyze calls and 5 OpenAI calls. Head made 7 Prompt Shield calls, 7 Azure analyze calls and 1 OpenAI call
Deterministic audit (dee9caf)
Same two
tests/integrationfiles as the PR below plus 6 Text Moderation cells, 33 cells totalType
🐛 Bug Fix
✅ Test
Caveats (if any)
Low
inputis scannedREVIEWER MUST KNOW BEFORE APPROVING
Link to Devin session: https://app.devin.ai/sessions/5c9051d5c1294188a325003170ceb7df
Open in Devin Desktop: https://app.devin.ai/desktop/session/5c9051d5c1294188a325003170ceb7df?variant=devin
Note
Medium Risk
Changes pre-call content moderation for
/v1/responses, adding Azure analyze calls and blocking behavior where requests previously reached the model; chat behavior is unchanged aside from log level.Overview
Azure Text Moderation now moderates
/v1/responsesthe same way as chat: the pre-call hook uses sharedget_user_prompt_from_requestso Responses calls readinput(string, structured list, streaming) and other routes still readmessages, including whenmessages: []would have skipped scanning before.The redundant
get_user_promptwrapper onAzureGuardrailBaseis removed. Scanned prompt text is logged at debug instead of info on all endpoints.Integration and unit tests cover opt-in moderation on Responses (including block above threshold), chat-only vs shadow
input, and the Azure wire mock fortext:analyze.Reviewed by Cursor Bugbot for commit 436904a. Bugbot is set up for automated code reviews on this repo. Configure here.