Repository navigation
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
9f75c34 to
f06c6c9
Compare
|
@greptileai please review this PR, which keeps a guardrail's masked text when it echoes multipart message rows back unchanged |
|
|
@greptileai please re-review the latest commit, which fixes the tuple comparison finding and adds a regression test for it |
The request model validates list message content lazily, so the dump that builds the POST body consumes it and a later read of the model rows sees empty content. apply_guardrail compared the guardrail's returned rows with those consumed rows, so no multipart row ever matched its echo. A guardrail that echoed every row unchanged and masked through texts had its rows taken as a rewrite, and the unmasked rows reached the LLM Compare against the JSON rows actually posted instead. An unchanged echo now falls back to texts again, and a partial echo restores the caller's row at each echoed index. structured_messages_from_response is renamed to structured_messages_from_json since it now reads request rows as well The dump also sends a row as content [] when one of its parts fails validation (guarded_text, a base64 document, flac audio), so the guardrail never saw that row's text. An echo of it restored the caller's unmasked row while another row's rewrite made the rows path win, which dropped the masked texts. Such a row now carries the caller's own content, so the guardrail sees it in full. Rows the model dumps in full are posted exactly as before, and the rest of the body is unchanged That content goes through a new as_json_value helper in _content_utils. It round-trips through the stdlib json codec because pydantic's serializer silently replaces anything nested past 254 levels with "..." GenericGuardrailAPI also takes an optional async_handler so tests can inject the HTTP client. The regression tests drive the real OpenAI chat and Anthropic Messages guardrail handlers through it
A row whose part list holds a part the request model rejects is now posted with the caller's content as is. The as_json_value round trip it went through is gone: the whole request is already dumped by pydantic before that point, so it never kept extra nesting depth, and no test needed it The two JSON builds in the row helpers carry the same mutable-ok as the POST body, and a new test covers /v1/responses input_text rows, which the fix already handled
A rejected part holding a non-JSON-native value, such as a tuple, was posted as a JSON array but compared against the caller's tuple. An unchanged echo then looked rewritten, so the masked texts were dropped and the original prompt reached the LLM Caller content now goes through the same stdlib codec httpx encodes the body with before it is posted and compared, and a regression test covers the tuple case
…drail_api Main dropped the LIT002 rule, so the mutable-ok comments on the JSON row builds no longer suppress anything and LIT013 now flags them. They are removed. The part-list guard narrows to Sequence[object] instead of list[object], since its callers only need len() and a read, which also keeps it clear of LIT001
597263c to
d1f31a6
Compare
|
@greptileai please re-review: rebased onto main and added a lint-only commit that drops unused suppressions and narrows one type guard |
A guardrail that re-serializes the rows it was sent often drops null fields, such as the thinking_blocks: null an Anthropic assistant row is posted with or the content: null of a chat tool-call row. The echo check compared rows exactly, so such an echo looked like a rewrite of every row: the masked texts were ignored and the unmasked rows written back, and the raw value reached the model. Rows now count as an echo when they match apart from null fields, in both the every-row check and the per-row restore. A field the guardrail sets to null where the caller had a value still counts as a change
|
@greptileai please review 5b61f8b, which treats rows echoed back without their null fields as an echo so masking still applies |
TLDR
Problem this solves:
nameguarded_textreach the guardrail emptyHow it solves it:
textsagainIntentional product change: for a /v1/chat/completions message that holds a part like
guarded_text, a base64documentorflacaudio, the guardrail endpoint now receives the real content where it used to receivecontent: [], so guardrail vendors now get full document and audio bytes and larger request bodies. Every other message is posted exactly as beforeUser Flow
Before: a prompt sent as content blocks on /v1/messages, /v1/chat/completions or /v1/responses reaches the model unmasked, even though the guardrail masked it
pre_callwhose server masks account numbersGUARDRAIL_INTERVENED, echoesstructured_messagesunchanged or without their null fields, and returns[REDACTED]intextsAfter: the same requests reach the model with the number masked
Relevant issues
Split out of #37055, as requested in its review
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
content: nulluv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or moremake checkpasses locally withBASE_REF=upstream/mainbudget-ratchetflags a basedpyright limit that fix(cost_calculator): bill ultrafast prompts above 272k at the ultrafast long-context rates #43764 raised on main, in a file this PR does not touch,proxy-behaviorfails only its Lens coverage upload step, which is missing tools on the runner, andrust-testwas cancelled at its time limit@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review)
Screenshots / Proof of Fix
All runs are against a live proxy on
localhost:4000started withpython litellm/proxy/proxy_cli.py --config proof_config.yaml --detailed_debug --use_v2_migration_resolver, backed by a real Postgres, and every completion is a real, billed call to Bedrockus.anthropic.claude-haiku-4-5-20251001-v1:0.$PROOF_KEYis a virtual key fromPOST /key/generatewith access tobedrock-haiku-4-5The guardrail server is a small FastAPI app on
127.0.0.1:8787that behaves like a masking vendor: it replaces SSN-shaped numbers intextswith[REDACTED], answersGUARDRAIL_INTERVENEDand echoes thestructured_messagesrows it was sent. Like many real servers it re-serializes those rows without their null fields. Nothing else is stubbedThe model refuses to repeat a number like this back, so every case asks it a question only the raw number can answer YES to. YES means the raw number reached Bedrock, NO means only the masked text did. Case 4 is a plain string control, which was masked correctly before this PR too. In case 5 the model answers NO either way, so that case shows what LiteLLM actually sent to Bedrock, read from the
--detailed_debuglog, and thethinking_blocks: nullthe assistant row was posted withBefore (a2bf67a)
Case 1: /v1/messages with a text content block
Case 2: /v1/chat/completions with list content
Case 3: /v1/responses with an input_text part
Case 4 (control): /v1/chat/completions with plain string content
Case 5: /v1/messages after a tool call, with the guardrail echoing rows without their null fields
After (5b61f8b)
Case 1: /v1/messages with a text content block
Case 2: /v1/chat/completions with list content
Case 3: /v1/responses with an input_text part
Case 4 (control): /v1/chat/completions with plain string content
Case 5: /v1/messages after a tool call, with the guardrail echoing rows without their null fields
Type
🐛 Bug Fix
Caveats (if any)
Medium
cache_controlLow
contentwas masked correctly before and still istextswere never affectedtextkey on a non-text part is not sent to the guardrailimage_url,file, wav audio,video_urlandthinkingpartsstructured_messages_from_responseis renamed tostructured_messages_from_jsonGenericGuardrailAPIgains an optionalasync_handlerfor injecting the HTTP clientFinal Attestation