Skip to content

fix(guardrails): keep masked text when a guardrail echoes multipart rows - #43679

Open
caduri wants to merge 5 commits into
BerriAI:mainfrom
caduri:bugfix/generic-guardrail-multipart-row-write-back
Open

caduri wants to merge 5 commits into
BerriAI:mainfrom
caduri:bugfix/generic-guardrail-multipart-row-write-back

Conversation

@caduri

@caduri caduri commented Sep 29, 2026 •

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Generic guardrail masking is dropped for messages with list content
  • The unmasked prompt, account number included, reaches the LLM
  • Messages the guardrail echoes back lose keys like name
  • Messages holding a part like guarded_text reach the guardrail empty
  • An echo that drops null fields counts as a rewrite, so the unmasked rows win

How it solves it:

  • Compare returned messages with the ones actually posted
  • An unchanged echo falls back to the masked texts again
  • A partly echoed request keeps the caller's own messages
  • Messages the guardrail used to get empty carry their real content
  • Rows that match apart from null fields count as an echo too

Intentional product change: for a /v1/chat/completions message that holds a part like guarded_text, a base64 document or flac audio, the guardrail endpoint now receives the real content where it used to receive content: [], so guardrail vendors now get full document and audio bytes and larger request bodies. Every other message is posted exactly as before

User Flow

Before: a prompt sent as content blocks on /v1/messages, /v1/chat/completions or /v1/responses reaches the model unmasked, even though the guardrail masked it

  1. The proxy admin adds a Generic Guardrail API guardrail on pre_call whose server masks account numbers
  2. The app sends POST https://litellm-domain/v1/messages with "my account number is 123-45-6789" in a text block
  3. The guardrail server answers GUARDRAIL_INTERVENED, echoes structured_messages unchanged or without their null fields, and returns [REDACTED] in texts
  4. The model's reply shows it saw the real number, as if no guardrail ran

After: the same requests reach the model with the number masked

  1. The proxy admin adds the same Generic Guardrail API guardrail
  2. The app sends the same POST https://litellm-domain/v1/messages
  3. The guardrail server gives the same answer
  4. The model's reply shows it only saw "my account number is [REDACTED]"

Relevant issues

Split out of #37055, as requested in its review

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
    • The 11 new regression tests fail at the merge base 2c9b0e0 with assertion errors, and the 2 controls for plain strings and rows the request model accepts pass there
    • The 2 null-echo regression tests fail at d1f31a6 with assertion errors: an Anthropic conversation after a tool call, and a chat tool-call row echoed without its content: null
    • Each of these mutations fails a test: comparing either echo check exactly again, keeping nulls, treating every pair or every key as equal, normalizing only one side, making the row restore a no-op, never or always restoring the caller's content, sending the caller's whole row, comparing echoes with the model rows again, posting the plain model dump, skipping the content length check, and skipping the posted-JSON normalization of the caller's content
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)
    • Greptile scored 5/5 on d1f31a6 after the rebase onto main, and Veria reported no security issues. Greptile scored 5/5 again on 5b61f8b
    • Bugbot has not run on this PR

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review)

Screenshots / Proof of Fix

All runs are against a live proxy on localhost:4000 started with python litellm/proxy/proxy_cli.py --config proof_config.yaml --detailed_debug --use_v2_migration_resolver, backed by a real Postgres, and every completion is a real, billed call to Bedrock us.anthropic.claude-haiku-4-5-20251001-v1:0. $PROOF_KEY is a virtual key from POST /key/generate with access to bedrock-haiku-4-5

model_list:
  - model_name: bedrock-haiku-4-5
    litellm_params:
      model: bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0
      aws_region_name: us-east-1

guardrails:
  - guardrail_name: ssn-masker
    litellm_params:
      guardrail: generic_guardrail_api
      mode: pre_call
      api_base: http://127.0.0.1:8787
      default_on: true

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL

The guardrail server is a small FastAPI app on 127.0.0.1:8787 that behaves like a masking vendor: it replaces SSN-shaped numbers in texts with [REDACTED], answers GUARDRAIL_INTERVENED and echoes the structured_messages rows it was sent. Like many real servers it re-serializes those rows without their null fields. Nothing else is stubbed

import json
import re
from pathlib import Path

from fastapi import FastAPI, Request

LOG = Path(__file__).with_name("masker.jsonl")
SSN = re.compile(r"\b\d{3}-\d{2}-\d{4}\b")
app = FastAPI()


def _without_nulls(rows: object) -> object:
    return json.loads(json.dumps(rows), object_hook=lambda row: {key: value for key, value in row.items() if value is not None})


@app.post("/beta/litellm_basic_guardrail_api")
async def guardrail(request: Request) -> dict[str, object]:
    payload = await request.json()
    with LOG.open("a") as log:
        log.write(json.dumps(payload) + "\n")
    return {
        "action": "GUARDRAIL_INTERVENED",
        "texts": [SSN.sub("[REDACTED]", text) for text in payload.get("texts") or []],
        "structured_messages": _without_nulls(payload.get("structured_messages")),
    }


@app.post("/_reset")
async def reset() -> dict[str, int]:
    LOG.write_text("")
    return {"records": 0}

The model refuses to repeat a number like this back, so every case asks it a question only the raw number can answer YES to. YES means the raw number reached Bedrock, NO means only the masked text did. Case 4 is a plain string control, which was masked correctly before this PR too. In case 5 the model answers NO either way, so that case shows what LiteLLM actually sent to Bedrock, read from the --detailed_debug log, and the thinking_blocks: null the assistant row was posted with

Before (a2bf67a)

Case 1: /v1/messages with a text content block

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/messages -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "messages": [{"role": "user", "content": [{"type": "text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: .content[0].text, usage}'
    
  2. Observe
    {"text":"YES","usage":{"input_tokens":40,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard"}}
    

Case 2: /v1/chat/completions with list content

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "messages": [{"role": "user", "content": [{"type": "text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}}'
    
  2. Observe
    {"text":"YES","usage":{"prompt_tokens":40,"completion_tokens":4}}
    

Case 3: /v1/responses with an input_text part

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/responses -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_output_tokens": 40, "input": [{"role": "user", "content": [{"type": "input_text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: [.output[] | select(.type == "message") | .content[0].text][0], usage: {input_tokens: .usage.input_tokens, output_tokens: .usage.output_tokens}}'
    
  2. Observe
    {"text":"YES","usage":{"input_tokens":40,"output_tokens":4}}
    

Case 4 (control): /v1/chat/completions with plain string content

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "messages": [{"role": "user", "content": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}' | jq -c '{text: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}}'
    
  2. Observe
    {"text":"NO","usage":{"prompt_tokens":38,"completion_tokens":4}}
    
  3. Run
    $ wc -l < masker.jsonl
    
  4. Observe
    1
    

Case 5: /v1/messages after a tool call, with the guardrail echoing rows without their null fields

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/messages -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "tools": [{"name": "lookup", "description": "Look up the account", "input_schema": {"type": "object", "properties": {}}}], "messages": [{"role": "user", "content": "Look up my account."}, {"role": "assistant", "content": [{"type": "text", "text": "Calling the lookup tool."}, {"type": "tool_use", "id": "toolu_1", "name": "lookup", "input": {}}]}, {"role": "user", "content": [{"type": "tool_result", "tool_use_id": "toolu_1", "content": "done"}, {"type": "text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: .content[0].text, usage}'
    
  2. Observe
    {"text":"NO","usage":{"input_tokens":636,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard"}}
    
  3. Run
    $ jq -c '[.structured_messages[] | select(.role == "assistant") | .thinking_blocks]' masker.jsonl
    
  4. Observe
    [null]
    
  5. Run
    $ grep -oE "account number is [^'\"]*" litellm.log | tail -1
    
  6. Observe
    account number is 123-45-6789
    

After (5b61f8b)

Case 1: /v1/messages with a text content block

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/messages -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "messages": [{"role": "user", "content": [{"type": "text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: .content[0].text, usage}'
    
  2. Observe
    {"text":"NO","usage":{"input_tokens":38,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard"}}
    

Case 2: /v1/chat/completions with list content

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "messages": [{"role": "user", "content": [{"type": "text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}}'
    
  2. Observe
    {"text":"NO","usage":{"prompt_tokens":38,"completion_tokens":4}}
    

Case 3: /v1/responses with an input_text part

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/responses -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_output_tokens": 40, "input": [{"role": "user", "content": [{"type": "input_text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: [.output[] | select(.type == "message") | .content[0].text][0], usage: {input_tokens: .usage.input_tokens, output_tokens: .usage.output_tokens}}'
    
  2. Observe
    {"text":"NO","usage":{"input_tokens":38,"output_tokens":4}}
    

Case 4 (control): /v1/chat/completions with plain string content

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "messages": [{"role": "user", "content": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}' | jq -c '{text: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}}'
    
  2. Observe
    {"text":"NO","usage":{"prompt_tokens":38,"completion_tokens":4}}
    
  3. Run
    $ wc -l < masker.jsonl
    
  4. Observe
    1
    

Case 5: /v1/messages after a tool call, with the guardrail echoing rows without their null fields

  1. Run
    $ curl -s -X POST http://localhost:4000/v1/messages -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 40, "tools": [{"name": "lookup", "description": "Look up the account", "input_schema": {"type": "object", "properties": {}}}], "messages": [{"role": "user", "content": "Look up my account."}, {"role": "assistant", "content": [{"type": "text", "text": "Calling the lookup tool."}, {"type": "tool_use", "id": "toolu_1", "name": "lookup", "input": {}}]}, {"role": "user", "content": [{"type": "tool_result", "tool_use_id": "toolu_1", "content": "done"}, {"type": "text", "text": "Does the sentence after the colon contain an actual number rather than a placeholder? Answer only YES or NO: my account number is 123-45-6789"}]}]}' | jq -c '{text: .content[0].text, usage}'
    
  2. Observe
    {"text":"NO","usage":{"input_tokens":634,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard"}}
    
  3. Run
    $ jq -c '[.structured_messages[] | select(.role == "assistant") | .thinking_blocks]' masker.jsonl
    
  4. Observe
    [null]
    
  5. Run
    $ grep -oE "account number is [^'\"]*" litellm.log | tail -1
    
  6. Observe
    account number is [REDACTED]
    

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Guardrail vendors now get full document and audio bytes for those messages
    • Those parts keep every caller key, including cache_control
    • A vendor with a body size limit could start rejecting such requests
    • With the default fail-closed setting, that rejection blocks the request

Low

  • A row that differs only in null fields now counts as an echo
    • A guardrail that sets a field to null where the caller had a value is still a rewrite
    • Nulling a key the request model removed before posting gets the caller's value back, since the guardrail never saw it
  • Plain string content was masked correctly before and still is
    • Guardrails that return only texts were never affected
  • The request model still validates list content lazily for its other readers
    • This PR works around it for the generic guardrail only
  • Assumed: only /v1/chat/completions sends parts the request model rejects
    • /v1/messages and /v1/responses convert or drop such blocks before the guardrail runs
    • If a route did not, it now gets the same fix
  • Pre-existing: a stray text key on a non-text part is not sent to the guardrail
    • Seen on image_url, file, wav audio, video_url and thinking parts
    • It can reach the provider unmasked when another message is rewritten
    • Providers do not read that key as prompt text
  • Assumed: response and streaming paths never send rows, so they are unaffected
    • If one did, it had the same bug and now gets the same fix
  • structured_messages_from_response is renamed to structured_messages_from_json
    • Nothing else in the repo imports it, and it only shipped in dev tags
  • GenericGuardrailAPI gains an optional async_handler for injecting the HTTP client
    • It defaults to the shared client, and config cannot set it

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

CLAassistant commented Sep 29, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing caduri:bugfix/generic-guardrail-multipart-row-write-back (5b61f8b) with main (7a7d27c)

Open in CodSpeed

@caduri

caduri commented Oct 1, 2026

Copy link
Copy Markdown
Author

@greptileai please review this PR, which keeps a guardrail's masked text when it echoes multipart message rows back unchanged

@caduri
caduri marked this pull request as ready for review October 1, 2026 09:01
@caduri
caduri requested a review from a team October 1, 2026 09:01
@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Guardrail response handling for multipart message content.

The PR appears safe to merge based on the reviewed changes

Summary

The PR preserves masked text when a generic guardrail echoes multipart messages, including echoes that omit null fields

  • Compares returned rows with the JSON actually posted and restores unchanged caller rows during partial rewrites
  • Adds regression coverage for chat, Anthropic Messages, and Responses inputs

Reviews (4) · Last reviewed commit: "fix(guardrails): treat a row echoed with..."

@caduri

caduri commented Oct 1, 2026

Copy link
Copy Markdown
Author

@greptileai please re-review the latest commit, which fixes the tuple comparison finding and adds a regression test for it

caduri added 4 commits October 4, 2026 15:01
The request model validates list message content lazily, so the dump that
builds the POST body consumes it and a later read of the model rows sees
empty content. apply_guardrail compared the guardrail's returned rows with
those consumed rows, so no multipart row ever matched its echo. A guardrail
that echoed every row unchanged and masked through texts had its rows taken
as a rewrite, and the unmasked rows reached the LLM

Compare against the JSON rows actually posted instead. An unchanged echo now
falls back to texts again, and a partial echo restores the caller's row at
each echoed index. structured_messages_from_response is renamed to
structured_messages_from_json since it now reads request rows as well

The dump also sends a row as content [] when one of its parts fails
validation (guarded_text, a base64 document, flac audio), so the guardrail
never saw that row's text. An echo of it restored the caller's unmasked row
while another row's rewrite made the rows path win, which dropped the masked
texts. Such a row now carries the caller's own content, so the guardrail sees
it in full. Rows the model dumps in full are posted exactly as before, and
the rest of the body is unchanged

That content goes through a new as_json_value helper in _content_utils. It
round-trips through the stdlib json codec because pydantic's serializer
silently replaces anything nested past 254 levels with "..."

GenericGuardrailAPI also takes an optional async_handler so tests can inject
the HTTP client. The regression tests drive the real OpenAI chat and Anthropic
Messages guardrail handlers through it
A row whose part list holds a part the request model rejects is now posted
with the caller's content as is. The as_json_value round trip it went
through is gone: the whole request is already dumped by pydantic before
that point, so it never kept extra nesting depth, and no test needed it

The two JSON builds in the row helpers carry the same mutable-ok as the
POST body, and a new test covers /v1/responses input_text rows, which the
fix already handled
A rejected part holding a non-JSON-native value, such as a tuple, was
posted as a JSON array but compared against the caller's tuple. An
unchanged echo then looked rewritten, so the masked texts were dropped and
the original prompt reached the LLM

Caller content now goes through the same stdlib codec httpx encodes the
body with before it is posted and compared, and a regression test covers
the tuple case
…drail_api

Main dropped the LIT002 rule, so the mutable-ok comments on the JSON row builds no longer suppress anything and LIT013 now flags them. They are removed. The part-list guard narrows to Sequence[object] instead of list[object], since its callers only need len() and a read, which also keeps it clear of LIT001
@caduri
caduri force-pushed the bugfix/generic-guardrail-multipart-row-write-back branch from 597263c to d1f31a6 Compare October 4, 2026 12:11
@caduri

caduri commented Oct 4, 2026

Copy link
Copy Markdown
Author

@greptileai please re-review: rebased onto main and added a lint-only commit that drops unused suppressions and narrows one type guard

A guardrail that re-serializes the rows it was sent often drops null fields,
such as the thinking_blocks: null an Anthropic assistant row is posted with or
the content: null of a chat tool-call row. The echo check compared rows
exactly, so such an echo looked like a rewrite of every row: the masked texts
were ignored and the unmasked rows written back, and the raw value reached the
model. Rows now count as an echo when they match apart from null fields, in
both the every-row check and the per-row restore. A field the guardrail sets
to null where the caller had a value still counts as a change
@caduri

caduri commented Oct 5, 2026

Copy link
Copy Markdown
Author

@greptileai please review 5b61f8b, which treats rows echoed back without their null fields as an echo so masking still applies

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants