Skip to content

fix(guardrails): scan each choice's tool-call arguments apart on n>1 streams and log why a rewrite was discarded - #40986

Merged
mateo-berri merged 7 commits into
mainfrom
litellm_lit_7346_multi_choice_stream_guardrails
Sep 20, 2026
Merged

mateo-berri merged 7 commits into
mainfrom
litellm_lit_7346_multi_choice_stream_guardrails

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • n>1 streamed tool calls are rebuilt with every choice's argument fragments merged into one string
  • tool_permission and the other post_call hooks reading that rebuild reject valid arguments as unparseable
  • logging integrations receive the same garbled arguments
  • a rewrite the stream cannot take back is dropped with no guardrail name and no reason

How it solves it:

  • the rebuilt response keys tool calls by (choice index, tool index)
  • every discarded rewrite logs the guardrail name and the exact reason
  • reasons name the choice count, the tool-call count change, or the legacy-hook limit

User Flow

Before: a developer streaming an n=2 tool call through a proxy whose post_call tool_permission guardrail allows lowercase fruit names gets HTTP 400 "arguments could not be parsed", their logging integration records one garbled tool call, and the proxy log never says why a mask was dropped

  1. The proxy admin attaches a post_call tool_permission guardrail named tool-arg-filter to gpt-4.1-mini that allows lookup_fruit when its fruit argument matches ^[a-z-]+$, plus a post_call content filter that masks the word persimmon
  2. The developer sends POST https://litellm-domain/v1/chat/completions with "stream": true, "n": 2, the user message "Look up the fruit blackberry-elderberry-gooseberry-huckleberry-lingonberry.", one lookup_fruit tool, tool_choice forcing that tool, and "guardrails": ["tool-arg-filter"]
  3. The response is HTTP 400 {"error":{"message":"Guardrail raised an exception, Guardrail: tool-arg-filter, Message: Tool 'lookup_fruit' arguments could not be parsed required by rule 'lookup_fruit_lowercase'","type":"None","param":"None","code":"400"}}, even though both choices asked for a lowercase fruit the rule allows
  4. They send the same request for "Look up the fruit persimmon." without the guardrails field and get HTTP 200, the two choices' argument fragments interleaved chunk by chunk, each assembling to {"fruit":"persimmon"}, unmasked
  5. The logging integration attached to the proxy records that request as one lookup_fruit call whose arguments read {"fruit":"{"persfruitim":"monpers"}immon"}
  6. The proxy admin's log says Masked keyword 'persimmon' in content twice, then only Pipeline: guardrail 'output-word-filter' rewrote the streamed response in a way this endpoint's streaming pipeline cannot deliver yet; the rewrite was discarded and the original stream released, with no reason

After: the same tool_permission request comes back HTTP 200 with both tool calls, the logging integration records each choice's arguments intact, and the proxy log names the guardrail and the reason the mask was dropped

  1. The proxy admin attaches a post_call tool_permission guardrail named tool-arg-filter to gpt-4.1-mini that allows lookup_fruit when its fruit argument matches ^[a-z-]+$, plus a post_call content filter that masks the word persimmon
  2. The developer sends POST https://litellm-domain/v1/chat/completions with "stream": true, "n": 2, the user message "Look up the fruit blackberry-elderberry-gooseberry-huckleberry-lingonberry.", one lookup_fruit tool, tool_choice forcing that tool, and "guardrails": ["tool-arg-filter"]
  3. The response is HTTP 200: one chunk on choice 0 carrying both samples' lookup_fruit calls, {"fruit":"blackberry-elderberry-gooseberry-huckleberry-lingonberry"} and the second sample's shorter {"fruit":"blackberry"}, then data: [DONE]
  4. They send the same request for "Look up the fruit persimmon." without the guardrails field and get HTTP 200, the two choices' argument fragments interleaved chunk by chunk, each assembling to {"fruit":"persimmon"}, still unmasked
  5. The logging integration attached to the proxy records that request as two lookup_fruit calls, each with arguments {"fruit":"persimmon"}
  6. The proxy admin's log says Masked keyword 'persimmon' in content twice, then Pipeline: guardrail 'output-word-filter' rewrote the streamed response but the rewrite could not be written back to the stream: the stream carries 2 choices and tool-call rewrites are only written back on single-choice streams. The whole rewrite, text rewrites included, was discarded and the original stream released

Design decisions

  • The rebuilt response keeps a single choice that carries every stream choice's tool calls, now one entry per (choice index, tool index) instead of one per tool index. Rebuilding one choice per stream choice changes the shape every consumer of the rebuilt response sees and is tracked as LIT-8224
  • The content-filter path that scans and masks streamed tool calls per choice landed on main in fix(policy_engine): deliver guardrail text rewrites on multi-choice, unfinished, and envelope-less streams #41933 (LIT-7178) and is in this PR's merge base; what this PR fixes is the rebuilt response that everything else reads: the legacy streaming hooks (tool_permission, bedrock, panw_prisma_airs, cisco_ai_defense, noma, model_armor, purview_dlp), the logging integrations, and the discard path
  • Writing a rewrite back into a multi-choice stream stays refused. The rebuilt response has one choice, so a mask on it cannot be attributed to choice 0 or choice 1 on the wire; this PR makes that refusal loud instead of lifting it. The write-back helper that pairs rewrites with stream fragments still groups by tool index only, on purpose, because it only runs once the stream is known to carry a single choice
  • Text content across choices is still joined into one string by the rebuilt response, exactly as before; per-choice text assembly for n>1 is part of LIT-8224 and is not touched here
  • The rebuild's tool-call selection and its text-only fast path now look at every choice in a chunk instead of choice 0 only. OpenAI sends one choice per chunk, but a provider that packs one choice per candidate into the same chunk had a tool call carried only by a later candidate dropped before the rebuild ever saw it (raised in review). A packed chunk now goes through the same assembly as a one-choice-per-chunk stream instead of the text-only shortcut
  • Refusal reasons are full sentences with counts so the warning alone explains what happened, and the warning says the text rewrites were discarded with the tool-call rewrite, which is the widened case from the ticket
  • OpenAI sends the same tool-call id on both choices under a forced tool_choice, so the two rebuilt tool calls in the proof share an id; that is provider behavior, not the rebuild

Relevant issues

Affected release

Linear ticket

Resolves LIT-7346

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs run the same config through a real proxy against the real OpenAI API with gpt-4.1-mini, the chat model that still honors n (the gpt-5.x chat models are bridged to the Responses API and ignore it). Each leg is its own checkout with PYTHONPATH pointing at it, one proxy per leg on its own random port, 2 uvicorn workers, no database, the master key and provider keys from the environment. Before is 1fcef68, the merge base 3df1593 plus the two test-only commits of #42048 (git diff --stat 3df159308d5 1fcef68ab7 -- . ':!tests' is empty). After is 2a35dc5; the three commits after it, 810acda (drops a comment, restores the OpenAPI snapshot), 2e83871, and 19d77e2 (type annotations on a test-only hook), cannot change behavior

Config (lit7346_config.yaml):

model_list:
  - model_name: gpt-4.1-mini
    litellm_params:
      model: openai/gpt-4.1-mini
      api_key: os.environ/OPENAI_API_KEY

guardrails:
  - guardrail_name: output-word-filter
    litellm_params:
      guardrail: litellm_content_filter
      mode: post_call
      blocked_words:
        - keyword: kumquat
          action: BLOCK
        - keyword: persimmon
          action: MASK
        - keyword: mangosteen
          action: BLOCK
  - guardrail_name: tool-arg-filter
    litellm_params:
      guardrail: tool_permission
      mode: post_call
      default_on: false
      rules:
        - id: lookup_fruit_lowercase
          tool_name: "^lookup_fruit$"
          decision: allow
          allowed_param_patterns:
            fruit: "^[a-z-]+$"
      default_action: deny
      on_disallowed_action: block

policies:
  response-governance:
    guardrails:
      add: [output-word-filter]
    pipeline:
      mode: post_call
      steps:
        - guardrail: output-word-filter
          on_pass: allow
          on_fail: block

policy_attachments:
  - policy: response-governance
    models: [gpt-4.1-mini]

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY

Proxy boot, each leg from its own checkout on its own port:

export PYTHONPATH=$PWD
python litellm/proxy/proxy_cli.py --config lit7346_config.yaml --port $PORT --num_workers 2 --detailed_debug > proxy.log 2>&1

Requests. toolperm.json is the tool_permission case; the content-filter cases only swap the user message and drop the guardrails field: persimmon (masked, streamed as pers, im, mon), mangosteen (blocked, streamed as m, ang, ost, een), kumquat (blocked, streamed whole). persimmon_n1.json is the persimmon request with "n": 1, nonstream_n2.json the persimmon request with "stream": false, and the two text requests carry no tools

{"model":"gpt-4.1-mini","stream":true,"n":2,"messages":[{"role":"user","content":"Look up the fruit blackberry-elderberry-gooseberry-huckleberry-lingonberry."}],"tools":[{"type":"function","function":{"name":"lookup_fruit","parameters":{"type":"object","properties":{"fruit":{"type":"string"}},"required":["fruit"]}}}],"tool_choice":{"type":"function","function":{"name":"lookup_fruit"}},"guardrails":["tool-arg-filter"]}
{"model":"gpt-4.1-mini","stream":true,"n":2,"messages":[{"role":"user","content":"Repeat exactly this sentence and nothing else: The persimmon is ripe."}]}

The curl and the one-liner that pulls each choice's argument fragments out of a saved stream run the same way on every leg:

curl -sS -N http://localhost:$PORT/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H "Content-Type: application/json" -d @toolperm.json -o toolperm.1.out -w 'HTTP %{http_code}\n'
grep -oE '"index":[01],"delta":\{[^}]*"arguments":"[^"]*"' persimmon.1.out | sed -E 's/"index":([01]),"delta".*"arguments":/c\1=/' | tr '\n' ' '

Before (1fcef68)

tool_permission on an n=2 tool call, three runs

  1. for run in 1 2 3; do curl ... -d @toolperm.json -o toolperm.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 400, run 2: HTTP 400, run 3: HTTP 400
  3. cat toolperm.1.out prints {"error":{"message":"Guardrail raised an exception, Guardrail: tool-arg-filter, Message: Tool 'lookup_fruit' arguments could not be parsed required by rule 'lookup_fruit_lowercase'","type":"None","param":"None","code":"400"}}
  4. The proxy log for each run reads Tool Permission Guardrail: Found 1 tool calls and then Tool 'lookup_fruit' arguments could not be parsed required by rule 'lookup_fruit_lowercase'

Masked word on an n=2 tool call (persimmon), two runs

  1. for run in 1 2; do curl ... -d @persimmon.json -o persimmon.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 200, run 2: HTTP 200
  3. The fragment one-liner on persimmon.1.out prints c0="" c0="{\" c0="fruit" c0="\" c1="" c1="{\" c0="pers" c1="fruit" c0="im" c1="\" c0="mon" c1="pers" c0="\" c1="im" c1="mon" c1="\": both choices assemble to {"fruit":"persimmon"}, unmasked
  4. The proxy log reads Masked keyword 'persimmon' in content twice, then Pipeline: guardrail 'output-word-filter' rewrote the streamed response in a way this endpoint's streaming pipeline cannot deliver yet; the rewrite was discarded and the original stream released

Blocked word streamed in pieces (mangosteen), three runs

  1. for run in 1 2 3; do curl ... -d @mangosteen.json -o mangosteen.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 400, run 2: HTTP 400, run 3: HTTP 400
  3. cat mangosteen.1.out prints data: {"error": {"message": "Content blocked: keyword 'mangosteen' detected", "type": "invalid_request_error", "param": null, "code": "400", "provider_specific_fields": {"error": "Content blocked: keyword 'mangosteen' detected", "keyword": "mangosteen", "description": null, "guardrail_name": "output-word-filter", "guardrail_mode": "post_call"}}} then data: [DONE]

Blocked word streamed whole (kumquat), three runs

  1. for run in 1 2 3; do curl ... -d @kumquat.json -o kumquat.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 400, run 2: HTTP 400, run 3: HTTP 400
  3. cat kumquat.1.out prints data: {"error": {"message": "Content blocked: keyword 'kumquat' detected", ... "guardrail_name": "output-word-filter", "guardrail_mode": "post_call"}} then data: [DONE]

n=1 masked tool call (persimmon_n1.json)

  1. curl ... -d @persimmon_n1.json -o persimmon_n1.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. The first chunk's tool call reads {"id":"call_NADqNnHdBxx82AtliWXmXplc","function":{"arguments":"{\"fruit\": \"[KEYWORD_REDACTED]\"}","name":"lookup_fruit"},"type":"function","index":0}

n=2 text (text_n2.json)

  1. curl ... -d @text_n2.json -o text_n2.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. grep -oE '"index":[01],"delta":\{"content":"[^"]+"' text_n2.1.out prints "index":0,"delta":{"content":"The [KEYWORD_REDACTED] is ripe." and "index":1,"delta":{"content":"The [KEYWORD_REDACTED] is ripe."

n=1 text (text_n1.json)

  1. curl ... -d @text_n1.json -o text_n1.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. grep -oE '"content":"[^"]+"' text_n1.1.out prints "content":"The [KEYWORD_REDACTED] is ripe."

Non-streamed n=2 tool call (nonstream_n2.json)

  1. curl ... -d @nonstream_n2.json -o nonstream_n2.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. python3 -c 'import json; d=json.load(open("nonstream_n2.1.out")); print([(c["index"], c["message"]["tool_calls"][0]["function"]["arguments"]) for c in d["choices"]])' prints [(0, '{"fruit": "[KEYWORD_REDACTED]"}'), (1, '{"fruit": "[KEYWORD_REDACTED]"}')]

After (2a35dc5)

tool_permission on an n=2 tool call, three runs

  1. for run in 1 2 3; do curl ... -d @toolperm.json -o toolperm.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 200, run 2: HTTP 200, run 3: HTTP 200
  3. cat toolperm.1.out prints one chunk, data: {"id":"chatcmpl-EQ0GZevHMSR4bQ5ArLDSfouwMdV3w","created":1789866259,"model":"gpt-4.1-mini","object":"chat.completion.chunk","choices":[{"finish_reason":"stop","index":0,"delta":{"role":"assistant","tool_calls":[{"id":"call_DSgGIBOsD74A8TiADNbS7gsV","function":{"arguments":"{\"fruit\":\"blackberry-elderberry-gooseberry-huckleberry-lingonberry\"}","name":"lookup_fruit"},"type":"function","index":0},{"id":"call_DSgGIBOsD74A8TiADNbS7gsV","function":{"arguments":"{\"fruit\":\"blackberry\"}","name":"lookup_fruit"},"type":"function","index":1}]}}],"usage":{"completion_tokens":25,"prompt_tokens":67,"total_tokens":92, ...}} then data: [DONE] (the shorter second argument is the second sample's own output; runs 2 and 3 carry ids call_anAgl88K7lNkj9ZZITFa6zOP and call_1i3P5gXo9H0kjrzt1arC23hN)
  4. The proxy log for each run reads Tool Permission Guardrail: Found 2 tool calls and then Tool Permission Guardrail Post-Call Hook: All tools allowed

Masked word on an n=2 tool call (persimmon), two runs

  1. for run in 1 2; do curl ... -d @persimmon.json -o persimmon.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 200, run 2: HTTP 200
  3. The fragment one-liner on persimmon.1.out prints c0="" c0="{\" c0="fruit" c0="\" c1="" c1="{\" c0="pers" c1="fruit" c0="im" c1="\" c0="mon" c1="pers" c0="\" c1="im" c1="mon" c1="\": both choices assemble to {"fruit":"persimmon"}, still unmasked on the wire
  4. The proxy log reads Masked keyword 'persimmon' in content twice, then Pipeline: guardrail 'output-word-filter' rewrote the streamed response but the rewrite could not be written back to the stream: the stream carries 2 choices and tool-call rewrites are only written back on single-choice streams. The whole rewrite, text rewrites included, was discarded and the original stream released

Blocked word streamed in pieces (mangosteen), three runs

  1. for run in 1 2 3; do curl ... -d @mangosteen.json -o mangosteen.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 400, run 2: HTTP 400, run 3: HTTP 400
  3. cat mangosteen.1.out prints data: {"error": {"message": "Content blocked: keyword 'mangosteen' detected", "type": "invalid_request_error", "param": null, "code": "400", "provider_specific_fields": {"error": "Content blocked: keyword 'mangosteen' detected", "keyword": "mangosteen", "description": null, "guardrail_name": "output-word-filter", "guardrail_mode": "post_call"}}} then data: [DONE]

Blocked word streamed whole (kumquat), three runs

  1. for run in 1 2 3; do curl ... -d @kumquat.json -o kumquat.$run.out -w "run $run: HTTP %{http_code}\n"; done
  2. run 1: HTTP 400, run 2: HTTP 400, run 3: HTTP 400
  3. cat kumquat.1.out prints data: {"error": {"message": "Content blocked: keyword 'kumquat' detected", ... "guardrail_name": "output-word-filter", "guardrail_mode": "post_call"}} then data: [DONE]

n=1 masked tool call (persimmon_n1.json)

  1. curl ... -d @persimmon_n1.json -o persimmon_n1.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. The first chunk's tool call reads {"id":"call_DxRVYMFHzB8K2CnOXudR8QKo","function":{"arguments":"{\"fruit\": \"[KEYWORD_REDACTED]\"}","name":"lookup_fruit"},"type":"function","index":0}

n=2 text (text_n2.json)

  1. curl ... -d @text_n2.json -o text_n2.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. grep -oE '"index":[01],"delta":\{"content":"[^"]+"' text_n2.1.out prints "index":0,"delta":{"content":"The [KEYWORD_REDACTED] is ripe." and "index":1,"delta":{"content":"The [KEYWORD_REDACTED] is ripe."

n=1 text (text_n1.json)

  1. curl ... -d @text_n1.json -o text_n1.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. grep -oE '"content":"[^"]+"' text_n1.1.out prints "content":"The [KEYWORD_REDACTED] is ripe."

Non-streamed n=2 tool call (nonstream_n2.json)

  1. curl ... -d @nonstream_n2.json -o nonstream_n2.1.out -w 'HTTP %{http_code}\n' prints HTTP 200
  2. python3 -c 'import json; d=json.load(open("nonstream_n2.1.out")); print([(c["index"], c["message"]["tool_calls"][0]["function"]["arguments"]) for c in d["choices"]])' prints [(0, '{"fruit": "[KEYWORD_REDACTED]"}'), (1, '{"fruit": "[KEYWORD_REDACTED]"}')]

Observed next to the fix, none caused or worsened by this PR:

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • n>1 tool-call streams still fail open on the wire: the mask is computed and logged, never written back (pre-existing, LIT-8224)
  • tool_permission, bedrock, panw_prisma_airs, cisco_ai_defense, noma, model_armor, purview_dlp, and presidio (its output masking and unmasking paths) still re-emit an n>1 stream as one choice, now carrying every choice's tool calls with their shared ids on choice 0 (LIT-8224)

Low

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

  • 2a35dc5 passes /live-pr-risk (810acda, 2e83871, and 19d77e2 after it are comment, snapshot, and test-typing only)


Note

Medium Risk
Changes streaming reassembly and post-call guardrail discard paths used on live proxy streams; behavior for successful single-choice streams should be unchanged, but n>1 and packed-chunk tool-call shapes differ.

Overview
Fixes garbled tool-call JSON on n>1 (and multi-choice-per-chunk) streams by assembling streaming tool-call fragments under (choice index, tool index) instead of tool index alone, so post-call guardrails scan each choice’s arguments separately.

stream_chunk_builder now considers every choice in a chunk when detecting tool calls and when deciding the text-only fast path (via shared delta helpers), so tool calls on later choices are not dropped.

When a guardrail rewrite cannot be written back to the stream, UndeliverableStreamRewrite carries a specific reason; the pipeline executor logs that reason and still releases the original chunks. OpenAI, Anthropic, and Responses guardrail paths emit clearer messages (multi-choice streams, tool-call count mismatches, missing stop/terminal envelope, legacy-hook limits).

Tests cover per-choice argument assembly and the new exception reason fields.

Reviewed by Cursor Bugbot for commit 65160a9. Bugbot is set up for automated code reviews on this repo. Configure here.

…streams and log why a rewrite was discarded

The rebuilt streamed response keyed tool-call fragments by tool index alone, so
on n>1 chat streams the two choices' argument fragments were concatenated into
one string and post_call guardrails scanned garbled JSON. Fragments are now
keyed by (choice index, tool index).

When a guardrail's rewrite cannot be written back to the stream (multi-choice
streams, a rewrite that adds or drops a tool call, legacy-hook shapes the
translation cannot rescan), the pipeline now logs a warning naming the
guardrail and the exact reason before releasing the original stream.

Also commits the regenerated dashboard API types that make check produced.
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 13, 2026 08:49
@codspeed

codspeed Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_lit_7346_multi_choice_stream_guardrails (19d77e2) with main (4e55e99)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with no outstanding findings or newly introduced actionable issues

Summary

Fixes streamed tool-call reconstruction by grouping fragments by both choice and tool-call index. It also adds specific diagnostic reasons when guardrail rewrites cannot be written back, while preserving the original stream

  • Scans each choice's reconstructed tool-call arguments independently
  • Detects tool calls in later choices of packed stream chunks
  • Logs why text or tool-call rewrites were discarded
  • Adds regression coverage for multi-choice assembly and rewrite refusal reasons

Reviews (6) · Last reviewed commit: "test(guardrails): type the recorder hook..."

@codecov

codecov Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.09804% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
.../llms/openai/chat/guardrail_translation/handler.py 60.00% 2 Missing ⚠️
...itellm_core_utils/streaming_chunk_builder_utils.py 94.73% 1 Missing ⚠️
.../openai/responses/guardrail_translation/handler.py 94.11% 1 Missing ⚠️
litellm/main.py 94.73% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

…d multi-choice chunk

The rebuild's tool-call selection and its text-only fast path only looked at
choice 0 of each chunk, so a chunk that packs several choices (Gemini with
candidateCount above 1) lost a tool call carried by a later candidate, and a
chunk whose later choice had no tool calls at all made the rebuild raise.
Both now consider every choice in the chunk.
@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 15, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread litellm/litellm_core_utils/streaming_chunk_builder_utils.py Outdated
@greptile-apps

greptile-apps Bot commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Comments Outside Diff

These findings sit on lines the diff does not cover, so they could not be posted inline. Each one leaves this list once its file changes.

  • P1 Nonzero choices remain unscanned litellm/litellm_core_utils/streaming_chunk_builder_utils.py:478 ▶

    When only a later choice contains tool calls, main.py drops that chunk before this loop, so its arguments bypass guardrail scanning

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 19d77e2. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 4011367 into main Sep 20, 2026
133 of 150 checks passed
@mateo-berri
mateo-berri deleted the litellm_lit_7346_multi_choice_stream_guardrails branch September 20, 2026 03:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant