Skip to content

fix(anthropic): backport #42152 and #42288 to stable/1.100.x for v1.100.2 - #42532

Merged
mateo-berri merged 5 commits into
stable/1.100.xfrom
litellm_cherrypick_safeguards_1_100_x
Sep 22, 2026
Merged

mateo-berri merged 5 commits into
stable/1.100.xfrom
litellm_cherrypick_safeguards_1_100_x

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Claude Code auto mode sends a safeguards field on /v1/messages
  • Every released version strips it, so the gateway reply carries no safeguard_results
  • Claude Code then reports auto mode as unavailable through the gateway
  • From Thu Sep 25 those users get a billing notice and are told to contact their gateway provider

How it solves it:

User Flow

Before: a developer running Claude Code through a LiteLLM gateway on this release line cannot use auto mode, because the gateway drops the field Claude Code needs for it

  1. They start claude with ANTHROPIC_BASE_URL=https://litellm-domain and switch to auto mode (--permission-mode auto, or shift+tab)
  2. Claude Code sends POST https://litellm-domain/v1/messages with a safeguards array in the body asking for the server-side classifier
  3. The 200 reply carries no safeguard_results, so /status shows Auto mode server: Disabled and every tool call is judged by the billed client-side classifier instead
  4. From Thu Sep 25 the same session prompts them to keep the paid client-side classifier and to contact their gateway provider

After: the same developer gets auto mode through the gateway, with the classifier answers coming back from the provider

  1. They start claude with ANTHROPIC_BASE_URL=https://litellm-domain and switch to auto mode (--permission-mode auto, or shift+tab)
  2. Claude Code sends POST https://litellm-domain/v1/messages with a safeguards array in the body asking for the server-side classifier
  3. The 200 reply carries safeguard_results (type: evaluated, outcome: not_flagged per tool call), so /status shows Auto mode server: Enabled and the server-side classifier judges every tool call
  4. On Thu Sep 25 nothing changes for them

Relevant issues

Backport of #42152 and #42288

Affected release

Linear ticket

Resolves LIT-8336

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Backport checklist

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Before (44d3c4f)

Proxy booted from 44d3c4f290 with --num_workers 2 on port 54616, no DB; Claude Code v2.1.280 started with ANTHROPIC_BASE_URL=http://localhost:54616 and --permission-mode auto, model anthropic-sonnet-5

Claude Code auto mode through the gateway

Read and Bash tool calls run in auto mode, but /status reports Auto mode server: Disabled: the gateway dropped the safeguards field, so Claude Code falls back to its billed client-side classifier

pr42532-8ccfd6a84e-before_1_100_x_status.png
pr42532-8ccfd6a84e-before_1_100_x_bash.png

POST /v1/messages with safeguards (non-streaming and streaming, three providers)

All six runs answer 200 with safeguard_results: []: the beta header reached Anthropic but the safeguards field did not

$ MODEL=anthropic-sonnet-5 BASE_URL=http://localhost:54616 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01Pma9ttcqMchfjA4smYSZBk']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01MJMuy948Qc45uW2CY9DzxF']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
$ MODEL=bedrock-sonnet-5 BASE_URL=http://localhost:54616 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_01XDDfmmQf7FAERNc1F2eERp']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_011yzyWCmt7Lo7EiMZXbpive']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
$ MODEL=vertex-sonnet-5 BASE_URL=http://localhost:54616 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_01HR7edErWL11Ah1Qw3XFT8J']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_01Uw84URFtGjsK17PXZXjdc2']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request

After (8ccfd6a)

Proxy booted from 8ccfd6a84e with --num_workers 2 on port 20049, no DB; Claude Code v2.1.280 started with ANTHROPIC_BASE_URL=http://localhost:20049 and --permission-mode auto, model anthropic-sonnet-5

Claude Code auto mode through the gateway

The same Read and Bash tool calls run in auto mode and /status reports Auto mode server: Enabled: the server-side classifier answers inside the main request

pr42532-8ccfd6a84e-after_1_100_x_status.png
pr42532-8ccfd6a84e-after_1_100_x_bash.png

POST /v1/messages with safeguards (non-streaming and streaming, three providers)

All six runs answer 200 with one dangerous_tool_use entry whose status.type is available and whose per-call result is evaluated

$ MODEL=anthropic-sonnet-5 BASE_URL=http://localhost:20049 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01Nr78vLiSf8SB3wGRLembby']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_01Nr78vLiSf8SB3wGRLembby": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_01Nr78vLiSf8SB3wGRLembby evaluated)
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01GTTQafSSmMgHkPfkGVZABh']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_01GTTQafSSmMgHkPfkGVZABh": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_01GTTQafSSmMgHkPfkGVZABh evaluated)
$ MODEL=bedrock-sonnet-5 BASE_URL=http://localhost:20049 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_01HYqhvTUYp5TvyzFjxcvxiz']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_bdrk_01HYqhvTUYp5TvyzFjxcvxiz": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_bdrk_01HYqhvTUYp5TvyzFjxcvxiz evaluated)
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_01MwmffLwWB8FZekuxhRXFdR']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_bdrk_01MwmffLwWB8FZekuxhRXFdR": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_bdrk_01MwmffLwWB8FZekuxhRXFdR evaluated)
$ MODEL=vertex-sonnet-5 BASE_URL=http://localhost:20049 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_01U3KwsiqQdgV6f8iBPacXam']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_vrtx_01U3KwsiqQdgV6f8iBPacXam": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_vrtx_01U3KwsiqQdgV6f8iBPacXam evaluated)
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_01WZeMLhre7Hwm7Teu65MUZV']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_vrtx_01WZeMLhre7Hwm7Teu65MUZV": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_vrtx_01WZeMLhre7Hwm7Teu65MUZV evaluated)

The QA proof above ran at 8ccfd6a84e. The only commit since, 4438739b46 (the #42048 test port), touches tests/proxy_e2e_anthropic_messages_tests/ only, so it cannot change runtime behavior and the proof stands at that hash

Type

🐛 Bug Fix

Caveats (if any)

  • Low (introduced by the first push of this PR, fixed by its tip commit 4438739b46): test_bedrock_invoke_messages_with_all_beta_headers[bedrock-claude-opus-4.5-bedrock] failed on the first CircleCI run. The picked beta-header map adds dangerous-tool-use-2026-09-03 for Bedrock, and Claude Opus 4.5 on Bedrock Invoke rejects it as an invalid beta flag. main has run those cases on Claude Fable 5.1 since fix(test): unbreak the integration-cost and proxy_e2e_anthropic_messages CircleCI jobs on main #42048 (7966f50c34), ported here as 4438739b46; the CircleCI run at the tip no longer lists it
  • Low (inherited, ignored): CircleCI llm_translation_testing, local_testing_part1 and local_testing_part2 fail on the TogetherAI cases only (the 18 tests/llm_translation/test_together_ai.py cases, test_completion_custom_provider_model_name, test_completion_together_ai_stream, test_customprompt_together_ai, test_async_text_completion_together_ai). TogetherAI retired the serverless openai/gpt-oss-20b those tests pin (its deprecation_date is 2026-09-14), so they fail the same way at this line's base; main fixed the suites in 1aa2e19ee4, 8c046e13bd, b478131701, 515bf8c9d5 and 02ced74540. Nothing in this PR touches TogetherAI
  • Low (inherited, ignored): e2e_ui_testing uploads no junit results and fails on tests/e2e/ui/tests/mcp/mcpTools.spec.ts (the MCP Tools tab no longer lists the ask_question card the spec hardcodes). main rewrote the spec to check against the upstream's own tools/list in test(e2e-ui): check the MCP Tools tab against the upstream's own tools/list #42397 (b96842f62c). No dashboard code changes here
  • Low (inherited, ignored): test_completion_bedrock_httpx_models[...-bedrock/cohere.command-r-plus-v1:0] (both parametrizations) and test_parallel_streaming_requests[cohere.command-r-plus-v1:0-False] fail because Bedrock retired that Cohere model. main dropped the cases in test(bedrock): drop EOL cohere.command-r-plus-v1:0 from local_testing #39608 (f87b9097ea). Unrelated to this change
  • Low (inherited, ignored): test_bad_request_error in tests/openai_endpoints_tests/test_e2e_openai_responses_api.py still expects a 400 for an unknown model on /v1/responses; OpenAI returns a 404 now and main rewrote the test in d33fe95d19. The route is untouched by this PR
  • Low (inherited, ignored): test_timeout_streaming and test_router_timeout are the known flaky timeout tests, green on main the same day and untouched by this PR
  • Low: the stable/** ruleset requires one approving review and has no required status checks, so the inherited CircleCI reds above do not gate the merge and are recorded here instead of merging main into the line; every GitHub Actions check is green at the tip

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

yassin-berriai and others added 4 commits September 22, 2026 10:16
…tive /v1/messages

Backport of #42152 to stable/1.100.x.
Cherry-picked from merge commit e912ebe (litellm_claude_code_safeguards_passthrough).
… beta to Bedrock Invoke and Vertex on /v1/messages

Backport of #42288 to stable/1.100.x.
Cherry-picked from merge commit fc82f6e (litellm_safeguards_bedrock_vertex_messages).
The line has no bedrock_mantle beta-header mapping and no Mantle /v1/messages route, so the Mantle mapping, its test file, and the bedrock_mantle test parameter are left out.
Hand-ported to stable/1.100.x from 47b2479 on main (fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag), the one prerequisite the #42288 handler tests need; the rest of that commit stays on main.
…ine's budgets

The picked TypedDict fields use read-only Sequence[Mapping[str, object]] annotations and the picked Vertex test carries a test-quality-ok marker, so stable/1.100.x's LIT001, LIT012 and TQ008 budgets hold. Static typing only, no runtime change.
@devin-ai-integration

devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 22, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 3 committers have signed the CLA.

✅ mateo-berri
❌ yassin-berriai
❌ kerry-berri
You have signed the CLA already but the status is still pending? Let us recheck it.

Backport of #42048 to stable/1.100.x.
Cherry-picked from 7966f50 (main). The safeguards backport maps the dangerous-tool-use-2026-09-03 beta for Bedrock, which Claude Opus 4.5 on Bedrock Invoke rejects as an invalid beta flag, so the all-beta-headers Bedrock cases run on Claude Fable 5.1 as they do on main.
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 22, 2026
@mateo-berri
mateo-berri marked this pull request as ready for review September 22, 2026 21:29
@mateo-berri
mateo-berri requested a review from a team September 22, 2026 21:29

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 35f5b7c into stable/1.100.x Sep 22, 2026
43 of 49 checks passed
@mateo-berri
mateo-berri deleted the litellm_cherrypick_safeguards_1_100_x branch September 22, 2026 21:29
@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The runtime backport appears correct, but the repository requirements must be satisfied by fixing the network-dependent teardown and removing prohibited comments before merging

Findings

  1. P2 Remote teardown reload ▶
  2. P2 Prohibited explanatory comments ▶

Summary

This backport adds Claude Code safeguards support across native Anthropic Messages, Bedrock Invoke, and Vertex, including beta-header handling and response preservation

  • Adds safeguards request typing and safeguard_results response typing
  • Injects the dangerous-tool-use beta for Bedrock Invoke and Vertex when safeguards are present
  • Preserves complete and streaming safeguard results through provider pass-through paths
  • Adds focused provider and proxy regression coverage
  • Requires cleanup of a network-dependent fixture teardown and repository-prohibited explanatory test comments

Reviews (1) · Last reviewed commit: "fix(test): run the all-beta-headers bedr..."

Comment on lines +216 to +217
finally:
monkeypatch.delenv("LITELLM_LOCAL_ANTHROPIC_BETA_HEADERS", raising=False)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Remote teardown reload

Deleting local mode before reloading triggers an unmocked GitHub request and leaves cached beta configuration inconsistent after environment restoration

Rule Used: What: prevent any tests from being added here that make real network calls - only mock tests can be added to this folder. Exception: tests/e2e folder, this folder can contain network calls. Why: ensure it runs correctly on github ci/cd + for all d... (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +1656 to +1663
def test_bedrock_messages_forwards_safeguards_with_dangerous_tool_use_beta(local_beta_headers_config, client_beta_header):
"""
Claude Code's server-side auto-mode classifier sends `safeguards` alongside the
dangerous-tool-use-2026-09-03 beta. Bedrock Invoke accepts the pair, answers
"safeguards: Extra inputs are not permitted" for the field alone, and returns
`safeguard_results: []` for the beta alone, so the field reaches it unchanged
and the beta rides along whether or not the client sent it, as every other
body-driven beta does here.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Prohibited explanatory comments

This test adds ordinary explanatory prose, violating the repository directive limiting comments. The same pattern appears elsewhere and must be removed before merging

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants