Skip to content

fix(anthropic): backport #42152 and #42288 to stable/1.101.x for v1.101.1 - #42533

Merged
mateo-berri merged 5 commits into
stable/1.101.xfrom
litellm_cherrypick_safeguards_1_101_x
Sep 22, 2026
Merged

mateo-berri merged 5 commits into
stable/1.101.xfrom
litellm_cherrypick_safeguards_1_101_x

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Claude Code auto mode sends a safeguards field on /v1/messages
  • Every released version strips it, so the gateway reply carries no safeguard_results
  • Claude Code then reports auto mode as unavailable through the gateway
  • From Thu Sep 25 those users get a billing notice and are told to contact their gateway provider

How it solves it:

User Flow

Before: a developer running Claude Code through a LiteLLM gateway on this release line cannot use auto mode, because the gateway drops the field Claude Code needs for it

  1. They start claude with ANTHROPIC_BASE_URL=https://litellm-domain and switch to auto mode (--permission-mode auto, or shift+tab)
  2. Claude Code sends POST https://litellm-domain/v1/messages with a safeguards array in the body asking for the server-side classifier
  3. The 200 reply carries no safeguard_results, so /status shows Auto mode server: Disabled and every tool call is judged by the billed client-side classifier instead
  4. From Thu Sep 25 the same session prompts them to keep the paid client-side classifier and to contact their gateway provider

After: the same developer gets auto mode through the gateway, with the classifier answers coming back from the provider

  1. They start claude with ANTHROPIC_BASE_URL=https://litellm-domain and switch to auto mode (--permission-mode auto, or shift+tab)
  2. Claude Code sends POST https://litellm-domain/v1/messages with a safeguards array in the body asking for the server-side classifier
  3. The 200 reply carries safeguard_results (type: evaluated, outcome: not_flagged per tool call), so /status shows Auto mode server: Enabled and the server-side classifier judges every tool call
  4. On Thu Sep 25 nothing changes for them

Relevant issues

Backport of #42152 and #42288

Affected release

Linear ticket

Resolves LIT-8336

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Backport checklist

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Before (1fb7e5a)

Proxy booted from 1fb7e5a9e9 with --num_workers 2 on port 50358, no DB; Claude Code v2.1.280 started with ANTHROPIC_BASE_URL=http://localhost:50358 and --permission-mode auto, model anthropic-sonnet-5

Claude Code auto mode through the gateway

Read and Bash tool calls run in auto mode, but /status reports Auto mode server: Disabled: the gateway dropped the safeguards field, so Claude Code falls back to its billed client-side classifier

pr42533-3c6a66dc48-before_1_101_x_status.png
pr42533-3c6a66dc48-before_1_101_x_bash.png

POST /v1/messages with safeguards (non-streaming and streaming, three providers)

All six runs answer 200 with safeguard_results: []: the beta header reached Anthropic but the safeguards field did not

$ MODEL=anthropic-sonnet-5 BASE_URL=http://localhost:50358 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01WhXgh5P3YRhZoaGtyyzZBb']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01Dww9HXxaDmCLZ4CSRGU6Gx']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
$ MODEL=bedrock-sonnet-5 BASE_URL=http://localhost:50358 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_01JJ44t1dUvJ2LVMaRXUStQT']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_01RVos7k2JzuRfqR6AgE4mSs']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
$ MODEL=vertex-sonnet-5 BASE_URL=http://localhost:50358 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_016s7VphNrZXrRhbcvEEd3g3']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_01GPpKM1ztSeNihbisuuGUdw']
safeguard_results: []
FAIL: safeguard_results is [] - header forwarded, but the safeguards field was dropped from the request

After (3c6a66d)

Proxy booted from 3c6a66dc48 with --num_workers 2 on port 42480, no DB; Claude Code v2.1.280 started with ANTHROPIC_BASE_URL=http://localhost:42480 and --permission-mode auto, model anthropic-sonnet-5

Claude Code auto mode through the gateway

The same Read and Bash tool calls run in auto mode and /status reports Auto mode server: Enabled: the server-side classifier answers inside the main request

pr42533-3c6a66dc48-after_1_101_x_status.png
pr42533-3c6a66dc48-after_1_101_x_bash.png

POST /v1/messages with safeguards (non-streaming and streaming, three providers)

All six runs answer 200 with one dangerous_tool_use entry whose status.type is available and whose per-call result is evaluated

$ MODEL=anthropic-sonnet-5 BASE_URL=http://localhost:42480 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01Mn8txgVsiaJbEwd16qQseV']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_01Mn8txgVsiaJbEwd16qQseV": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_01Mn8txgVsiaJbEwd16qQseV evaluated)
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_01QeiNVuJVsKAkmJrXFwJXoR']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_01QeiNVuJVsKAkmJrXFwJXoR": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_01QeiNVuJVsKAkmJrXFwJXoR evaluated)
$ MODEL=bedrock-sonnet-5 BASE_URL=http://localhost:42480 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_01ASoipB9WQfh8N8nrFyQFir']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_bdrk_01ASoipB9WQfh8N8nrFyQFir": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_bdrk_01ASoipB9WQfh8N8nrFyQFir evaluated)
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_bdrk_0154fyxWQqkQthx5oMLbQVLq']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_bdrk_0154fyxWQqkQthx5oMLbQVLq": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_bdrk_0154fyxWQqkQthx5oMLbQVLq evaluated)
$ MODEL=vertex-sonnet-5 BASE_URL=http://localhost:42480 KEY=$LITELLM_MASTER_KEY AUTH=bearer python3 check_gateway.py
== non-streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_01Vy3pgMsfRQqdY61oLajYCa']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_vrtx_01Vy3pgMsfRQqdY61oLajYCa": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_vrtx_01Vy3pgMsfRQqdY61oLajYCa evaluated)
== streaming
anthropic-beta: dangerous-tool-use-2026-09-03
HTTP 200  request-id: (none - the gateway did not forward it)
tool_use ids: ['toolu_vrtx_01MN87ZF8oZeMdWd9TZs12Lg']
safeguard_results: [
  {
    "type": "dangerous_tool_use",
    "status": {
      "type": "available",
      "tool_uses": {
        "toolu_vrtx_01MN87ZF8oZeMdWd9TZs12Lg": {
          "type": "evaluated",
          "outcome": "not_flagged"
        }
      }
    }
  }
]
PASS (toolu_vrtx_01MN87ZF8oZeMdWd9TZs12Lg evaluated)

The QA proof above ran at 3c6a66dc48. The only commit since, d4ca007b6c (the #42048 test port), touches tests/proxy_e2e_anthropic_messages_tests/ only, so it cannot change runtime behavior and the proof stands at that hash

Type

🐛 Bug Fix

Caveats (if any)

  • Low (introduced by the first push of this PR, fixed by its tip commit d4ca007b6c): test_bedrock_invoke_messages_with_all_beta_headers[bedrock-claude-opus-4.5-bedrock] failed on the first CircleCI run. The picked beta-header map adds dangerous-tool-use-2026-09-03 for Bedrock, and Claude Opus 4.5 on Bedrock Invoke rejects it as an invalid beta flag. main has run those cases on Claude Fable 5.1 since fix(test): unbreak the integration-cost and proxy_e2e_anthropic_messages CircleCI jobs on main #42048 (7966f50c34), ported here as d4ca007b6c; the CircleCI run at the tip no longer lists it
  • Low (inherited, ignored): CircleCI llm_translation_testing, local_testing_part1 and local_testing_part2 fail on the TogetherAI cases only (the 18 tests/llm_translation/test_together_ai.py cases, test_completion_custom_provider_model_name, test_completion_together_ai_stream, test_customprompt_together_ai, test_async_text_completion_together_ai). TogetherAI retired the serverless openai/gpt-oss-20b those tests pin (its deprecation_date is 2026-09-14), so they fail the same way at this line's base; main fixed the suites in 1aa2e19ee4, 8c046e13bd, b478131701, 515bf8c9d5 and 02ced74540. Nothing in this PR touches TogetherAI
  • Low (inherited, ignored): e2e_ui_testing uploads no junit results and fails on tests/e2e/ui/tests/mcp/mcpTools.spec.ts (the MCP Tools tab no longer lists the ask_question card the spec hardcodes). main rewrote the spec to check against the upstream's own tools/list in test(e2e-ui): check the MCP Tools tab against the upstream's own tools/list #42397 (b96842f62c). No dashboard code changes here
  • Low (inherited, ignored): test_litellm_gateway_from_sdk_embedding[True] are the known flaky timeout tests, green on main the same day and untouched by this PR
  • Low: the stable/** ruleset requires one approving review and has no required status checks, so the inherited CircleCI reds above do not gate the merge and are recorded here instead of merging main into the line; every GitHub Actions check is green at the tip

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

yassin-berriai and others added 4 commits September 22, 2026 10:16
…tive /v1/messages

Backport of #42152 to stable/1.101.x.
Cherry-picked from merge commit e912ebe (litellm_claude_code_safeguards_passthrough).
… beta to Bedrock Invoke and Vertex on /v1/messages

Backport of #42288 to stable/1.101.x.
Cherry-picked from merge commit fc82f6e (litellm_safeguards_bedrock_vertex_messages).
The line has no bedrock_mantle beta-header mapping and no Mantle /v1/messages route, so the Mantle mapping, its test file, and the bedrock_mantle test parameter are left out.
Hand-ported to stable/1.101.x from 47b2479 on main (fix(bedrock): gate Invoke tool search on the model map's supports_tool_search flag), the one prerequisite the #42288 handler tests need; the rest of that commit stays on main.
…ine's budgets

The picked TypedDict fields use read-only Sequence[Mapping[str, object]] annotations and the picked Vertex test carries a test-quality-ok marker, so stable/1.101.x's LIT001, LIT012 and TQ008 budgets hold. Static typing only, no runtime change.
@devin-ai-integration

devin-ai-integration Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 22, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 3 committers have signed the CLA.

✅ mateo-berri
❌ yassin-berriai
❌ kerry-berri
You have signed the CLA already but the status is still pending? Let us recheck it.

Backport of #42048 to stable/1.101.x.
Cherry-picked from 7966f50 (main). The safeguards backport maps the dangerous-tool-use-2026-09-03 beta for Bedrock, which Claude Opus 4.5 on Bedrock Invoke rejects as an invalid beta flag, so the all-beta-headers Bedrock cases run on Claude Fable 5.1 as they do on main.
@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 22, 2026
@mateo-berri
mateo-berri marked this pull request as ready for review September 22, 2026 21:29
@mateo-berri
mateo-berri requested a review from a team September 22, 2026 21:29

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 4b432a9 into stable/1.101.x Sep 22, 2026
44 of 49 checks passed
@mateo-berri
mateo-berri deleted the litellm_cherrypick_safeguards_1_101_x branch September 22, 2026 21:29
@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The runtime change appears correct, but the repository’s dependency-injection testing requirement must be satisfied before merging

Findings

  1. P2 Avoid Patched Vertex Methods ▶
  2. P2 Cache Reload Uses Wrong State ▶

Summary

This backport forwards Anthropic safeguards through direct Anthropic, Bedrock Invoke, and Vertex requests while preserving safeguard results in complete and streamed responses

  • Adds the dangerous-tool-use beta mapping and provider-specific beta injection
  • Extends request and response type contracts for safeguards
  • Adds regression coverage across direct Anthropic, Bedrock, and Vertex paths
  • Updates Bedrock all-beta-header E2E coverage to use Claude Fable 5.1
  • Introduces two test-maintainability concerns involving dependency patching and global beta-config cleanup

Reviews (1) · Last reviewed commit: "fix(test): run the all-beta-headers bedr..."

Comment on lines +1578 to +1580

safeguards, safeguard_results = _claude_code_auto_mode_request()
captured: dict[str, object] = {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Avoid patched Vertex methods

These tests patch Vertex methods instead of injecting dependencies, violating a repository requirement that must be satisfied before merging; the Vertex config test does likewise

Context Used: CLAUDE.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +216 to +218
finally:
monkeypatch.delenv("LITELLM_LOCAL_ANTHROPIC_BETA_HEADERS", raising=False)
reload_beta_headers_config()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Cache reload uses wrong state

Teardown deletes the local override before reloading the global cache, triggering network access and leaving later tests with cached state that mismatches the environment

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants