Skip to content

fix(bedrock): forward anthropic-beta headers verbatim on the Claude platform messages path - #42275

Merged
mateo-berri merged 1 commit into
mainfrom
litellm_claude_platform_messages_beta_passthrough
Sep 21, 2026
Merged

mateo-berri merged 1 commit into
mainfrom
litellm_claude_platform_messages_beta_passthrough

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • The Claude Platform messages route now states outright: never filter betas
  • The two contrast tests assert filtering on Azure AI, which really filters
  • A new request-level test fails if this route ever filters again

User Flow

Before: a developer calling Claude Platform on AWS through the gateway with a beta feature gets a 200 today, but the tests describing this route say the gateway should strip that beta, and the CI fix that follows those tests (#42260) turns the same call into a 400

  1. They add bedrock/claude_platform/claude-haiku-4-5-20251001 with their AWS credentials and workspace_id to the proxy config
  2. They send POST https://litellm-domain/v1/messages with the header anthropic-beta: mcp-client-2025-11-20 and an mcp_servers entry in the body
  3. They get HTTP 200 with the model's reply
  4. The repo's CI on main shows two failing tests about this route, both expecting the gateway to strip betas Bedrock does not know
  5. With those tests satisfied the way fix(bedrock): keep filtering anthropic-beta headers on the Claude platform messages path #42260 does it, the same POST returns HTTP 400 mcp_servers: this parameter requires anthropic-beta: mcp-client-2026-09-15 (or mcp-client-2025-11-20), and a cache_control with scope: global returns HTTP 400 system.0.cache_control.ephemeral.scope: Extra inputs are not permitted

After: the same call keeps working, CI is green, and a change that starts stripping betas on this route fails CI

  1. They add bedrock/claude_platform/claude-haiku-4-5-20251001 with their AWS credentials and workspace_id to the proxy config
  2. They send POST https://litellm-domain/v1/messages with the header anthropic-beta: mcp-client-2025-11-20 and an mcp_servers entry in the body
  3. They get HTTP 200 with the model's reply, the anthropic-beta header reaching the gateway exactly as sent
  4. The repo's CI on main is green: the two tests now check the stripping against a route that really strips (Azure AI)
  5. Any change that turns stripping on for this route fails a test that sends a real request shape through it

Relevant issues

Supersedes #42260

Affected release

Linear ticket

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Setup shared by every leg. One proxy process per side, each booted from its own worktree and venv with --num_workers 2 on its own port, the same config on all of them, real SigV4 calls to aws-external-anthropic.us-west-2.api.aws (the Claude Platform on AWS gateway), real spend

model_list:
  - model_name: claude-haiku-4-5
    litellm_params:
      model: bedrock/claude_platform/claude-haiku-4-5-20251001
      aws_region_name: us-west-2
      aws_profile_name: <SigV4 profile>
      workspace_id: wrkspc_<id>
general_settings:
  master_key: sk-1234
litellm --config config.yaml --port <port> --num_workers 2 --detailed_debug

The four requests, run in this order on every side (<port> is that side's port):

# 1. cache_control scope=global, gated by the prompt-caching-scope-2026-01-05 beta
curl -s -w '\nHTTP %{http_code}\n' http://127.0.0.1:<port>/v1/messages \
  -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' \
  -H 'anthropic-version: 2023-06-01' -H 'anthropic-beta: prompt-caching-scope-2026-01-05' \
  -d '{"model":"claude-haiku-4-5","max_tokens":8,"system":[{"type":"text","text":"Answer with one word.","cache_control":{"type":"ephemeral","scope":"global"}}],"messages":[{"role":"user","content":"Say ok"}]}'

# 2. mcp_servers, gated by the mcp-client-2025-11-20 beta
curl -s -w '\nHTTP %{http_code}\n' http://127.0.0.1:<port>/v1/messages \
  -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' \
  -H 'anthropic-version: 2023-06-01' -H 'anthropic-beta: mcp-client-2025-11-20' \
  -d '{"model":"claude-haiku-4-5","max_tokens":8,"mcp_servers":[{"type":"url","url":"https://mcp.deepwiki.com/mcp","name":"deepwiki"}],"tools":[{"type":"mcp_toolset","mcp_server_name":"deepwiki"}],"messages":[{"role":"user","content":"Say ok"}]}'

# 3. the Anthropic Python SDK (anthropic 0.84.0) sending the same mcp_servers request
client = anthropic.Anthropic(base_url="http://127.0.0.1:<port>", api_key="sk-1234")
client.beta.messages.create(model="claude-haiku-4-5", max_tokens=8, betas=["mcp-client-2025-11-20"],
    mcp_servers=[{"type": "url", "url": "https://mcp.deepwiki.com/mcp", "name": "deepwiki"}],
    tools=[{"type": "mcp_toolset", "mcp_server_name": "deepwiki"}],
    messages=[{"role": "user", "content": "Say ok"}])

# 4. control, no beta at all
curl -s -w '\nHTTP %{http_code}\n' http://127.0.0.1:<port>/v1/messages \
  -H 'Authorization: Bearer sk-1234' -H 'content-type: application/json' \
  -H 'anthropic-version: 2023-06-01' \
  -d '{"model":"claude-haiku-4-5","max_tokens":8,"messages":[{"role":"user","content":"Say ok"}]}'

Before (36b8be7)

The two contrast tests on main

  1. uv run pytest tests/unit/llms/github_copilot/messages/test_github_copilot_messages_transformation.py::test_github_copilot_config_disables_anthropic_beta_filtering tests/unit/llms/openai_like/messages/test_openai_like_anthropic_messages_transformation.py::test_passthrough_disables_anthropic_beta_filtering
  2. Observed:
>       assert AnthropicMessagesConfig().should_filter_anthropic_beta_headers() is True
E       AssertionError: assert False is True
tests/unit/llms/github_copilot/messages/test_github_copilot_messages_transformation.py:281: AssertionError
>       assert AnthropicMessagesConfig().should_filter_anthropic_beta_headers() is True
E       AssertionError: assert False is True
tests/unit/llms/openai_like/messages/test_openai_like_anthropic_messages_transformation.py:276: AssertionError
2 failed in 1.11s

cache_control scope=global with the prompt-caching-scope beta

  1. Request 1 above
  2. Observed (the live behavior on main already passes the beta through, nothing guards it):
{"model":"claude-haiku-4-5","id":"msg_011CfH5r8dekBUcvYsgFjXYQ","type":"message","role":"assistant","content":[{"type":"text","text":"Ok"}],"container":null,"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":14,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard","inference_geo":"not_available"}}
HTTP 200

mcp_servers with the mcp-client beta

  1. Request 2 above
  2. Observed:
{"model":"claude-haiku-4-5","id":"msg_011CfH5rR38KbWPzRnwJxM7y","type":"message","role":"assistant","content":[{"type":"text","text":"ok"}],"container":null,"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":828,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard","inference_geo":"not_available"}}
HTTP 200

Anthropic Python SDK, mcp_servers with the mcp-client beta

  1. Request 3 above
  2. Observed: HTTP 200 ok input_tokens 828

Control, no beta

  1. Request 4 above
  2. Observed:
{"model":"claude-haiku-4-5","id":"msg_011CfH5rUcvt7CdkAG6HSFfB","type":"message","role":"assistant","content":[{"type":"text","text":"ok"}],"container":null,"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":9,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard","inference_geo":"not_available"}}
HTTP 200

After (e51ccbc)

The two contrast tests on main

  1. uv run pytest tests/unit/llms/github_copilot/messages/test_github_copilot_messages_transformation.py::test_github_copilot_config_disables_anthropic_beta_filtering tests/unit/llms/openai_like/messages/test_openai_like_anthropic_messages_transformation.py::test_passthrough_disables_anthropic_beta_filtering tests/test_litellm/llms/bedrock/test_claude_platform_provider.py::test_anthropic_messages_bedrock_claude_platform_forwards_anthropic_beta_verbatim
  2. Observed: 3 passed, 1 warning in 0.72s (the third is the new request-level regression test; it fails when this route's filter is switched on, which is what fix(bedrock): keep filtering anthropic-beta headers on the Claude platform messages path #42260 did)

cache_control scope=global with the prompt-caching-scope beta

  1. Request 1 above
  2. Observed:
{"model":"claude-haiku-4-5","id":"msg_011CfH5gZYKfGM5ExNtqaDPR","type":"message","role":"assistant","content":[{"type":"text","text":"Ok"}],"container":null,"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":14,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard","inference_geo":"not_available"}}
HTTP 200

mcp_servers with the mcp-client beta

  1. Request 2 above
  2. Observed:
{"model":"claude-haiku-4-5","id":"msg_011CfH5hAbU8PSijk3tSHXZs","type":"message","role":"assistant","content":[{"type":"text","text":"ok"}],"container":null,"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":828,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard","inference_geo":"not_available"}}
HTTP 200

Anthropic Python SDK, mcp_servers with the mcp-client beta

  1. Request 3 above
  2. Observed: HTTP 200 ok input_tokens 828

Control, no beta

  1. Request 4 above
  2. Observed:
{"model":"claude-haiku-4-5","id":"msg_011CfH5hdCruLbwa5qsiju3G","type":"message","role":"assistant","content":[{"type":"text","text":"ok"}],"container":null,"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":9,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":0},"output_tokens":4,"service_tier":"standard","inference_geo":"not_available"}}
HTTP 200

Contrast: the same four requests at #42260's head (cac3628)

This is what the new regression test guards against. Same config, same two-worker boot, same order

  1. Request 1: HTTP 400 with system.0.cache_control.ephemeral.scope: Extra inputs are not permitted (the prompt-caching-scope-2026-01-05 beta was stripped while the body kept scope)
  2. Request 2: HTTP 400 with mcp_servers: this parameter requires anthropic-beta: mcp-client-2026-09-15 (or mcp-client-2025-11-20) (the mcp-client-2025-11-20 beta was stripped)
  3. Request 3 (SDK): HTTP 400 Error code: 400 ... mcp_servers: this parameter requires anthropic-beta: mcp-client-2026-09-15 (or mcp-client-2025-11-20)
  4. Request 4 (control): HTTP 200

Observations from the run:

  • /v1/chat/completions on this deployment 400s on both sides: workspace_id lands in the body (LIT-8279)
  • That chat 400 is pre-existing on main; this PR leaves it alone
  • No legacy CircleCI suite reaches this route, so run-ci adds nothing here

Type

🐛 Bug Fix
✅ Test

Caveats (if any)

Low

  • Live behavior on main already passed these betas; the override makes it explicit instead of inherited, so the before leg's failure is the CI red, not a live 4xx
  • /v1/chat/completions on the same deployment 400s on main and here alike (LIT-8279)
  • The existing header merge sorts anthropic-beta values, so "verbatim" is the set, not the order; pre-existing and left alone
  • The new test patches AsyncHTTPHandler.post like its five siblings in the file instead of injecting a client; rewriting the file's capture pattern is out of scope for a three-line fix

Live PR risk

  • Breaking: none observed. Base and tip answered every /v1/messages case and a /v1/chat/completions parity case with the same status and body shape
  • Backward incompatible: none. The filter switch already answered "no" for this route on the base; the override pins that answer
  • Regression risk: the outbound anthropic-beta header was not read off the wire, since SigV4 signs the Host header and a forwarding recorder cannot sit in front of the gateway. The gateway's own beta-gated 400 vs 200 answers, plus the new test's capture of the outbound headers, stand in for it
  • Dependency graph: one caller of the switch (the shared messages HTTP handler), verified live; five implementations of the switch (base, Anthropic, GitHub Copilot, OpenAI-like, and this one), the other four untouched and covered by their own suites; no subclass of the changed config; the factory that returns it is reached only by an explicit bedrock/claude_platform/ model; main moved 17 commits since the merge base, none touching this surface
  • Not verified: whether /v1/chat/completions on this route filters betas (every request on it 400s on workspace_id first, see caveats)

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

  • e51ccbc passes /live-pr-risk

@devin-ai-integration

devin-ai-integration Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_claude_platform_messages_beta_passthrough (e51ccbc) with main (36b8be7)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (4b2e96a) during the generation of this report, so 36b8be7 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@greptile-apps

greptile-apps Bot commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with focused provider behavior and regression coverage matching the Claude Platform Messages route

Summary

This PR disables Bedrock-specific Anthropic beta filtering for the Claude Platform Messages route, allowing the complete caller-provided beta-token set to reach its Anthropic-compatible endpoint. It adds request-level regression coverage and updates existing passthrough tests to use Azure as the positive filtering control

Reviews (1) · Last reviewed commit: "fix(bedrock): forward anthropic-beta hea..."

@codecov

codecov Bot commented Sep 21, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit e51ccbc. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit 9a90ada into main Sep 21, 2026
95 of 97 checks passed
@mateo-berri
mateo-berri deleted the litellm_claude_platform_messages_beta_passthrough branch September 21, 2026 19:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant