Skip to content

fix(bedrock): never emit Converse cachePoint for OpenAI-family models - #41419

Merged
mateo-berri merged 2 commits into
mainfrom
litellm_bedrock_openai_no_cachepoint
Sep 17, 2026
Merged

mateo-berri merged 2 commits into
mainfrom
litellm_bedrock_openai_no_cachepoint

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Bedrock global.openai.gpt-6-astra calls carrying cache_control markers fail with 403
  • Bedrock answers "unsupported model or your request did not allow prompt caching"
  • Converse translation sends cachePoint blocks to OpenAI-family models
  • OpenAI models on Bedrock only cache implicitly and reject explicit cachePoint

How it solves it:

  • bedrock_model_accepts_cache_points returns False for any openai. model id
  • Applies whether or not the model is in the cost map
  • cache_control markers are dropped for those models like other non-caching models

User Flow

Before: a developer whose OpenAI-compatible app marks its system prompt with cache_control gets a 403 from every global.openai.gpt-6-astra call, while the same call to global.openai.gpt-5.6-sol works

  1. The proxy admin adds global.openai.gpt-6-astra (bedrock/global.openai.gpt-6-astra) next to the working global.openai.gpt-5.6-sol entry and restarts the proxy
  2. The app sends POST https://litellm-domain/v1/chat/completions with "model": "global.openai.gpt-6-astra" and a system message carrying "cache_control": {"type": "ephemeral"}, the same markers it sends so Claude models on the gateway get prompt caching
  3. The proxy answers 403 with BedrockException - {"message":"You invoked an unsupported model or your request did not allow prompt caching. See the documentation for more information."}
  4. The same request with "model": "global.openai.gpt-5.6-sol" answers 200 with the assistant reply, so astra only works if the app strips the markers for that one model

After: the same astra request answers 200 and the app needs no per-model special case

  1. The proxy admin adds global.openai.gpt-6-astra (bedrock/global.openai.gpt-6-astra) next to the working global.openai.gpt-5.6-sol entry and restarts the proxy
  2. The app sends POST https://litellm-domain/v1/chat/completions with "model": "global.openai.gpt-6-astra" and a system message carrying "cache_control": {"type": "ephemeral"}
  3. The proxy answers 200 with the assistant reply and usage, and the x-litellm-response-cost header carries the call's cost
  4. The same request with "model": "global.openai.gpt-5.6-sol" still answers 200, and Claude models on the gateway still report cache_creation_input_tokens on their cached requests

Relevant issues

Affected release

Linear ticket

Resolves LIT-7861

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Live proxies against real Bedrock (us-east-1), two uvicorn workers each (--num_workers 2), LITELLM_LOCAL_MODEL_COST_MAP=True, no DB. Before runs main at the merge base on :47080, After runs this PR's tip on :53760. Only the proxy commit differs between the two legs

Shared setup:

# proxy_config.yaml
model_list:
  - model_name: global.openai.gpt-6-astra
    litellm_params:
      model: bedrock/global.openai.gpt-6-astra
      aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
      aws_region_name: us-east-1
  - model_name: global.openai.gpt-5.6-sol
    litellm_params:
      model: bedrock/global.openai.gpt-5.6-sol
      aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
      aws_region_name: us-east-1
  - model_name: claude-control
    litellm_params:
      model: bedrock/global.anthropic.claude-fable-5-1
      aws_access_key_id: os.environ/AWS_ACCESS_KEY_ID
      aws_secret_access_key: os.environ/AWS_SECRET_ACCESS_KEY
      aws_region_name: us-east-1

general_settings:
  master_key: sk-1234
LITELLM_LOCAL_MODEL_COST_MAP=True python litellm/proxy/proxy_cli.py --config proxy_config.yaml --port <port> --num_workers 2 --detailed_debug

The Claude regression-control payload (req_claude_cache.json) is a system message of about 11k tokens marked cache_control, so a Claude model has to report cache creation when the gateway still places its cachePoint:

python3 - <<'EOF'
import json
rules = [f"Rule {i}: when the user asks about topic number {i}, answer with the {i}th canonical explanation, keep it under {20 + (i % 7)} words, and cite section {i % 13} of the internal handbook before moving on." for i in range(1, 181)]
system_text = "You are a terse assistant. Follow every rule below.\n" + "\n".join(rules)
json.dump({"model": "claude-control", "max_tokens": 30,
           "messages": [{"role": "system", "content": [{"type": "text", "text": system_text, "cache_control": {"type": "ephemeral"}}]},
                        {"role": "user", "content": "Say hello."}]}, open("req_claude_cache.json", "w"))
EOF

Before (1dd4c13)

cache_control on the system prompt, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "messages": [
    {"role": "system", "content": [{"type": "text", "text": "You are a terse assistant. Answer in one short sentence.", "cache_control": {"type": "ephemeral"}}]},
    {"role": "user", "content": "Say hello."}
  ],
  "max_tokens": 50
}'
  1. Observed:
{"error":{"message":"litellm.PermissionDeniedError: BedrockException - {\"message\":\"You invoked an unsupported model or your request did not allow prompt caching. See the documentation for more information.\"}. Received Model Group=global.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"permission_error","param":null,"code":"403"}}
HTTP 403

tools carrying cache_control, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"global.openai.gpt-6-astra","max_tokens":80,
 "messages":[{"role":"user","content":"What is the weather in Paris? Use the tool."}],
 "tools":[{"type":"function","function":{"name":"get_weather","description":"Get the weather for a city","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}},"cache_control":{"type":"ephemeral"}}]}'
  1. Observed:
{"error":{"message":"litellm.PermissionDeniedError: BedrockException - {\"message\":\"You invoked an unsupported model or your request did not allow prompt caching. See the documentation for more information.\"}. Received Model Group=global.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"permission_error","param":null,"code":"403"}}
HTTP 403

cache_control_injection_points on the system message, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model": "global.openai.gpt-6-astra", "max_tokens": 20, "messages": [{"role": "system", "content": "You are terse. Answer in one word."}, {"role": "user", "content": "Say hello."}], "cache_control_injection_points": [{"location": "message", "role": "system"}]}'
  1. Observed:
{"error":{"message":"litellm.PermissionDeniedError: BedrockException - {\"message\":\"You invoked an unsupported model or your request did not allow prompt caching. See the documentation for more information.\"}. Received Model Group=global.openai.gpt-6-astra\nAvailable Model Group Fallbacks=None","type":"permission_error","param":null,"code":"403"}}
HTTP 403

no cache_control, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "messages": [
    {"role": "system", "content": "You are a terse assistant. Answer in one short sentence."},
    {"role": "user", "content": "Say hello."}
  ],
  "max_tokens": 50
}'
  1. Observed:
{"id":"chatcmpl-dae24dd1-9343-4943-ae72-bc9d35587315","created":1789681742,"model":"global.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello!","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":25,"total_tokens":31,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

reasoning_effort=low with the cost header, gpt-6-astra

  1. Command:
curl -sS -D headers.txt -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "messages": [{"role": "user", "content": "What is 17*23? Answer with just the number."}],
  "max_tokens": 200,
  "reasoning_effort": "low"
}' && grep -i -E '^(HTTP/|x-litellm-response-cost|x-litellm-model-id|llm_provider-x-amzn-bedrock)' headers.txt
  1. Observed:
{"id":"chatcmpl-cc6379be-1675-4316-ace1-ee6183f0977f","created":1789681848,"model":"global.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"391","role":"assistant"}}],"usage":{"completion_tokens":5,"prompt_tokens":19,"total_tokens":24,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":5},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200
HTTP/1.1 200 OK
x-litellm-model-id: e721c6d840c5e85b24f285b5aabf81efc542a6a426b2da5f8b8f8cfbae592d49
x-litellm-response-cost: 0.00044
x-litellm-response-cost-original: 0.00044
x-litellm-response-cost-discount-amount: 0.0
x-litellm-response-cost-margin-amount: 0.0
x-litellm-response-cost-margin-percent: 0.0
x-litellm-response-cost-input: 0.00019
x-litellm-response-cost-output: 0.00025
x-litellm-response-cost-tool-usage: 0.0

cache_control on the system prompt, gpt-5.6-sol (the customer's working model)

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-5.6-sol",
  "messages": [
    {"role": "system", "content": [{"type": "text", "text": "You are a terse assistant. Answer in one short sentence.", "cache_control": {"type": "ephemeral"}}]},
    {"role": "user", "content": "Say hello."}
  ],
  "max_tokens": 50
}'
  1. Observed:
{"id":"chatcmpl-7b8048ee-450e-4448-b338-dd2aefd0f029","created":1789681803,"model":"global.openai.gpt-5.6-sol","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello!","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":25,"total_tokens":31,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

/v1/messages with cache_control on the system prompt, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/messages -H 'x-api-key: sk-1234' -H 'anthropic-version: 2023-06-01' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "max_tokens": 50,
  "system": [{"type": "text", "text": "You are a terse assistant. Answer in one short sentence.", "cache_control": {"type": "ephemeral"}}],
  "messages": [{"role": "user", "content": "Say hello."}]
}'
  1. Observed:
{"id":"chatcmpl-5656a5c5-3def-4bc0-8bad-187824818918","type":"message","role":"assistant","model":"global.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":25,"output_tokens":6},"content":[{"type":"text","text":"Hello!"}],"stop_reason":"end_turn","stop_details":null}
HTTP 200

/v1/responses, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"global.openai.gpt-6-astra","input":"Say hello in one word.","max_output_tokens":30}' | python3 -c 'import sys,json; raw=sys.stdin.read(); body,_,code=raw.rpartition("\nHTTP "); d=json.loads(body); print(json.dumps({"id": d.get("id"), "model": d.get("model"), "status": d.get("status"), "output_text": "".join(c.get("text","") for o in d.get("output",[]) for c in o.get("content",[]) if isinstance(c,dict)), "usage": d.get("usage")})); print("HTTP "+code.strip())'
  1. Observed:
{"id": "resp_Gj1LlWM-HZ31LJDbIBzI5bGq5XPkZY6TpCS4ugESWSFetc5yut3nrwG7ZB3GfXPpI7V3Z1z2djwhrzfmLMyOoXkz9DGGcIFPZ5OL7_Q71S19XbkGVsVHGw4z8aU1FJJR6HOMU7JRMaQLtRZSiQPNnN-dyUefonvnfCNoMIQmZQOymYQXMkALYQi9jYPnyjX04kipEYWqb2ja-hhcsmfwj-Nsa5qJeaY-oE7tEi3WWoZe7ni2bYKfKo7G-vW8ulc-mhS4BQVtWbZKB6h8gWhaiFCeYHKR0TGK7ne_oT37qFrhsN-apt7gVmBzv3BgXLcCgs_6puyGoB3YjjxA6VTjLA1h44bi-Ru8byLCT6i-NrX5MK5GLiCduZWxGf1SS34xi2djpS-WSzN-tGh3uDW4B3yF4ZdmtUb4ta3_-VyAW54N21mR85WnIn8lM94_xTCMlaFtY4lP", "model": "global.openai.gpt-6-astra", "status": "completed", "output_text": "Hello", "usage": {"input_tokens": 12, "input_tokens_details": {"audio_tokens": null, "cached_tokens": 0, "cached_tokens_details": null, "text_tokens": 12, "cache_write_tokens": 0}, "output_tokens": 5, "output_tokens_details": {"audio_tokens": null, "reasoning_tokens": 0, "text_tokens": 5}, "total_tokens": 17, "cost": null}}
HTTP 200

Claude regression control: cache_control on an 11k-token system prompt, claude-control

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:47080/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_claude_cache.json | python3 -c 'import sys,json; raw=sys.stdin.read(); body,_,code=raw.rpartition("\nHTTP "); d=json.loads(body); print(json.dumps({"model": d.get("model"), "usage": d.get("usage"), "content": d["choices"][0]["message"]["content"]})); print("HTTP "+code.strip())'
  1. Observed:
{"model": "claude-control", "usage": {"completion_tokens": 6, "prompt_tokens": 11011, "total_tokens": 11017, "completion_tokens_details": {"reasoning_tokens": 0, "text_tokens": 6}, "prompt_tokens_details": {"cached_tokens": 0, "text_tokens": 11, "cache_write_tokens": 11000, "cache_creation_tokens": 11000, "cache_creation_token_details": {"ephemeral_5m_input_tokens": 11000, "ephemeral_1h_input_tokens": 0}}, "cache_creation_input_tokens": 11000, "cache_read_input_tokens": 0}, "content": "Hello."}
HTTP 200

After (36b471f)

cache_control on the system prompt, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "messages": [
    {"role": "system", "content": [{"type": "text", "text": "You are a terse assistant. Answer in one short sentence.", "cache_control": {"type": "ephemeral"}}]},
    {"role": "user", "content": "Say hello."}
  ],
  "max_tokens": 50
}'
  1. Observed:
{"id":"chatcmpl-af81db35-edd6-4202-81ec-5361172a9ab1","created":1789682083,"model":"global.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello!","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":25,"total_tokens":31,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

tools carrying cache_control, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"global.openai.gpt-6-astra","max_tokens":80,
 "messages":[{"role":"user","content":"What is the weather in Paris? Use the tool."}],
 "tools":[{"type":"function","function":{"name":"get_weather","description":"Get the weather for a city","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}},"cache_control":{"type":"ephemeral"}}]}'
  1. Observed:
{"id":"chatcmpl-6e02137a-b024-4f20-b23a-8d5d3be68394","created":1789682100,"model":"global.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"tool_calls","index":0,"message":{"content":"","role":"assistant","tool_calls":[{"index":0,"function":{"arguments":"{\"city\": \"Paris\"}","name":"get_weather"},"id":"call_e73c727ae1235a5e96caa687e8b27be0","type":"function"}]}}],"usage":{"completion_tokens":18,"prompt_tokens":54,"total_tokens":72,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":18},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":54,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

cache_control_injection_points on the system message, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model": "global.openai.gpt-6-astra", "max_tokens": 20, "messages": [{"role": "system", "content": "You are terse. Answer in one word."}, {"role": "user", "content": "Say hello."}], "cache_control_injection_points": [{"location": "message", "role": "system"}]}'
  1. Observed:
{"id":"chatcmpl-20327fe1-7a3b-47a4-ae8a-6bd06737e95e","created":1789684033,"model":"global.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello.","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":22,"total_tokens":28,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":22,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

no cache_control, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "messages": [
    {"role": "system", "content": "You are a terse assistant. Answer in one short sentence."},
    {"role": "user", "content": "Say hello."}
  ],
  "max_tokens": 50
}'
  1. Observed:
{"id":"chatcmpl-ee791eee-d03b-4561-aa38-732e0a495ad9","created":1789682090,"model":"global.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello!","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":25,"total_tokens":31,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

reasoning_effort=low with the cost header, gpt-6-astra

  1. Command:
curl -sS -D headers.txt -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "messages": [{"role": "user", "content": "What is 17*23? Answer with just the number."}],
  "max_tokens": 200,
  "reasoning_effort": "low"
}' && grep -i -E '^(HTTP/|x-litellm-response-cost|x-litellm-model-id|llm_provider-x-amzn-bedrock)' headers.txt
  1. Observed:
{"id":"chatcmpl-7aaa33aa-3bb1-440e-ae03-987e7ba6c21b","created":1789682109,"model":"global.openai.gpt-6-astra","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"391","role":"assistant"}}],"usage":{"completion_tokens":5,"prompt_tokens":19,"total_tokens":24,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":5},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":19,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200
HTTP/1.1 200 OK
x-litellm-model-id: e721c6d840c5e85b24f285b5aabf81efc542a6a426b2da5f8b8f8cfbae592d49
x-litellm-response-cost: 0.00044
x-litellm-response-cost-original: 0.00044
x-litellm-response-cost-discount-amount: 0.0
x-litellm-response-cost-margin-amount: 0.0
x-litellm-response-cost-margin-percent: 0.0
x-litellm-response-cost-input: 0.00019
x-litellm-response-cost-output: 0.00025
x-litellm-response-cost-tool-usage: 0.0

cache_control on the system prompt, gpt-5.6-sol (the customer's working model)

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-5.6-sol",
  "messages": [
    {"role": "system", "content": [{"type": "text", "text": "You are a terse assistant. Answer in one short sentence.", "cache_control": {"type": "ephemeral"}}]},
    {"role": "user", "content": "Say hello."}
  ],
  "max_tokens": 50
}'
  1. Observed:
{"id":"chatcmpl-ef7d0723-d15c-432f-9bd4-7ed2ccea6330","created":1789682094,"model":"global.openai.gpt-5.6-sol","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello!","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":25,"total_tokens":31,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25,"cache_write_tokens":0,"cache_creation_tokens":0},"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}
HTTP 200

/v1/messages with cache_control on the system prompt, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/messages -H 'x-api-key: sk-1234' -H 'anthropic-version: 2023-06-01' -H 'Content-Type: application/json' -d '{
  "model": "global.openai.gpt-6-astra",
  "max_tokens": 50,
  "system": [{"type": "text", "text": "You are a terse assistant. Answer in one short sentence.", "cache_control": {"type": "ephemeral"}}],
  "messages": [{"role": "user", "content": "Say hello."}]
}'
  1. Observed:
{"id":"chatcmpl-3be52804-9b78-4435-812c-a8a03d66e55d","type":"message","role":"assistant","model":"global.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":25,"output_tokens":6},"content":[{"type":"text","text":"Hello!"}],"stop_reason":"end_turn","stop_details":null}
HTTP 200

/v1/responses, gpt-6-astra

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/responses -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d '{"model":"global.openai.gpt-6-astra","input":"Say hello in one word.","max_output_tokens":30}' | python3 -c 'import sys,json; raw=sys.stdin.read(); body,_,code=raw.rpartition("\nHTTP "); d=json.loads(body); print(json.dumps({"id": d.get("id"), "model": d.get("model"), "status": d.get("status"), "output_text": "".join(c.get("text","") for o in d.get("output",[]) for c in o.get("content",[]) if isinstance(c,dict)), "usage": d.get("usage")})); print("HTTP "+code.strip())'
  1. Observed:
{"id": "resp_h7Ykbf8CfK2Nde-AxEBPpIA2YF4k_FomvI3M6eKJOLg5PsL9xJr1pNQJUI00r1YAdOxS6gJqNWrKtpgmgWpQeq2LPLnJEkXG1RiCFBKoy0LS1ATse8gGRHOhY3LFIInWStCY8VdSP0nUOOyGHdjSnxMqXnjvTZyuLsmDCNR3cVa5_iCI67qfoPMVRQSoxgDgLVyRKi_R802_T523ER8y5l-UG47OGtRyIehMIrAX9kvuazZXvnabGcB33eUog6oTkrvCM4gjzPG_7SVYCKrlxIFg38UyY5WvQAARVj_Jy72hmwixX4GNgzAW2nPOcFyGogbJcY9VZvEiFcnpVPBnleWzKzcotOWxB2lQtGp2sKg_aBl2aLtYogbHMAePNsaygouUmvxoOlZIxqdI_ddHc5jNXUaYz9QQbknojq39xPj1AwHd9T5z1CZTnd-Mmzyb9gzE29nl", "model": "global.openai.gpt-6-astra", "status": "completed", "output_text": "Hello!", "usage": {"input_tokens": 12, "input_tokens_details": {"audio_tokens": null, "cached_tokens": 0, "cached_tokens_details": null, "text_tokens": 12, "cache_write_tokens": 0}, "output_tokens": 6, "output_tokens_details": {"audio_tokens": null, "reasoning_tokens": 0, "text_tokens": 6}, "total_tokens": 18, "cost": null}}
HTTP 200

Claude regression control: cache_control on an 11k-token system prompt, claude-control

  1. Command:
curl -sS -w '\nHTTP %{http_code}\n' http://localhost:53760/v1/chat/completions -H 'Authorization: Bearer sk-1234' -H 'Content-Type: application/json' -d @req_claude_cache.json | python3 -c 'import sys,json; raw=sys.stdin.read(); body,_,code=raw.rpartition("\nHTTP "); d=json.loads(body); print(json.dumps({"model": d.get("model"), "usage": d.get("usage"), "content": d["choices"][0]["message"]["content"]})); print("HTTP "+code.strip())'
  1. Observed:
{"model": "claude-control", "usage": {"completion_tokens": 6, "prompt_tokens": 11011, "total_tokens": 11017, "completion_tokens_details": {"reasoning_tokens": 0, "text_tokens": 6}, "prompt_tokens_details": {"cached_tokens": 11000, "text_tokens": 11, "cache_write_tokens": 0, "cache_creation_tokens": 0}, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 11000}, "content": "Hello."}
HTTP 200

Claude Code on /v1/messages (both legs)

Claude Code v2.1.274 in tmux with ANTHROPIC_BASE_URL pointed at each proxy and ANTHROPIC_MODEL=global.openai.gpt-6-astra, prompted with "Reply with exactly one word: ready". Before (1dd4c13) then After (36b471f):

Before: Claude Code before
After: Claude Code after

A Claude-Code-shaped /v1/messages body answers 200 on v1.101.0 and on this tip

cache_control sits on the system block, the tool, and the user block, the places Claude Code puts it. The v1.101.0 proxy ran with the remote cost map, the default a released proxy uses

curl -sS -w '
HTTP %{http_code}
' http://localhost:<port>/v1/messages -H 'x-api-key: sk-1234' -H 'anthropic-version: 2023-06-01' -H 'Content-Type: application/json' -d '{"model": "global.openai.gpt-6-astra", "max_tokens": 64, "system": [{"type": "text", "text": "You are a terse assistant.", "cache_control": {"type": "ephemeral"}}], "tools": [{"name": "get_weather", "description": "Get the weather for a city", "input_schema": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}, "cache_control": {"type": "ephemeral"}}], "messages": [{"role": "user", "content": [{"type": "text", "text": "Reply with exactly one word: ready", "cache_control": {"type": "ephemeral"}}]}]}'
### v1.101.0-remote-map :48208 /v1/messages Claude-Code-shaped body (cache_control on system, tools, and the user block), no thinking
{"id":"chatcmpl-0d2c2298-78d8-44ac-812a-fdd9249ad8ba","type":"message","role":"assistant","model":"global.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":60,"output_tokens":5},"content":[{"type":"text","text":"ready"}],"stop_reason":"end_turn"}
HTTP 200
### pr-tip-36b471ff24 :53760 /v1/messages Claude-Code-shaped body (cache_control on system, tools, and the user block), no thinking
{"id":"chatcmpl-a932085e-da3e-4317-b3ae-d9e979071b4a","type":"message","role":"assistant","model":"global.openai.gpt-6-astra","stop_sequence":null,"usage":{"input_tokens":60,"output_tokens":5},"content":[{"type":"text","text":"ready"}],"stop_reason":"end_turn","stop_details":null}
HTTP 200
v1.101.0 with model_info.supports_prompt_caching set to false answers 200 for all three shapes
model_list:
  - model_name: astra-workaround
    litellm_params:
      model: bedrock/global.openai.gpt-6-astra
      aws_region_name: us-east-1
    model_info:
      supports_prompt_caching: false
### v1.101.0 (remote cost map) workaround deployment model_info.supports_prompt_caching=false :48208 /v1/chat/completions body=req_cache
{"id":"chatcmpl-47537320-c890-4771-b0d5-54e254040a92","created":1789684886,"model":"astra-workaround","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello!","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":25,"total_tokens":31,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":25 ...<trimmed>
HTTP 200

### v1.101.0 (remote cost map) workaround deployment model_info.supports_prompt_caching=false :48208 /v1/chat/completions body=req_tools_cache
{"id":"chatcmpl-79db9b62-73ff-4e69-ba2f-dccd97e3fdf4","created":1789684888,"model":"astra-workaround","object":"chat.completion","choices":[{"finish_reason":"tool_calls","index":0,"message":{"content":"","role":"assistant","tool_calls":[{"index":0,"function":{"arguments":"{\"city\": \"Paris\"}","name":"get_weather"},"id":"call_94c3e842d499587492eaed82c421a838","type":"function"}]}}],"usage":{"completion_tokens":18,"p ...<trimmed>
HTTP 200

### v1.101.0 (remote cost map) workaround deployment model_info.supports_prompt_caching=false :48208 /v1/chat/completions body=req_injection
{"id":"chatcmpl-75b16f4a-1c51-4380-a687-fe70ec85c1cf","created":1789684890,"model":"astra-workaround","object":"chat.completion","choices":[{"finish_reason":"stop","index":0,"message":{"content":"Hello.","role":"assistant"}}],"usage":{"completion_tokens":6,"prompt_tokens":22,"total_tokens":28,"completion_tokens_details":{"reasoning_tokens":0,"text_tokens":6},"prompt_tokens_details":{"cached_tokens":0,"text_tokens":22 ...<trimmed>
HTTP 200

Observations from the run:

  • Claude Code on /v1/messages answered on both legs
  • cache_control_injection_points reaches the same gate: 403 Before, 200 After
  • That route already drops cache_control for non-Claude models
  • Claude control: cache created Before, read back After
  • /v1/responses unaffected on both legs; it carries no cache markers
  • gpt-5.6-sol answered 200 on both legs: its map entry has no supports_prompt_caching, so it was never sent a cachePoint
  • On v1.101.0, model_info: supports_prompt_caching: false on the deployment turns the 403 into a 200 for all three request shapes, for anyone who cannot upgrade yet

Type

🐛 Bug Fix

Caveats (if any)

Low

  • Detection is by the openai. id segment, since Converse rejects cachePoint for the whole family:
    • Among mapped Converse ids only us.openai.gpt-6-astra and global.openai.gpt-6-astra change; the gpt-oss and gpt-5.6 entries never carried the flag
    • An unmapped or future OpenAI id now drops cachePoint instead of sending it
    • bedrock_mantle/openai.* entries carry the flag but never reach this Converse gate
  • A bare application inference profile ARN that wraps an OpenAI model has no openai. segment, so it still gets cachePoint (unmapped ids keep emitting since fix(bedrock): gate Converse cachePoint emission on model prompt caching support #39210). The documented model: bedrock/<openai id> plus model_id: <ARN> form is covered, because the gate reads model. Not exercised live
  • The reporter never named their client. A Claude-Code-shaped /v1/messages body with cache_control on system, tools, and the user block returns 200 on v1.101.0 and on this tip, so the proof drives /v1/chat/completions, the route that sends cachePoint
  • Cost map keeps supports_prompt_caching: true for these models so implicit cache-read usage is still priced
  • If AWS later adds explicit cachePoint for OpenAI models, this gate must be revisited
  • /v1/messages callers were never affected: that route already drops cache_control for non-Claude models
  • Non-required reds at this tip come from main, not this diff:

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Link to Devin session: https://app.devin.ai/sessions/079b374b62fa4ca0b0b788df204a1ae9
Open in Devin Desktop: https://app.devin.ai/desktop/session/079b374b62fa4ca0b0b788df204a1ae9?variant=devin


Note

Low Risk
Narrow Bedrock Converse gating change with tests; only suppresses explicit cachePoint for OpenAI-family ids and fixes prior 403s without altering Claude caching behavior.

Overview
Fixes 403 errors when Bedrock Converse calls use cache_control / cachePoint with OpenAI-family model IDs (e.g. global.openai.gpt-6-astra). Those models only support implicit caching and reject explicit cachePoint blocks.

bedrock_model_accepts_cache_points now returns False for any model id matching an openai. segment (via new _OPENAI_FAMILY_MODEL_RE), before cost-map lookup—so unmapped regional OpenAI ids are suppressed too. Converse translation stops injecting cachePoint for those models while Claude and other caching-capable models are unchanged.

Tests cover OpenAI-family suppression for tool-config cache injection and the prompt-caching emission matrix.

Reviewed by Cursor Bugbot for commit 36b471f. Bugbot is set up for automated code reviews on this repo. Configure here.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot requested a review from a team September 16, 2026 13:22
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

CLAassistant commented Sep 16, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mateo-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

@codspeed

codspeed Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_bedrock_openai_no_cachepoint (36b471f) with main (356b8d4)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge with no outstanding correctness or repository-rule issues.

Findings

  1. P2 Hardcoded Model Capability ▶

Summary

This PR prevents Bedrock Converse requests from emitting explicit cachePoint blocks for OpenAI-family model identifiers while preserving implicit prompt-caching accounting.

  • Detects OpenAI-family identifiers across standard, regional, and unmapped model forms.
  • Drops cache points from messages, system content, and tool configuration for those models.
  • Adds regression coverage for mapped and unmapped OpenAI-family models while retaining cache points for supported Anthropic models.

Reviews (2) · Last reviewed commit: "Merge commit '1dd4c13815' into litellm_b..."



_ERROR_REQUEST_URL: Final = "https://docs.litellm.ai/docs"
_OPENAI_FAMILY_MODEL_RE: Final = re.compile(r"(^|[./])openai\.")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Hardcoded Model Capability

This regex bypasses model metadata, violating the requirement that model-specific flags use model_prices_and_context_window.json. This requirement must be satisfied before merging

Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Converse rejects cachePoint for every Bedrock OpenAI-family model: a provider API contract, not a per-model capability. The map's supports_prompt_caching stays true for implicit cache-read pricing

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — I conflated two separate capabilities. supports_prompt_caching in the cost map controls cached-token accounting/pricing, including implicit cache reads; it is not a declaration that the Converse wire format accepts an explicit cachePoint. AWS documents OpenAI Bedrock caching as a separate model-family/API contract, so encoding this gate in pricing metadata would incorrectly couple protocol serialization to billing data. The early return in litellm/llms/bedrock/common_utils.py is therefore appropriate, and I withdraw the hardcoded-capability finding.

@codecov

codecov Bot commented Sep 16, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 36b471f. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit fec8231 into main Sep 17, 2026
138 of 144 checks passed
@mateo-berri
mateo-berri deleted the litellm_bedrock_openai_no_cachepoint branch September 17, 2026 22:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants