Repository navigation
fix(bedrock): never emit Converse cachePoint for OpenAI-family models - #41419
Conversation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
|
|
|
|
||
|
|
||
| _ERROR_REQUEST_URL: Final = "https://docs.litellm.ai/docs" | ||
| _OPENAI_FAMILY_MODEL_RE: Final = re.compile(r"(^|[./])openai\.") |
There was a problem hiding this comment.
This regex bypasses model metadata, violating the requirement that model-specific flags use model_prices_and_context_window.json. This requirement must be satisfied before merging
Rule Used: What: Do not hardcode model-specific flags in the codebase. Instead, put them in model_prices_and_context_window.json and then read them in via get_model_info Why: Prevents need for users to upgrade litellm each time a new model supports this featu... (source)
There was a problem hiding this comment.
Converse rejects cachePoint for every Bedrock OpenAI-family model: a provider API contract, not a per-model capability. The map's supports_prompt_caching stays true for implicit cache-read pricing
There was a problem hiding this comment.
You're right — I conflated two separate capabilities. supports_prompt_caching in the cost map controls cached-token accounting/pricing, including implicit cache reads; it is not a declaration that the Converse wire format accepts an explicit cachePoint. AWS documents OpenAI Bedrock caching as a separate model-family/API contract, so encoding this gate in pricing metadata would incorrectly couple protocol serialization to billing data. The early return in litellm/llms/bedrock/common_utils.py is therefore appropriate, and I withdraw the hardcoded-capability finding.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 36b471f. Configure here.
TLDR
Problem this solves:
global.openai.gpt-6-astracalls carryingcache_controlmarkers fail with 403cachePointblocks to OpenAI-family modelsHow it solves it:
bedrock_model_accepts_cache_pointsreturns False for anyopenai.model idcache_controlmarkers are dropped for those models like other non-caching modelsUser Flow
Before: a developer whose OpenAI-compatible app marks its system prompt with
cache_controlgets a 403 from everyglobal.openai.gpt-6-astracall, while the same call toglobal.openai.gpt-5.6-solworksglobal.openai.gpt-6-astra(bedrock/global.openai.gpt-6-astra) next to the workingglobal.openai.gpt-5.6-solentry and restarts the proxy"model": "global.openai.gpt-6-astra"and a system message carrying"cache_control": {"type": "ephemeral"}, the same markers it sends so Claude models on the gateway get prompt cachingBedrockException - {"message":"You invoked an unsupported model or your request did not allow prompt caching. See the documentation for more information."}"model": "global.openai.gpt-5.6-sol"answers 200 with the assistant reply, so astra only works if the app strips the markers for that one modelAfter: the same astra request answers 200 and the app needs no per-model special case
global.openai.gpt-6-astra(bedrock/global.openai.gpt-6-astra) next to the workingglobal.openai.gpt-5.6-solentry and restarts the proxy"model": "global.openai.gpt-6-astra"and a system message carrying"cache_control": {"type": "ephemeral"}x-litellm-response-costheader carries the call's cost"model": "global.openai.gpt-5.6-sol"still answers 200, and Claude models on the gateway still reportcache_creation_input_tokenson their cached requestsRelevant issues
Affected release
Linear ticket
Resolves LIT-7861
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Live proxies against real Bedrock (us-east-1), two uvicorn workers each (
--num_workers 2),LITELLM_LOCAL_MODEL_COST_MAP=True, no DB. Before runsmainat the merge base on :47080, After runs this PR's tip on :53760. Only the proxy commit differs between the two legsShared setup:
The Claude regression-control payload (
req_claude_cache.json) is a system message of about 11k tokens markedcache_control, so a Claude model has to report cache creation when the gateway still places itscachePoint:Before (1dd4c13)
cache_control on the system prompt, gpt-6-astra
tools carrying cache_control, gpt-6-astra
cache_control_injection_points on the system message, gpt-6-astra
no cache_control, gpt-6-astra
reasoning_effort=low with the cost header, gpt-6-astra
cache_control on the system prompt, gpt-5.6-sol (the customer's working model)
/v1/messages with cache_control on the system prompt, gpt-6-astra
/v1/responses, gpt-6-astra
Claude regression control: cache_control on an 11k-token system prompt, claude-control
After (36b471f)
cache_control on the system prompt, gpt-6-astra
tools carrying cache_control, gpt-6-astra
cache_control_injection_points on the system message, gpt-6-astra
no cache_control, gpt-6-astra
reasoning_effort=low with the cost header, gpt-6-astra
cache_control on the system prompt, gpt-5.6-sol (the customer's working model)
/v1/messages with cache_control on the system prompt, gpt-6-astra
/v1/responses, gpt-6-astra
Claude regression control: cache_control on an 11k-token system prompt, claude-control
Claude Code on /v1/messages (both legs)
Claude Code v2.1.274 in tmux with
ANTHROPIC_BASE_URLpointed at each proxy andANTHROPIC_MODEL=global.openai.gpt-6-astra, prompted with "Reply with exactly one word: ready". Before (1dd4c13) then After (36b471f):Before:

After:
A Claude-Code-shaped /v1/messages body answers 200 on v1.101.0 and on this tip
cache_controlsits on the system block, the tool, and the user block, the places Claude Code puts it. The v1.101.0 proxy ran with the remote cost map, the default a released proxy usesv1.101.0 with model_info.supports_prompt_caching set to false answers 200 for all three shapes
Observations from the run:
cache_control_injection_pointsreaches the same gate: 403 Before, 200 Aftercache_controlfor non-Claude modelssupports_prompt_caching, so it was never sent a cachePointmodel_info: supports_prompt_caching: falseon the deployment turns the 403 into a 200 for all three request shapes, for anyone who cannot upgrade yetType
🐛 Bug Fix
Caveats (if any)
Low
openai.id segment, since Converse rejects cachePoint for the whole family:us.openai.gpt-6-astraandglobal.openai.gpt-6-astrachange; the gpt-oss and gpt-5.6 entries never carried the flagbedrock_mantle/openai.*entries carry the flag but never reach this Converse gateopenai.segment, so it still gets cachePoint (unmapped ids keep emitting since fix(bedrock): gate Converse cachePoint emission on model prompt caching support #39210). The documentedmodel: bedrock/<openai id>plusmodel_id: <ARN>form is covered, because the gate readsmodel. Not exercised livecache_controlon system, tools, and the user block returns 200 on v1.101.0 and on this tip, so the proof drives /v1/chat/completions, the route that sends cachePointsupports_prompt_caching: truefor these models so implicit cache-read usage is still pricedcache_controlfor non-Claude modelsmain, not this diff:llm_translation_testingandlitellm_utils_testingfail only onfireworks_ai/minimax-m3vision tests, since price sync chore(prices): sync Together AI prices: 6 models, 6 deprecated [sync failed: Google Gemini] #41570 setsupports_vision: falsetest_bedrock_gpt_oss.pyincludedocr_testingfails on the Azure AI OCR URL test, which flips between green and red across unrelated pipelinesosv-scanflags soupsieve 2.8.4 inuv.lockon every PR; build(deps): bump soupsieve from 2.8.4 to 2.9.2 #41679 bumps itFinal Attestation
Link to Devin session: https://app.devin.ai/sessions/079b374b62fa4ca0b0b788df204a1ae9
Open in Devin Desktop: https://app.devin.ai/desktop/session/079b374b62fa4ca0b0b788df204a1ae9?variant=devin
Note
Low Risk
Narrow Bedrock Converse gating change with tests; only suppresses explicit cachePoint for OpenAI-family ids and fixes prior 403s without altering Claude caching behavior.
Overview
Fixes 403 errors when Bedrock Converse calls use
cache_control/cachePointwith OpenAI-family model IDs (e.g.global.openai.gpt-6-astra). Those models only support implicit caching and reject explicitcachePointblocks.bedrock_model_accepts_cache_pointsnow returnsFalsefor any model id matching anopenai.segment (via new_OPENAI_FAMILY_MODEL_RE), before cost-map lookup—so unmapped regional OpenAI ids are suppressed too. Converse translation stops injectingcachePointfor those models while Claude and other caching-capable models are unchanged.Tests cover OpenAI-family suppression for tool-config cache injection and the prompt-caching emission matrix.
Reviewed by Cursor Bugbot for commit 36b471f. Bugbot is set up for automated code reviews on this repo. Configure here.