Repository navigation
fix(caching): stand default cache points down when extra_body hides a direct client mark - #43342
Conversation
… direct client mark On native /v1/messages the extra_body envelope is dropped, so a client tool mark or root cache_control reaches Anthropic even when extra_body overrides it. The stand-down check only counted the envelope-merged view and injected two default marks on top of the client's. (cherry picked from commit 01ef7c4)
|
| + AnthropicCacheControlHook.count_external_cache_breakpoints_on_messages_route( | ||
| tools, cache_control, request_kwargs | ||
| ) |
There was a problem hiding this comment.
Chat caching can disappear When a
/chat/completions request has a marked tool but extra_body.tools replaces it with an unmarked tool, this check sees the direct mark and skips the automatic cache breakpoints. The chat path then sends the replacement tool, so Anthropic receives no breakpoint and the request loses automatic prompt caching.
… for the default stand-down Chat completions merge extra_body over the request, so a direct tool mark that extra_body replaces never reaches the provider there. Only /v1/messages, where the native transforms drop the envelope, needs to count marks on both sides. (cherry picked from commit a9c5fa7)
TLDR
Problem this solves:
/v1/messageswhenextra_bodyoverrides a client-marked tool or rootcache_controlextra_body, so the client's mark still reaches AnthropicHow it solves it:
/v1/messages, the automatic defaults stand down if the client marked anything, with or withoutextra_bodyapplied/chat/completionskeeps counting only the marks that reach the wire, sinceextra_bodyis merged therecache_control_injection_pointsare unchangedUser Flow
Before: a client that caches its own tools gets extra cache breakpoints it never asked for
cache_control, and anextra_body.toolscopy without itAfter: the proxy leaves a client that marked its own cache breakpoints alone
extra_bodyRelevant issues
Backport of #43341
Affected release
Regression on rc/1.103.0 since #43331, not in v1.103.0-rc.1
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)The cherry-pick applied cleanly. The new test fails on rc/1.103.0 for both hidden-mark cases (2 marks injected where 0 are expected) and passes here. Its third case, with no client mark anywhere, still injects on both sides, so the test cannot pass by never injecting. A second test covers
/chat/completions: a marked tool thatextra_bodyreplaces with an unmarked one still gets the defaults, in both the injected points and the router cache-affinity prediction. It fails if the/v1/messagescount is used on chat.test_anthropic_cache_control_hook.pypasses excepttest_non_anthropic_providers_never_injected[gemini-2.0-flash-gemini], which also fails on the rc tip because the remote cost map on main dropped that model. It passes withLITELLM_LOCAL_MODEL_COST_MAP=TrueType
🐛 Bug Fix
Caveats (if any)
Medium
Low
/v1/messagesfrom chat, so it keeps the chat count/v1/messages, affinity predicts defaults the request will not carry