Skip to content

fix(caching): stand default cache points down when extra_body hides a direct client mark - #43342

Merged
yuneng-berri merged 2 commits into
rc/1.103.0from
litellm_rc_1_103_0_fix_cache_default_stand_down
Sep 26, 2026
Merged

yuneng-berri merged 2 commits into
rc/1.103.0from
litellm_rc_1_103_0_fix_cache_default_stand_down

Conversation

@yuneng-berri

@yuneng-berri yuneng-berri commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Automatic prompt caching adds 2 cache marks on top of the client's own
  • Happens on /v1/messages when extra_body overrides a client-marked tool or root cache_control
  • The native Anthropic path drops extra_body, so the client's mark still reaches Anthropic

How it solves it:

  • On /v1/messages, the automatic defaults stand down if the client marked anything, with or without extra_body applied
  • /chat/completions keeps counting only the marks that reach the wire, since extra_body is merged there
  • Configured cache_control_injection_points are unchanged

User Flow

Before: a client that caches its own tools gets extra cache breakpoints it never asked for

  1. The proxy admin turns on automatic prompt caching for an Anthropic model
  2. A client sends POST http://localhost:4000/v1/messages with a tool carrying cache_control, and an extra_body.tools copy without it
  3. Anthropic receives the client's tool mark plus 2 more on the system prompt and the last turn, and bills extra cache writes

After: the proxy leaves a client that marked its own cache breakpoints alone

  1. The proxy admin keeps the same setting
  2. The client sends the same POST http://localhost:4000/v1/messages
  3. Anthropic receives only the client's own tool mark, the same as when there is no extra_body

Relevant issues

Backport of #43341

Affected release

Regression on rc/1.103.0 since #43331, not in v1.103.0-rc.1

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

The cherry-pick applied cleanly. The new test fails on rc/1.103.0 for both hidden-mark cases (2 marks injected where 0 are expected) and passes here. Its third case, with no client mark anywhere, still injects on both sides, so the test cannot pass by never injecting. A second test covers /chat/completions: a marked tool that extra_body replaces with an unmarked one still gets the defaults, in both the injected points and the router cache-affinity prediction. It fails if the /v1/messages count is used on chat. test_anthropic_cache_control_hook.py passes except test_non_anthropic_providers_never_injected[gemini-2.0-flash-gemini], which also fails on the rc tip because the remote cost map on main dropped that model. It passes with LITELLM_LOCAL_MODEL_COST_MAP=True

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Live proxy proof against Anthropic is not captured yet

Low

  • Router cache affinity cannot tell /v1/messages from chat, so it keeps the chat count
    • For this rare shape on /v1/messages, affinity predicts defaults the request will not carry
    • Worst case is that request missing its sticky deployment and paying a fresh cache write

… direct client mark

On native /v1/messages the extra_body envelope is dropped, so a client tool mark or root cache_control reaches Anthropic even when extra_body overrides it. The stand-down check only counted the envelope-merged view and injected two default marks on top of the client's.

(cherry picked from commit 01ef7c4)
@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Fixes cache control logic for Anthropic API integration.

The PR should not merge until chat completions retains automatic caching when extra_body removes a direct mark.

Findings

  1. P1 Chat caching can disappear ▶

Summary

The PR changes automatic cache-mark detection to retain direct client marks when extra_body hides them, and adds /v1/messages regression cases.

  • The shared default-injection decision also serves /chat/completions, where the new census can suppress caching after the direct mark is replaced.

Reviews (1) · Last reviewed commit: "fix(caching): stand default cache points..."

Comment on lines +709 to +711
+ AnthropicCacheControlHook.count_external_cache_breakpoints_on_messages_route(
tools, cache_control, request_kwargs
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Chat caching can disappear When a /chat/completions request has a marked tool but extra_body.tools replaces it with an unmarked tool, this check sees the direct mark and skips the automatic cache breakpoints. The chat path then sends the replacement tool, so Anthropic receives no breakpoint and the request loses automatic prompt caching.

… for the default stand-down

Chat completions merge extra_body over the request, so a direct tool mark that extra_body replaces never reaches the provider there. Only /v1/messages, where the native transforms drop the envelope, needs to count marks on both sides.

(cherry picked from commit a9c5fa7)
@yuneng-berri
yuneng-berri merged commit 10e77ce into rc/1.103.0 Sep 26, 2026
9 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_rc_1_103_0_fix_cache_default_stand_down branch September 26, 2026 23:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant