Skip to content

fix(agent): enable prompt caching for LiteLLM OpenAI-compatible wire (#84506) - #84550

Closed
Enough1122 wants to merge 2 commits into
NousResearch:mainfrom
Enough1122:fix/84506-litellm-prompt-caching
Closed

Enough1122 wants to merge 2 commits into
NousResearch:mainfrom
Enough1122:fix/84506-litellm-prompt-caching

Conversation

@Enough1122

Copy link
Copy Markdown

Fixes #84506

Root cause

anthropic_prompt_cache_policy() granted Anthropic cache_control markers only on the native Anthropic wire (api_mode == "anthropic_messages"). LiteLLM deployments exposing the OpenAI-compatible surface (/v1/chat/completions) matched no branch → fell through to (False, False) → zero cache hits, full prompt re-billed every turn.

Fix

Detect LiteLLM by provider id substring (custom:litellm, litellm) or base-URL host substring (litellm), combined with Claude model detection: grant cache_control on the OpenAI-compatible wire exactly like OpenRouter. Anthropic-wire LiteLLM keeps the existing (True, True) behavior; non-Claude models on LiteLLM stay uncached.

Tests

  • 5 new policy tests: custom:litellm → (True, False); bare litellm → (True, False); bare custom + litellm host → (True, False); LiteLLM on anthropic wire → (True, True); non-Claude on LiteLLM → (False, False)
  • tests/run_agent/test_anthropic_prompt_cache_policy.py: 39/39 pass
  • Related suites (cache-disabled stubs, switch-model pool reload, fallback reasoning override, fallback credential isolation): 30/30 pass
  • 3 pre-existing failures verified unrelated (missing anthropic package in shared venv; fail identically with changes stashed)

@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/anthropic Anthropic native Messages API P3 Low — cosmetic, nice to have sweeper:risk-caching Sweeper risk: may break/degrade prompt caching or cache-key stability (invariant) labels Aug 12, 2026
Enough1122 and others added 2 commits August 13, 2026 15:25
…n unrelated hosts (NousResearch#84506)

The LiteLLM detection granted cache_control when 'litellm' appeared anywhere
in the full base URL, including a path segment on an unrelated host (e.g.
https://api.corp.example.com/v1/litellm-failover). A strict OpenAI-wire host
that rejects unknown keys would then fail with HTTP 400. Switch to
base_url_hostname so only a 'litellm' substring in the actual host grants the
marker, matching how every other host check in this function works
(OpenRouter/MiniMax/Anthropic). Added a regression test for the path-segment
false positive.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Enough1122

Copy link
Copy Markdown
Author

Superseded by #87039 — the combined fix that absorbs both this PR (#84550) and #84982 into one comprehensive solution (per the synthesis documented in #87039's body). Closing this one so maintainers see a single, more complete PR for the same issue. (#84506)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have provider/anthropic Anthropic native Messages API sweeper:risk-caching Sweeper risk: may break/degrade prompt caching or cache-key stability (invariant) type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Prompt caching never engages for LiteLLM on the OpenAI-compatible wire

2 participants