Repository navigation
Conversation
Claude Code puts cache_control on the live trailing turn every request. Treating that marker as the Anthropic cache-prefix end protected the entire conversation, so Headroom never called /v1/compress (BerriAI#42939). Ignore a trailing-turn-only breakpoint (or defer to the earlier one when present), and log when nothing is left to compress.
|
| prefix_end: Final = ( | ||
| breakpoints[-2] | ||
| if breakpoints[-1] == len(messages) - 1 and len(breakpoints) >= 2 | ||
| else -1 | ||
| if breakpoints[-1] == len(messages) - 1 | ||
| else breakpoints[-1] | ||
| ) |
There was a problem hiding this comment.
Final tool result loses protection If a cache-marked tool result is the final message, this change excludes its breakpoint, and role protection does not cover tool results. With a tight compression budget,
compress() can replace the result with a stub while retaining its top-level marker. The changed content can turn a cache read into a write
Knowledge Base Used: Protect cache-marked history during compression
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| # Match the other early returns: leave a breadcrumb when compression | ||
| # is a no-op so Cost Optimization / DEBUG is not left with a silent | ||
| # "0 tokens compressed" and no explanation (#42939). |
There was a problem hiding this comment.
Unnecessary source comments This comment repeats the debug log rather than explaining complex logic. The new assertion comments do the same. The repository permits source comments only for necessary complex logic, tool input, or justified TODOs and FIXMEs; this requirement must be satisfied before merging
Context Used: AGENTS.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
lets-order-some-fries
left a comment
There was a problem hiding this comment.
I reproduced #42939 independently at main 571ada0b0f before this PR and re-ran the same script against this head (5787c877f5). On main, the Claude Code shape (50 rows, cache_control on rows 0 and 49):
"litellm_get_protected_indices_len": 50, "protected": 50, "compressible": 0,
"compress_calls": 0, "headroom_log_lines": [], "applied_guardrails_in_request_data": null
At this head, same input:
"litellm_get_protected_indices_len": 3, "protected": 26, "compressible": 24,
"compress_calls": 1, "applied_guardrails_in_request_data": ["headroom-compression"]
The control case (system breakpoint only) is unchanged in both.
Root cause: yes for that shape — _cached_prefix_indices (litellm/compression/compress.py:217) stops treating a breakpoint on the final row as the prefix end.
On the cross-turn side I simulated two consecutive turns through headroom._protected_indices on this head and on untouched main:
head turn N : n=13 prefix=(0,) protected=8/13
head turn N+1: n=15 prefix=(0,) protected=9/15
rows sent verbatim at turn N but compressible at turn N+1: [11]
main turn N/N+1: protected 13/13 and 15/15, rows flipped: []
Row 11 is inside the span turn N's trailing breakpoint asked the provider to cache, so turn N's cache write cannot be read at turn N+1. That is the failure mode #39519 reported and #41161 landed to stop. Does the trailing-marker case also need the previous turn's cached span held intact, or is the per-turn compression expected to outweigh the lost read?
Two narrower shapes from the same run: with two adjacent trailing breakpoints, prefix_end = breakpoints[-2] = len-2 and protection is still 13/13, compressible [] — does that shape need handling too? And when the trailing marker is the only breakpoint, _cached_prefix_indices returns (), so from turn 2 on the previously cached prefix gets no protection (compressible: [1, 3, 5, 7, 9]) — intended?
get_protected_indices is also called from guardrail_hooks/typesafe/typesafe.py:141 and compression/compress.py:470; I read both but did not test them.
Tests at this head: tests/unit/compression/test_compress.py 15 passed, tests/test_litellm/proxy/guardrails/guardrail_hooks/test_headroom.py 106 passed.
The failing lint check is not a lint rule — it is the diff-scoped Check ruff format step (ruff 0.15.3): Would reformat: litellm/compression/compress.py, exit 123. The only hunk is compress.py:232-234, which the formatter wants on one line (105 chars, ruff.toml line-length 120):
breakpoints: Final = [index for index, msg in enumerate(messages) if _message_has_cache_control(msg)]With that single change, ruff 0.15.3 format --check reports 1 file already formatted.
…x_indices Unblocks the Check ruff format lint step on BerriAI#42948.
|
Thanks @lets-order-some-fries — really solid repro, and the numbers match what I saw. On the cross-turn point: yes, dropping the trailing marker can leave a row that turn N wrote into cache unprotected at turn N+1. That's intentional for the Claude Code shape. That trailing Two adjacent trailing breakpoints: I haven't seen Claude Code emit that (usually system + one live trailing turn). With the current logic Sole trailing breakpoint returning Pushed the ruff one-liner on |
|
Your ruff change did land — I ran the repo's own gate on Changing that one line to a tuple clears it: breakpoints: Final = tuple(index for index, msg in enumerate(messages) if _message_has_cache_control(msg))With that applied the gate reports 102 — identical to base, so the delta is zero. 110 chars, under the 120 limit. Everything downstream ( Thanks for the detailed answers on the cross-turn question — the write-marker-for-the-next-call framing makes sense, and the two narrower shapes being deliberate follow-ups is a clearer boundary than I had. |
LIT002 flags the listcomp; tuple(gen) clears the budget delta vs base.
|
Thanks @lets-order-some-fries — good catch on the LIT002 delta. Pushed the change to a tuple: breakpoints: Final = tuple(index for index, msg in enumerate(messages) if _message_has_cache_control(msg))
|
|
My tuple suggestion was wrong and I'm sorry — it cleared the type-discipline gate and breached the next one. Those 2 are mine. I ran basedpyright 1.40.1 on
Both are the same diagnostic: basedpyright doesn't narrow a variadic What does pass both gates is keeping your list and using the escape hatch the LIT002 message itself offers ( breakpoints: Final = [ # mutable-ok: local to this function, never returned or stored
index for index, msg in enumerate(messages) if _message_has_cache_control(msg)
]Measured on your head with that applied:
I should have checked the later gates before suggesting the first change — the earlier failures were masking them, and I only looked at the step that was red at the time. Apologies for the extra round trip. One thing I have not verified: |
|
Thanks @lets-order-some-fries — yeah, the tuple cleared LIT002 then blew basedpyright on Noted on |
TLDR
Problem this solves:
cache_controlevery requestHow it solves it:
cache_controlwhen picking the cached prefixUser Flow
Before: a proxy admin enables Headroom for Claude Code traffic and sees zero compression forever
guardrail: headroomwithdefault_on: trueand a Headroom sidecar athttp://headroom:8787ANTHROPIC_BASE_URL=https://litellm-domainand chats for a few turns (Claude Code putscache_controlon the system prompt and on the live trailing turn)GET http://headroom:8787/statsand see/v1/compressstill at 0LITELLM_LOG=DEBUGhas no line explaining the skipAfter: the same Claude Code traffic actually reaches Headroom, with mid-history after the stable breakpoint compressible
cache_controlGET http://headroom:8787/statsshows/v1/compressincrementing on those turnsHeadroom: nothing compressible (protected=N/N)Relevant issues
Fixes #42939
Affected release
regression since the cache-prefix protection landed in #41161 (present in v1.104.0-dev.1)
Pre-Submission checklist
uv run pytest tests/unit/compression/test_compress.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Screenshots / Proof of Fix
Shared setup: same 8-message Claude Code–shaped conversation (system + trailing
cache_control, two tool turns in between).Before (ddc7ee6)
get_protected_indices(messages)on that conversationAfter (5787c87)
Also locally:
python -m pytest tests/unit/compression/test_compress.py -q→ 15 passed.Type
🐛 Bug Fix
Caveats (if any)
Medium
frozen_message_count/ session-aware Headroom API wiring from the issue is still follow-up; this PR only stops the silent no-op.Low
litellm[proxy]extras; I covered the policy intests/unit/compression/test_compress.pyand added a Headroom fixture test for CI.Final Attestation