Skip to content

fix(headroom): don't let trailing cache_control make compression a no-op - #42948

Open
tanvir-ux wants to merge 4 commits into
BerriAI:mainfrom
tanvir-ux:fix/headroom-trailing-cache-control-noop
Open

tanvir-ux wants to merge 4 commits into
BerriAI:mainfrom
tanvir-ux:fix/headroom-trailing-cache-control-noop

Conversation

@tanvir-ux

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Claude Code marks the live turn with cache_control every request
  • That trailing marker made Headroom protect 100% of messages
  • Compression never ran; Cost Optimization stayed at 0 with no DEBUG line

How it solves it:

  • Ignore a trailing-turn-only cache_control when picking the cached prefix
  • If an earlier breakpoint exists, that one stays the prefix end
  • Log when nothing is left to compress, like the other early returns

User Flow

Before: a proxy admin enables Headroom for Claude Code traffic and sees zero compression forever

  1. They set guardrail: headroom with default_on: true and a Headroom sidecar at http://headroom:8787
  2. A developer points Claude Code at ANTHROPIC_BASE_URL=https://litellm-domain and chats for a few turns (Claude Code puts cache_control on the system prompt and on the live trailing turn)
  3. They hit GET http://headroom:8787/stats and see /v1/compress still at 0
  4. Admin UI → Cost Optimization → Prompt Compression shows 0 tokens compressed, and LITELLM_LOG=DEBUG has no line explaining the skip

After: the same Claude Code traffic actually reaches Headroom, with mid-history after the stable breakpoint compressible

  1. Same Headroom config and sidecar
  2. Same Claude Code conversation with trailing cache_control
  3. GET http://headroom:8787/stats shows /v1/compress incrementing on those turns
  4. Cost Optimization shows non-zero tokens compressed; if a request is still fully protected, DEBUG logs Headroom: nothing compressible (protected=N/N)

Relevant issues

Fixes #42939

Affected release

regression since the cache-prefix protection landed in #41161 (present in v1.104.0-dev.1)

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/compression/test_compress.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Screenshots / Proof of Fix

Shared setup: same 8-message Claude Code–shaped conversation (system + trailing cache_control, two tool turns in between).

Before (ddc7ee6)

  1. Evaluated get_protected_indices(messages) on that conversation
  2. Observed:
messages=8
protected=[0, 1, 2, 3, 4, 5, 6, 7]
protected_count=8/8
compressible=[]

After (5787c87)

  1. Same call on the same conversation
  2. Observed:
messages=8
protected=[0, 5, 7]
protected_count=3/8
compressible=[1, 2, 3, 4, 6]

Also locally: python -m pytest tests/unit/compression/test_compress.py -q → 15 passed.

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • Trailing-turn markers no longer extend the Anthropic cache-prefix protection; mid-history after the earlier breakpoint can be rewritten. That is what makes Headroom useful for Claude Code, but operators who relied on "trailing CC ⇒ never rewrite anything" get compression again.
  • Full frozen_message_count / session-aware Headroom API wiring from the issue is still follow-up; this PR only stops the silent no-op.

Low

  • Headroom proxy integration tests need litellm[proxy] extras; I covered the policy in tests/unit/compression/test_compress.py and added a Headroom fixture test for CI.

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

Claude Code puts cache_control on the live trailing turn every request.
Treating that marker as the Anthropic cache-prefix end protected the
entire conversation, so Headroom never called /v1/compress (BerriAI#42939).

Ignore a trailing-turn-only breakpoint (or defer to the earlier one when
present), and log when nothing is left to compress.
@tanvir-ux
tanvir-ux requested a review from a team September 24, 2026 12:46
@CLAassistant

CLAassistant commented Sep 24, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@greptile-apps

greptile-apps Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

The PR is not ready to merge until final cache-marked tool results remain protected and the repository's comment requirement is satisfied

Findings

  1. P1 Final tool result loses protection ▶
  2. P2 Unnecessary source comments ▶

Summary

The PR excludes a trailing cache-control marker when selecting the protected cached prefix, allowing eligible history to reach Headroom, and logs when every message remains protected

  • Adds unit and guardrail tests for trailing-marker conversations
  • The shared compressor still needs to protect a final cache-marked tool result

Reviews (1) · Last reviewed commit: "fix(headroom): don't let trailing cache_..."

Comment on lines +237 to 243
prefix_end: Final = (
breakpoints[-2]
if breakpoints[-1] == len(messages) - 1 and len(breakpoints) >= 2
else -1
if breakpoints[-1] == len(messages) - 1
else breakpoints[-1]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Final tool result loses protection If a cache-marked tool result is the final message, this change excludes its breakpoint, and role protection does not cover tool results. With a tight compression budget, compress() can replace the result with a stub while retaining its top-level marker. The changed content can turn a cache read into a write

Knowledge Base Used: Protect cache-marked history during compression

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Comment on lines +810 to +812
# Match the other early returns: leave a breadcrumb when compression
# is a no-op so Cost Optimization / DEBUG is not left with a silent
# "0 tokens compressed" and no explanation (#42939).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Unnecessary source comments This comment repeats the debug log rather than explaining complex logic. The new assertion comments do the same. The repository permits source comments only for necessary complex logic, tool input, or justified TODOs and FIXMEs; this requirement must be satisfied before merging

Context Used: AGENTS.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@codspeed

codspeed Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing tanvir-ux:fix/headroom-trailing-cache-control-noop (b588587) with main (88fd153)

Open in CodSpeed

@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 87.50000% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/compression/compress.py 85.71% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@lets-order-some-fries lets-order-some-fries left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reproduced #42939 independently at main 571ada0b0f before this PR and re-ran the same script against this head (5787c877f5). On main, the Claude Code shape (50 rows, cache_control on rows 0 and 49):

"litellm_get_protected_indices_len": 50, "protected": 50, "compressible": 0,
"compress_calls": 0, "headroom_log_lines": [], "applied_guardrails_in_request_data": null

At this head, same input:

"litellm_get_protected_indices_len": 3, "protected": 26, "compressible": 24,
"compress_calls": 1, "applied_guardrails_in_request_data": ["headroom-compression"]

The control case (system breakpoint only) is unchanged in both.

Root cause: yes for that shape — _cached_prefix_indices (litellm/compression/compress.py:217) stops treating a breakpoint on the final row as the prefix end.

On the cross-turn side I simulated two consecutive turns through headroom._protected_indices on this head and on untouched main:

head  turn N  : n=13 prefix=(0,)  protected=8/13
head  turn N+1: n=15 prefix=(0,)  protected=9/15
      rows sent verbatim at turn N but compressible at turn N+1: [11]
main  turn N/N+1: protected 13/13 and 15/15, rows flipped: []

Row 11 is inside the span turn N's trailing breakpoint asked the provider to cache, so turn N's cache write cannot be read at turn N+1. That is the failure mode #39519 reported and #41161 landed to stop. Does the trailing-marker case also need the previous turn's cached span held intact, or is the per-turn compression expected to outweigh the lost read?

Two narrower shapes from the same run: with two adjacent trailing breakpoints, prefix_end = breakpoints[-2] = len-2 and protection is still 13/13, compressible [] — does that shape need handling too? And when the trailing marker is the only breakpoint, _cached_prefix_indices returns (), so from turn 2 on the previously cached prefix gets no protection (compressible: [1, 3, 5, 7, 9]) — intended?

get_protected_indices is also called from guardrail_hooks/typesafe/typesafe.py:141 and compression/compress.py:470; I read both but did not test them.

Tests at this head: tests/unit/compression/test_compress.py 15 passed, tests/test_litellm/proxy/guardrails/guardrail_hooks/test_headroom.py 106 passed.

The failing lint check is not a lint rule — it is the diff-scoped Check ruff format step (ruff 0.15.3): Would reformat: litellm/compression/compress.py, exit 123. The only hunk is compress.py:232-234, which the formatter wants on one line (105 chars, ruff.toml line-length 120):

    breakpoints: Final = [index for index, msg in enumerate(messages) if _message_has_cache_control(msg)]

With that single change, ruff 0.15.3 format --check reports 1 file already formatted.

…x_indices

Unblocks the Check ruff format lint step on BerriAI#42948.
@tanvir-ux

Copy link
Copy Markdown
Author

Thanks @lets-order-some-fries — really solid repro, and the numbers match what I saw.

On the cross-turn point: yes, dropping the trailing marker can leave a row that turn N wrote into cache unprotected at turn N+1. That's intentional for the Claude Code shape. That trailing cache_control is a write marker for the next call, not a claim that the whole history is a stable prefix. Leaving Headroom permanently off (#42939) felt worse than accepting that those trailing-only writes may miss on the next turn. Mid-history / system breakpoints from #41161 still pin the real prefix.

Two adjacent trailing breakpoints: I haven't seen Claude Code emit that (usually system + one live trailing turn). With the current logic breakpoints[-2] would still protect almost everything. Happy to walk back past a run of trailing markers in a follow-up if that's a shape we care about — didn't want to widen this PR past #42939.

Sole trailing breakpoint returning (): also intentional. It doesn't establish a stable prefix yet, so _cached_prefix_indices has nothing to hold; role protection (last user / last assistant / system) still applies through get_protected_indices.

Pushed the ruff one-liner on compress.py so the format check should go green.

@lets-order-some-fries

Copy link
Copy Markdown

Your ruff change did land — Check ruff format passes at c2b7c987. The lint job just fails one step later now, at Check type-discipline budget (step 20, scripts/type_discipline_gate.py --base), which is why it still shows red.

I ran the repo's own gate on compress.py at your head and at the gate base (ddc7ee6838): 103 findings vs 102. The single new one is on the line the format fix produced:

LIT002 mutable list comprehension: ... Build it in one shot and freeze it --
a tuple/frozenset wrapping a generator (`tuple(f(x) for x in xs)`) ...

Changing that one line to a tuple clears it:

breakpoints: Final = tuple(index for index, msg in enumerate(messages) if _message_has_cache_control(msg))

With that applied the gate reports 102 — identical to base, so the delta is zero. 110 chars, under the 120 limit. Everything downstream (breakpoints[-1], breakpoints[-2], len(...), if not breakpoints) works unchanged on a tuple; I re-ran the four shapes and got the same results: no breakpoints (), mid-history (0, 1), trailing-only (), mid+trailing (0,).

Thanks for the detailed answers on the cross-turn question — the write-marker-for-the-next-call framing makes sense, and the two narrower shapes being deliberate follow-ups is a clearer boundary than I had.

LIT002 flags the listcomp; tuple(gen) clears the budget delta vs base.
@tanvir-ux

Copy link
Copy Markdown
Author

Thanks @lets-order-some-fries — good catch on the LIT002 delta.

Pushed the change to a tuple:

breakpoints: Final = tuple(index for index, msg in enumerate(messages) if _message_has_cache_control(msg))

tests/unit/compression/test_compress.py still 15/15. That should zero the type-discipline budget vs base.

@lets-order-some-fries

Copy link
Copy Markdown

My tuple suggestion was wrong and I'm sorry — it cleared the type-discipline gate and breached the next one. lint at e67ca822 now fails at step 23, Check basedpyright budget:

FAIL: basedpyright errors exceed the per-rule limit:
  reportGeneralTypeIssues: total 103 over limit 101 (this change added 2)

Those 2 are mine. I ran basedpyright 1.40.1 on compress.py both ways:

form reportGeneralTypeIssues
your original list 0
tuple(...) 2

Both are the same diagnostic:

L239: Index -1 is out of range for type tuple[()]
L240: Index -1 is out of range for type tuple[()]

basedpyright doesn't narrow a variadic tuple[int, ...] to non-empty after if not breakpoints: return (), so breakpoints[-1] and [-2] look unsafe to it. A list[int] narrows fine. Annotating Final[tuple[int, ...]] explicitly does not help — still 2, and the line goes to 127 chars.

What does pass both gates is keeping your list and using the escape hatch the LIT002 message itself offers (suppress: # mutable-ok: <reason>). The checker requires it on the offending line, so the comprehension has to wrap — which is stable, because the collapsed form is 141 chars and ruff will not join it:

    breakpoints: Final = [  # mutable-ok: local to this function, never returned or stored
        index for index, msg in enumerate(messages) if _message_has_cache_control(msg)
    ]

Measured on your head with that applied:

  • scripts/check_type_discipline.py litellm/compression/compress.py → 102 findings, identical to the gate base ddc7ee6838, so the LIT delta is zero
  • basedpyright reportGeneralTypeIssues → 0, back to the base count
  • longest line 90 chars
  • _cached_prefix_indices unchanged on all four shapes: no breakpoints (), mid-history (0, 1), trailing-only (), mid+trailing (0,)

I should have checked the later gates before suggesting the first change — the earlier failures were masking them, and I only looked at the step that was red at the time. Apologies for the extra round trip.

One thing I have not verified: rust-lint is also red at this head, and mcp-integration is failing repo-wide at the moment for an unrelated reason (I filed #43192 about that — it fails on ~19% of main commits since #42904, so it is very likely not yours).

@tanvir-ux

Copy link
Copy Markdown
Author

Thanks @lets-order-some-fries — yeah, the tuple cleared LIT002 then blew basedpyright on breakpoints[-1] / [-2]. Switched back to the list with # mutable-ok as you suggested (b588587). lint is green on this head.

Noted on rust-lint / mcp-integration and #43192 — leaving those alone unless a maintainer says they're on this PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Headroom guardrail compresses nothing when the client sets a trailing cache_control breakpoint (protected == 100% of messages)

3 participants