Skip to content

fix: backport five regression fixes to rc/1.103.0 - #43331

Merged
yuneng-berri merged 5 commits into
rc/1.103.0from
litellm_backport_regression_fixes_rc_1_103_0
Sep 26, 2026
Merged

yuneng-berri merged 5 commits into
rc/1.103.0from
litellm_backport_regression_fixes_rc_1_103_0

Conversation

@yuneng-berri

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Five regressions present in v1.103.0-rc.1 would ship into the next stable
  • Fixes already merged on main but were missing from rc/1.103.0

How it solves it:

User Flow

Before: a customer on v1.103.0-rc.1 hits one of five breaks that worked on an earlier release

  1. An operator with cache_control_injection_points on an Anthropic deployment sends POST http://localhost:4000/v1/chat/completions with a client-marked system block, and cached_tokens stays at the system block only on every turn
  2. A Claude Code user on a group with prompt_caching replays a redacted_thinking block to POST http://localhost:4000/v1/messages, and each turn lands on a different deployment, paying a fresh cache write
  3. A script calls GET http://localhost:4000/v1/mcp/tools and reads inputSchema, which is missing because the route now returns input_schema
  4. A Claude Code user on a proxy with LITELLM_RUST=1 sends a compaction edit to POST http://localhost:4000/v1/messages and gets 400 context_management: Extra inputs are not permitted
  5. An admin with DB-stored config and Slack alerting watches the proxy's background task count grow by two every 30 seconds

After: the same five flows behave the way they did before the regression

  1. The configured tail checkpoint lands next to the client's mark, so later turns read the conversation from cache
  2. Every replayed turn stays on the deployment that served the first one
  3. GET http://localhost:4000/v1/mcp/tools returns inputSchema and outputSchema
  4. The compaction request returns 200
  5. The task count stays flat across reloads

Affected release

Regressions in v1.103.0-rc.1: #42352 and #42517 since v1.103.0-rc.1, #42069 and #42784 since v1.102.0, #41956 since v1.95.0

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Every touched test file passes on this branch (1873 tests). Each pick's own tests were also run against rc/1.103.0 with its source change removed, and they fail there: #41956 48 failures, #42069 2, #42352 1, #42517 189, #42784 2

Adaptations from the main versions:

Type

🐛 Bug Fix

Caveats (if any)

Low

mateo-berri and others added 5 commits September 26, 2026 11:39
…_caching keeps pinning (#42069)

(cherry picked from commit 0e7cf51)
…th is ready (#42517)

rc/1.103.0 has no token counter or tokenizer routes in the catalog, so only the Messages rule moves to PYTHON_ONLY

(cherry picked from commit 85ed18e)
@yuneng-berri
yuneng-berri merged commit 6397b36 into rc/1.103.0 Sep 26, 2026
9 of 10 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_backport_regression_fixes_rc_1_103_0 branch September 26, 2026 18:47
@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[Medium risk] Backports five regression fixes across integrations and routing.

The PR is not ready to merge because router updates can lose queued RPM increments or leave routing arguments stale.

Findings

  1. P1 Queued increments are lost ▶
  2. P1 Mutated routing arguments stay stale ▶

Summary

This PR backports fixes for cache-control injection, replayed redacted-thinking token counting, MCP tool field aliases, Messages dispatch, and background-task growth. The router retirement and settings changes need correction before merging.

Reviews (1) · Last reviewed commit: "fix(proxy): stop leaking periodic tasks ..."

Comment on lines +51 to +55
try:
loop: Final = asyncio.get_running_loop()
except RuntimeError:
return
loop.create_task(self._push_in_memory_increments_to_redis())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Queued increments are lost

If a synchronous Router.update_settings() call replaces a usage-based selector while it has Redis increments queued, retire() cancels its sync task and returns because there is no running event loop. Those increments are never flushed, so later routing decisions can use understated RPM usage.

Comment thread litellm/router.py
rebuild_routing_groups = True
elif var == "routing_strategy_args":
routing_args_updated = True
routing_args_updated = value != self.routing_strategy_args

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Mutated routing arguments stay stale

If a caller changes the dictionary previously passed as routing_strategy_args and passes that same dictionary to update_settings(), this comparison checks the dictionary against itself. It skips rebuilding the selector, so the router reports the new arguments while routing still uses the old values.

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mateo-berri
❌ devin-ai-integration[bot]
You have signed the CLA already but the status is still pending? Let us recheck it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants