Repository navigation
fix: backport five regression fixes to rc/1.103.0 - #43331
Conversation
|
| try: | ||
| loop: Final = asyncio.get_running_loop() | ||
| except RuntimeError: | ||
| return | ||
| loop.create_task(self._push_in_memory_increments_to_redis()) |
There was a problem hiding this comment.
If a synchronous Router.update_settings() call replaces a usage-based selector while it has Redis increments queued, retire() cancels its sync task and returns because there is no running event loop. Those increments are never flushed, so later routing decisions can use understated RPM usage.
| rebuild_routing_groups = True | ||
| elif var == "routing_strategy_args": | ||
| routing_args_updated = True | ||
| routing_args_updated = value != self.routing_strategy_args |
There was a problem hiding this comment.
Mutated routing arguments stay stale
If a caller changes the dictionary previously passed as routing_strategy_args and passes that same dictionary to update_settings(), this comparison checks the dictionary against itself. It skips rebuilding the selector, so the router reports the new arguments while routing still uses the old values.
|
|
TLDR
Problem this solves:
How it solves it:
cache_control_injection_pointsapply beside client marksredacted_thinkingblocks no longer breakprompt_cachingpinning/v1/mcp/toolsreturns camelCaseinputSchemaandoutputSchemaagainLITELLM_RUST=1,/v1/messagesstays on PythonUser Flow
Before: a customer on v1.103.0-rc.1 hits one of five breaks that worked on an earlier release
cache_control_injection_pointson an Anthropic deployment sends POST http://localhost:4000/v1/chat/completions with a client-marked system block, andcached_tokensstays at the system block only on every turnprompt_cachingreplays aredacted_thinkingblock to POST http://localhost:4000/v1/messages, and each turn lands on a different deployment, paying a fresh cache writeinputSchema, which is missing because the route now returnsinput_schemaLITELLM_RUST=1sends a compaction edit to POST http://localhost:4000/v1/messages and gets 400context_management: Extra inputs are not permittedAfter: the same five flows behave the way they did before the regression
inputSchemaandoutputSchemaAffected release
Regressions in v1.103.0-rc.1: #42352 and #42517 since v1.103.0-rc.1, #42069 and #42784 since v1.102.0, #41956 since v1.95.0
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Every touched test file passes on this branch (1873 tests). Each pick's own tests were also run against rc/1.103.0 with its source change removed, and they fail there: #41956 48 failures, #42069 2, #42352 1, #42517 189, #42784 2
Adaptations from the main versions:
patch()in its test carries atest-quality-okreason, since rc/1.103.0 is already over its TQ008 budgetType
🐛 Bug Fix
Caveats (if any)
Low