Skip to content

perf(router): fetch cooldown state and usage counters in one Redis round trip - #43320

Merged
yassin-berriai merged 11 commits into
mainfrom
litellm_router_single_redis_round_trip
Sep 29, 2026
Merged

yassin-berriai merged 11 commits into
mainfrom
litellm_router_single_redis_round_trip

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Every usage-based-routing-v2 request paid two Redis round trips before the provider call
  • The cooldown filter and the tpm/rpm selector each own an MGET

How it solves it:

  • RoutingReadBatch fetches cooldown keys and tpm/rpm counters in one MGET
  • The strategy reuses those counters instead of reading Redis again
  • DualCache.async_batch_get_cache_shared keeps each cache's memory tier and failure handling

User Flow

Before: a request on a usage-based-routing-v2 group spends two Redis round trips on routing before the provider is called

  1. They send POST https://litellm-domain/v1/chat/completions (or /v1/messages, /v1/responses) for a model group with two deployments
  2. Redis sees MGET deployment:*:cooldown, then a second MGET <id>:<model>:tpm:<minute>, <id>:<model>:rpm:<minute>, then the response-cache GET
  3. The provider call starts only after both routing reads returned

After: the same request spends one Redis round trip on routing

  1. They send the same POST https://litellm-domain/v1/chat/completions
  2. Redis sees one MGET carrying the cooldown keys and the tpm/rpm counter keys together, then the response-cache GET
  3. The provider call starts one Redis round trip earlier; the same deployment is chosen and the response is identical

Relevant issues

Follow-up to #40841 (post-call I/O consolidation), whose body names the router cache reads as the remaining pre-call round trips. Fleet numbers for the perf track: #40539 and https://berriai.grafana.net/d/litellm-perf-1k-rps?from=1789028100000&to=1789033500000

Linear ticket

Resolves LIT-8693

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Design notes

Which path fires the usage MGET

Traced on a local proxy with redis-cli monitor and a call-site tracer on RedisCache:

  • simple-shuffle with no rpm/tpm on the deployments: one cooldown MGET, zero usage reads. Unchanged by this PR
  • usage-based-routing-v2 (router-wide, per routing group, or per-key routing_strategy override through _get_routing_context): cooldown MGET from CooldownCache.async_get_active_cooldowns plus the counter MGET from LowestTPMLoggingHandler_v2.async_get_available_deployments. This is the two-read shape in the prod traces and the path this PR collapses
  • Deployments with rpm/tpm set under simple-shuffle do not read counters at routing time; they only write them (INCRBYFLOAT global_router:* after the call). The model-group limit read (get_model_group_usage) only fires when model_group_rpm/tpm limits are configured

Root cause in one sentence: the cooldown filter and the usage-based selector each own their own MGET because they live in different objects (CooldownCache over its own DualCache, the strategy over the router cache), and the response-cache GET is issued serially after routing although it does not depend on it.

Option picked: (b) a per-request RoutingReadBatch

Router.async_get_available_deployment creates a RoutingReadBatch right after _get_routing_context when the resolved strategy is usage-based-routing-v2, and passes it down:

strategy, selector = self._get_routing_context(model, request_kwargs)
batch = RoutingReadBatch.for_strategy(strategy, selector)          # None for every other strategy
healthy = await self.async_get_healthy_deployments(..., routing_read_batch=batch)
    -> batch.async_get_cooldown_deployments(router, healthy_deployments)   # ONE MGET: cooldown keys + tpm/rpm keys
deployment = await self._select_deployment_async(..., prefetched_usage=batch.prefetched_usage)
    -> selector.async_get_available_deployments(..., prefetched_usage=...)  # no second read when it covers the keys

Why not (a) inside _async_get_cooldown_deployments: the cooldown helper has no view of the strategy, and async_get_healthy_deployments has several callers (fallbacks, health probes, _select_deployment_async callers outside the main path) that must keep the old single read. The batch object is created in exactly one place and everything else keeps its existing signature with an optional parameter, so every other caller is byte-for-byte the old path.

The two key sets live in two different DualCache instances (cooldown store vs router cache) that share one RedisCache. DualCache.async_batch_get_cache_shared lets each cache consult its own in-memory tier and redis_batch_cache_expiry reservation independently, merges only the keys that still need Redis into one RedisCache.async_batch_get_cache, then hands each cache its slice for the normal backfill into memory. If the two caches do not share a Redis client it falls back to two independent reads.

Semantics preserved:

  • Same deployment for the same cache contents: the selector still receives [tpm..., rpm...] values in its own key order and runs the unchanged _common_checks_available_deployment
  • Cooldown filtering: active_cooldowns_from_results is the old tail of async_get_active_cooldowns, factored out so both paths interpret results identically (expiry check, stale-key delete)
  • Redis failure: a raised MGET error propagates from async_get_healthy_deployments exactly like the old cooldown MGET did, ending in the same RateLimitError: No deployments available; a circuit-breaker-open read serves memory-only for both caches, like before
  • Minute rollover between the prefetch and selection: PrefetchedUsage.covers() fails and the strategy issues its own read, so it can never use stale keys
  • The prefetch is keyed on the deployments before cooldown filtering, a superset of what selection asks for, so covers() holds after filtering

Response-cache GET: left serial, not made concurrent

The cache lookup happens in litellm.utils.client (_async_get_cache) before Router routing starts for the call itself, in a different layer: routing runs inside the routed acompletion, so the two are not siblings that could be gathered without restructuring the client wrapper or moving cache lookup into the router. The key is independent of the chosen deployment (it hashes the request), so the reorder is possible in principle, but it touches the cache handler for every call type and is out of scope here; noted as a follow-up on LIT-8693.

/v1/responses sync get_cache: left in place

The prod trace's get_cache <- _sync_get_cache comes from aresponses running the sync responses() wrapper in a worker thread (_worker <- _bootstrap_inner in the call-site trace below), so it is not a blocking call on the event loop. An earlier revision of this PR skipped it for aresponses, and tests/unit/test_utils.py::test_wrapper_async_replays_cached_converted_responses_stream_as_stream failed: for native Responses API models the sync read is the one whose key matches what the write stores, so removing it turns every cached /v1/responses stream into a provider call. The async aresponses read and the write are keyed differently (the live trace below shows the async GET on one key and the SET on another). That key mismatch is the actual defect, it is outside routing, and it is recorded as a follow-up on LIT-8693. This PR does not touch litellm/utils.py.

Alerting file touched by the type gate

DualCache.async_batch_get_cache used to return an untyped list (its element type was Unknown to basedpyright); the shared read gives it a typed list[object | None], and the type-check gate then surfaced six latent diagnostics in SlackAlerting.send_daily_reports, which compared and sorted those elements as numbers. The values are narrowed there with isinstance(val, (int, float)) before use; non-numeric or missing entries become None, exactly what the all_none and placeholder logic already handled. No behavior change on numeric values; tests/unit/integrations/SlackAlerting/test_slack_alerting.py and tests/logging_callback_tests/test_alerting.py -k daily pass (11 tests, the Redis one against a local Redis)

Out of scope, recorded on LIT-8693

  • redis async_get_cache <- _get_routed_model <- async_pre_call_hook (parallel request limiter) is a separate read on a separate path; not touched, per the ticket's instruction not to widen into parallel_request_limiter_v3 without a proven duplicate
  • /v1/responses: the async aresponses cache read and the cache write use different keys, so the effective read for native Responses models is the sync one on the worker thread; aligning the keys would remove one GET per request
  • Making the response-cache GET concurrent with routing (see above)

Screenshots / Proof of Fix

Admin UI, head fedb248

Logs page of the proxy started from this branch (PYTHONPATH=<checkout>, usage-based-routing-v2, Redis 7, real Anthropic claude-sonnet-4-6). The row is request chatcmpl-da0eae5d-057d-4acf-bb79-0024bd021144; the redis-cli monitor trace for that request contains exactly one routing MGET, carrying the cooldown key and the tpm/rpm keys together:

$ grep -c '"MGET"' monitor.log
1
"MGET" "deployment:cd5dd740...:cooldown" "cd5dd740...:anthropic/claude-sonnet-4-6:tpm:09-01" "cd5dd740...:anthropic/claude-sonnet-4-6:rpm:09-01"

Logs page showing the request served by the single-MGET routing path

Local proxy, real Redis 7 on 127.0.0.1:6379, real Anthropic claude-sonnet-4-6 via a key from the environment, model group claude-group with two deployments (claude-dep-a, claude-dep-b), plus claude-rpm with two deployments carrying rpm/tpm limits. Both arms run the same config_v2.yaml (routing_strategy: usage-based-routing-v2, cache: true with cache_params.type: redis, optional_pre_call_checks: [prompt_caching]) and the same fixture bytes, started with PYTHONPATH=<checkout>, redis-cli monitor capturing to a file, and each request wrapped between SET __marker:<arm>:<case>:start|end so the per-request command list is exact. Each request sleeps 11 s first so the 10 s redis_batch_cache_expiry throttle cannot serve it from memory.

Shared setup:

$ cat config_v2.yaml
model_list:
  - model_name: claude-group
    litellm_params: {model: anthropic/claude-sonnet-4-6, api_key: os.environ/ANTHROPIC_API_KEY}
    model_info: {id: claude-dep-a}
  - model_name: claude-group
    litellm_params: {model: anthropic/claude-sonnet-4-6, api_key: os.environ/ANTHROPIC_API_KEY}
    model_info: {id: claude-dep-b}
  - model_name: claude-rpm
    litellm_params: {model: anthropic/claude-sonnet-4-6, api_key: os.environ/ANTHROPIC_API_KEY, rpm: 1000, tpm: 1000000}
    model_info: {id: claude-rpm-a}
  - model_name: claude-rpm
    litellm_params: {model: anthropic/claude-sonnet-4-6, api_key: os.environ/ANTHROPIC_API_KEY, rpm: 1000, tpm: 1000000}
    model_info: {id: claude-rpm-b}
litellm_settings:
  cache: true
  cache_params: {type: redis, host: 127.0.0.1, port: 6379}
router_settings:
  routing_strategy: usage-based-routing-v2
  redis_host: 127.0.0.1
  redis_port: 6379
  optional_pre_call_checks: ["prompt_caching"]
general_settings:
  master_key: sk-perf-1234

$ cat chat.json
{"model":"claude-group","messages":[{"role":"user","content":"Reply with the single word: pong"}],"max_tokens":5}
$ cat messages.json
{"model":"claude-group","max_tokens":5,"messages":[{"role":"user","content":"Reply with the single word: pong"}]}
(the *_stream.json fixtures add "stream":true, responses.json uses "input" instead of "messages", *_rpm.json use "model":"claude-rpm")

$ PYTHONPATH=$PWD python litellm/proxy/proxy_cli.py --config config_v2.yaml --port 4000
$ redis-cli monitor > monitor_<arm>.log &
$ for each case: redis-cli set __marker:<arm>:<case>:start 1; curl -s -w '%{http_code}' -H 'Authorization: Bearer sk-perf-1234' -H 'Content-Type: application/json' http://127.0.0.1:4000<endpoint> -d @<fixture>; redis-cli set __marker:<arm>:<case>:end 1

The lists below are the Redis commands between the two markers, in order, until the first post-call write (INCRBYFLOAT/SET). Everything after that is the same on both arms: response-cache SET, spend MGET/INCRBYFLOAT/EXPIRE, tpm counter increments, and the key/user/model token EVALSHA/INCRBYFLOAT.

Before (14f4c34, main at the merge base)

POST /v1/chat/completions

  1. curl ... /v1/chat/completions -d @chat.json -> 200, body "model":"claude-group", content "pong"
  2. Redis commands before the provider call:
 1. MGET 6 keys: deployment:claude-dep-a:cooldown, deployment:claude-dep-b:cooldown, deployment:claude-rpm-a:cooldown, deployment:claude-rpm-b:cooldown, ...
 2. MGET 4 keys: claude-dep-a:anthropic/claude-sonnet-4-6:tpm:16-49, claude-dep-b:anthropic/claude-sonnet-4-6:tpm:16-49, claude-dep-a:anthropic/claude-sonnet-4-6:rpm:16-49, claude-dep-b:anthropic/claude-sonnet-4-6:rpm:16-49
 3. GET 37d6dd61fc3cbefe42412229ad50af8549a9bf44fcbe531012e2cdccc4293850        <- response cache (miss)
 4. INCRBYFLOAT claude-dep-b:anthropic/claude-sonnet-4-6:rpm:16-49              <- chosen deployment: claude-dep-b
  1. Same shape for "stream":true: 200, SSE chunks with "model":"claude-group"; commands 1 and 2 are the two MGETs (4 counter keys), then the response-cache GET

POST /v1/messages

  1. curl ... /v1/messages -d @messages.json -> 200, body "model":"claude-group", content[0].text = "pong"
  2. Redis commands before the provider call:
 1. MGET 6 keys: deployment:claude-dep-a:cooldown, deployment:claude-dep-b:cooldown, deployment:claude-rpm-a:cooldown, deployment:claude-rpm-b:cooldown, ...
 2. MGET 3 keys: claude-dep-a:anthropic/claude-sonnet-4-6:tpm:16-50, claude-dep-b:anthropic/claude-sonnet-4-6:tpm:16-50, claude-dep-b:anthropic/claude-sonnet-4-6:rpm:16-50
 3. GET 1cf2e64507a2993ac5958fd8c4611646909a888d65e0f72f34369afa254f8e00
 4. INCRBYFLOAT claude-dep-b:anthropic/claude-sonnet-4-6:rpm:16-50

(the second MGET carries fewer keys when the in-memory tier already holds some counters; the round trip still happens)
3. "stream":true: 200, event: message_start with "model":"claude-group"; two MGETs (6 + 2 keys) then the GET

POST /v1/responses

  1. curl ... /v1/responses -d @responses.json -> 200, "id":"resp_..."
  2. Redis commands before the provider call:
 1. MGET 6 keys: deployment:claude-dep-a:cooldown, deployment:claude-dep-b:cooldown, ...
 2. MGET 2 keys: claude-dep-a:anthropic/claude-sonnet-4-6:tpm:16-50, claude-dep-b:anthropic/claude-sonnet-4-6:tpm:16-50
 3. GET 6bf0b05bf6d423a419e196bcf0d3957c921c1e402dafba59d3b08c2eddfe147b        <- async response-cache read (aresponses wrapper)
 4. GET 69fa301f44491579448a2214f00c048cbc0936b6c5c65cf4920c06ee7b192828        <- sync response-cache read from the worker thread (_sync_get_cache)
 5. GET 1a626d0bc3100565acf99ebe599fd44df7a6562b7951e7311fc336e1cf590f59        <- prompt-caching pre-call check key
  1. Call-site tracer over the whole run: [REDIS get_cache SYNC] :: _sync_get_cache <- _worker <- _bootstrap_inner <- _bootstrap x3 (one per non-stream /v1/responses request, on a thread), async_get_available_deployments <- _select_deployment_async x18, async_get_active_cooldowns <- _async_get_cooldown_deployments x18

rpm-limited group (claude-rpm), chat and /v1/messages

  1. 200 on both; two MGETs (6 cooldown keys, then 4 or 2 counter keys for claude-rpm-a/claude-rpm-b) before the response-cache GET

After (branch tip, first measured at 078b26d)

POST /v1/chat/completions

  1. curl ... /v1/chat/completions -d @chat.json -> 200, body "model":"claude-group", content "pong"
  2. Redis commands before the provider call:
 1. MGET 10 keys: deployment:claude-dep-a:cooldown, deployment:claude-dep-b:cooldown, deployment:claude-rpm-a:cooldown, deployment:claude-rpm-b:cooldown, ..., claude-dep-a:anthropic/claude-sonnet-4-6:tpm:17-10, claude-dep-b:...:tpm:17-10, claude-dep-a:...:rpm:17-10, claude-dep-b:...:rpm:17-10
 2. GET 37d6dd61fc3cbefe42412229ad50af8549a9bf44fcbe531012e2cdccc4293850        <- same response-cache key as before (same fixture bytes)
 3. INCRBYFLOAT claude-dep-b:anthropic/claude-sonnet-4-6:rpm:17-10              <- chosen deployment: claude-dep-b, same class (anthropic/claude-sonnet-4-6)
  1. "stream":true: 200, SSE "model":"claude-group"; one MGET (8 keys), then the GET

POST /v1/messages

  1. curl ... /v1/messages -d @messages.json -> 200, "model":"claude-group", content[0].text = "pong"
  2. Redis commands before the provider call:
 1. MGET 10 keys: deployment:claude-dep-a:cooldown, ..., claude-dep-a:anthropic/claude-sonnet-4-6:tpm:17-11, claude-dep-b:...:tpm:17-11, claude-dep-a:...:rpm:17-11, claude-dep-b:...:rpm:17-11
 2. GET 1cf2e64507a2993ac5958fd8c4611646909a888d65e0f72f34369afa254f8e00
 3. INCRBYFLOAT claude-dep-b:anthropic/claude-sonnet-4-6:rpm:17-11
  1. "stream":true: 200, event: message_start with "model":"claude-group"; one MGET (9 keys), then the GET

POST /v1/responses

Re-captured at 8d55b19 (the tip that keeps the worker-thread read; later commits only rename a helper, narrow the alerting values and type/de-mutate the two DualCache helpers, no routing or Redis-traffic change), same fixture, redis-cli flushdb before the run:

  1. curl ... /v1/responses -d @responses.json -> 200, "id":"resp_...", "model":"claude-group"
  2. Redis commands before the provider call:
 1. MGET 11 keys: deployment:claude-dep-a:cooldown, ..., claude-dep-a:anthropic/claude-sonnet-4-6:tpm:17-39, claude-dep-b:...:tpm:17-39, claude-dep-a:...:rpm:17-39, claude-dep-b:...:rpm:17-39
 2. GET 6bf0b05bf6d423a419e196bcf0d3957c921c1e402dafba59d3b08c2eddfe147b        <- async response-cache read (aresponses wrapper)
 3. GET 69fa301f44491579448a2214f00c048cbc0936b6c5c65cf4920c06ee7b192828        <- sync response-cache read from the worker thread, unchanged
 4. GET 1a626d0bc3100565acf99ebe599fd44df7a6562b7951e7311fc336e1cf590f59        <- prompt-caching pre-call check key

The SET after the call writes 6bf0... (the bridge's chat-completion key), never 69fa...; the second request in the same run served the response from cache with only the routing MGET reaching Redis
3. "stream":true: 200, "model":"claude-sonnet-4-6" in the SSE events; one MGET, then GET fa8c... (async), GET 5521... (sync, worker thread), GET 5814... (bridge chat key, the one that is written and hit on the second request)
4. Call-site tracer over the original after run: async_get_cooldown_deployments <- async_get_healthy_deployments x18, async_get_available_deployments <- _select_deployment_async x0

rpm-limited group (claude-rpm), chat and /v1/messages

  1. 200 on both; one MGET (10 keys: 6 cooldown + claude-rpm-a/claude-rpm-b tpm and rpm) before the response-cache GET

simple-shuffle control (config_shuffle.yaml: same file without routing_strategy), before and after

Identical on both arms for every case: 200, then MGET 6 keys: deployment:*:cooldown, then the response-cache GET, and no counter read (the claude-rpm group only writes INCRBYFLOAT global_router:*:tpm after the call). This PR does not create a RoutingReadBatch for simple-shuffle.

Local latency A/B

Same box, same config_v2.yaml, 50 concurrent closed-loop clients for 60 s against the mock-group model (mock_response, so no provider network), one run per arm, back to back:

$ PYTHONPATH=<checkout> python litellm/proxy/proxy_cli.py --config config_v2.yaml --port 4000
$ python load.py http://127.0.0.1:4000/v1/chat/completions chat_mock.json 50 60
before (14f4c34c, litellm from base_wt):  {"requests": 5954, "rps": 98.4, "codes": {"200": 5954}, "p50_ms": 470.9, "p95_ms": 847.0, "p99_ms": 934.3, "mean_ms": 506.5}
after  (078b26dd, litellm from checkout): {"requests": 5947, "rps": 98.6, "codes": {"200": 5947}, "p50_ms": 467.9, "p95_ms": 833.7, "p99_ms": 983.4, "mean_ms": 505.8}

Regime: a single 8-vCPU box running the proxy, Redis and the load generator together, CPU-bound on the proxy at ~98 rps, with the fixture repeated so it is a response-cache hit after the first request (cmdstat_get did not move on either arm: the hit is served from memory). This run measures nothing about routing and is reported only because the ticket asked for the A/B; a second run with "cache": {"no-cache": true} so every request routes is below.

Second run, same rig, fixture with "cache": {"no-cache": true} so every request routes (and the counter MGET fires when the 10 s redis_batch_cache_expiry throttle allows a Redis read):

$ python load.py http://127.0.0.1:4000/v1/chat/completions chat_mock_nocache.json 50 60
before (14f4c34c): {"requests": 5050, "rps": 83.6, "codes": {"200": 5050}, "p50_ms": 565.0, "p95_ms": 963.7, "p99_ms": 1087.2, "mean_ms": 595.9}
after  (078b26dd): {"requests": 5014, "rps": 82.9, "codes": {"200": 5014}, "p50_ms": 561.8, "p95_ms": 996.0, "p99_ms": 1196.5, "mean_ms": 600.4}

No measurable difference at this regime: Redis is on loopback (usec_per_call about 3 us for MGET on both arms), the proxy is CPU-bound at ~83 rps on the shared box, and the in-memory tier plus the 10 s batch throttle mean only about one MGET per second reaches Redis on either arm. The saving this PR buys is one network round trip on the requests that do reach Redis (1.7 to 4.3 ms per MGET in the production traces that motivated the ticket), which a loopback box cannot show. Do not read these numbers as the production effect either way.

Production traces (tip 275197b deployed)

Chart 0.0.0-branch-litellm-router-single-redis-round-275197b built from this branch (https://github.com/BerriAI/litellm-ops/actions/runs/36264994770) and pinned on the production cluster via https://github.com/BerriAI/litellm-ops/pull/189: Argo Synced/Healthy at 19:56 UTC, migrations clean, gateway 2/2 with 0 restarts, real-model smoke 200 on chat, /v1/responses, /v1/messages and key/list. The DB router_settings on that cluster resolve to usage-based-routing-v2

Fresh Tempo traces right after the old pods drained, one per surface, from the same tagged run (otelverify1790452780, 19:59:40 UTC):

chat-stream      df7919e390606f12a648d220525d9903  9 spans
  redis async_batch_get_cache  3.0 ms  caller=async_batch_get_cache_shared <- async_get_cooldown_deployments
  redis async_get_cache        1.3 ms  caller=async_get_cache <- async_get_cache   (response cache)
responses        b662fe8a239004a9ce649b815ac9cea2  7 spans
  redis async_batch_get_cache  3.2 ms  caller=async_batch_get_cache_shared <- async_get_cooldown_deployments
  redis async_get_cache        1.3 ms  (response cache)
  redis get_cache              1.5 ms  caller=get_cache <- _sync_get_cache        (native Responses cache read, kept on purpose)
messages-stream  43f409b447b2df0f4b7640fd690a3804  7 spans
  redis async_batch_get_cache  3.0 ms  caller=async_batch_get_cache_shared <- async_get_cooldown_deployments

Before this build the same three requests carried two routing MGETs each (async_get_active_cooldowns at 4.1 ms plus async_get_available_deployments at 1.6 ms, trace 418e25ae328841906888e2fbee8959c1 on the previous build). A TraceQL search for span.litellm.service.caller=~"async_get_active_cooldowns.*|async_get_available_deployments.*" since the new pods took traffic returns zero spans, name=~".*<-.*" also returns zero, and every post-response Redis/Postgres write is still a linked root trace (5, 4 and 2 detached spans respectively), so the #43237 phase-based detach shipped with this branch is intact

Perf-track context

This change is one of the pre-call items behind the 1k rps perf track (#40539, Grafana https://berriai.grafana.net/d/litellm-perf-1k-rps?from=1789028100000&to=1789033500000). Those fleet numbers were measured on the integration branch as a whole and do not isolate this PR; the numbers above are the only measurements of this change on its own.

Tests and mutation check

$ uv run pytest tests/unit/router_utils/test_routing_read_batch.py tests/unit/caching/test_dual_cache.py tests/unit/router_strategy/test_lowest_tpm_rpm.py tests/unit/router_utils/test_cooldown_cache.py -q
81 passed

test_routing_read_batch.py drives the real Router with a MagicMock(spec=RedisCache) that records every async_batch_get_cache call and asserts exactly one Redis round trip per async_get_available_deployment on usage-based-routing-v2, still one on simple-shuffle, plus parity: same deployment chosen as the unbatched selector for fixed counter values, cooled deployments still excluded, Redis raising still ends in RateLimitError: No deployments available.

Mutants (each reverted with cp backups, restore verified with cmp -s):

  • RoutingReadBatch.for_strategy returns None: round-trip test sees 2 reads, 2 failed
  • PrefetchedUsage.covers forced False: 1 failed
  • shared read grouped per cache instead of per Redis client: 1 failed
  • per-cache prepare/backfill exceptions in the shared read left unhandled: 1 failed
  • cooldown results ignored in the batch: 1 failed

Live re-check at a71e6ff

This head adds the merge of main 7f95b5f into 8432eae. A read-count harness, run from this checkout with PYTHONPATH=<checkout>, one warm request then one request per case, redis_reads_processed from INFO stats minus the three marker commands

$ bash run_arm2.sh <arm> <checkout>  # redis-cli flushall, proxy on :4000 with config_v2.yaml, real Redis 7, real Postgres, Anthropic claude-sonnet-4-6
key/generate 200
warm 200
warm redis_reads_processed=15
chat_master 200
chat_master redis_reads_processed=8
chat_vk 200
chat_vk redis_reads_processed=17
messages_vk 200
messages_vk redis_reads_processed=19

Every response body was pong from claude-group with usage populated, and the proxy log has no ERROR or Traceback lines

Type

🚄 Infrastructure

Caveats (if any)

Low

  • The counters are now read at the cooldown step, before the blocked-deployment, pre-call-check callback, token-count and tag filters run, instead of after them. Selection therefore sees values that are older by the duration of those in-process filters (the only I/O among them is the optional pre-call-check cache read). Read-then-choose was already racy against concurrent requests in the old order (the counters are minute buckets incremented after the call), and the near-call RPM check in LowestTPMLoggingHandler_v2.async_pre_call_check still runs on the chosen deployment as before; TPM had no near-call check on either version
  • Only usage-based-routing-v2 is batched; other counter-reading strategies keep their own read
  • The response-cache GET stays serial to routing; follow-up noted on LIT-8693

CI Status

Head a71e6ff = 8432eae + merge of origin/main 7f95b5f. The only failing check is misc / Run tests, four tests in tests/unit/interactions/test_openapi_compliance.py that read a remote OpenAPI spec; the same job fails the same way on main at 7f95b5f

Link to Devin session: https://app.devin.ai/sessions/f3b64ab771264ecf9b474add1ce663de
Open in Devin Desktop: https://app.devin.ai/desktop/session/f3b64ab771264ecf9b474add1ce663de?variant=devin
Requested by: @yassin-berriai


Note

Medium Risk
Changes pre-call routing and shared Redis read failure semantics for usage-based-routing-v2; behavior is heavily tested for parity but touches core deployment selection on every routed request.

Overview
Collapses two Redis MGETs into one for usage-based-routing-v2 by batching cooldown keys and tpm/rpm counter keys before deployment selection.

Adds DualCache.async_batch_get_cache_shared so separate DualCache instances (cooldown store vs router cache) still use their own in-memory tier, throttling, and backfill while sharing a single Redis MGET. RoutingReadBatch runs that shared read during healthy-deployment resolution and passes counters to LowestTPMLoggingHandler_v2 via PrefetchedUsage, so selection skips a second cache read when keys are already covered. simple-shuffle and other strategies keep the existing single cooldown read.

async_batch_get_cache is refactored into prepare/apply helpers; CooldownCache.active_cooldowns_from_results is extracted so batched cooldown results parse the same way as before. Slack daily reports narrow batch cache values to numeric types for typing only.

Reviewed by Cursor Bugbot for commit a71e6ff. Bugbot is set up for automated code reviews on this repo. Configure here.

…und trip

The cooldown filter (CooldownCache) and usage-based-routing-v2 selection
(LowestTPMLoggingHandler_v2) each issued their own MGET on every request
because they live in different objects. RoutingReadBatch fetches both key
sets through DualCache.async_batch_get_cache_shared while the healthy
deployments are resolved and hands the usage slice to the strategy, so
selection does not read again. Each cache keeps its own memory tier,
throttling, reservation rollback and circuit-breaker handling, and the
strategy falls back to its own read when the prefetch does not cover its
keys. simple-shuffle keeps reading only cooldowns.

aresponses no longer issues a second, blocking response-cache read from
the worker thread that runs the sync wrapper.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Optimizes router cache reads to batch cooldown and usage counters together.

The PR appears safe to merge; no outstanding findings or new actionable issues remain.

Summary

This PR combines cooldown and usage-counter Redis reads for asynchronous usage-based routing while retaining separate cache-tier handling.

  • Reuses prefetched counters during deployment selection.
  • Adds tests for routing parity, cache failures, and the single-read path.

Reviews (7) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

Comment thread litellm/router_utils/routing_read_batch.py Outdated
Comment thread litellm/caching/dual_cache.py Outdated
@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.62162% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...tellm/integrations/SlackAlerting/slack_alerting.py 0.00% 4 Missing ⚠️
litellm/router.py 85.71% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@codspeed

codspeed Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_router_single_redis_round_trip (a71e6ff) with main (7f95b5f)

Open in CodSpeed

yassin-berriai and others added 2 commits September 26, 2026 17:37
Wrap the memory-tier prepare and backfill steps of DualCache.async_batch_get_cache_shared
so a failing tier degrades that cache's read to None the way async_batch_get_cache does,
instead of escaping into routing. Drop the aresponses sync-cache guard: for native
Responses models the worker-thread read is the one whose key matches the write, so
skipping it broke cached /v1/responses replays.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…eck reads it as a key helper

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai review 31f81c5 please: 8d55b19 wraps the per-cache tier steps of the shared read and drops the aresponses sync-cache guard, 31f81c5 renames the usage key builder for the async cache-call check

…ison

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai review 5bb2c36 please: 5bb2c36 narrows the daily-report cache values in SlackAlerting.send_daily_reports with isinstance(val, (int, float)) before the numeric comparisons, which the basedpyright gate flagged once DualCache.async_batch_get_cache returns a typed list[object | None]; no routing or cache change since 8d55b19

Comment thread litellm/caching/dual_cache.py Outdated
Comment thread litellm/caching/dual_cache.py Outdated
Comment thread litellm/caching/dual_cache.py Outdated
… results without mutation

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai review 464b584 please: 464b584 types _prepare_batch_get / _apply_batch_get, merges the Redis slice into a new list instead of writing into pending.result, and drops the three carried-over comments; no routing or Redis-traffic change

yassin-berriai and others added 2 commits September 27, 2026 08:58
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai please re-review at fedb248. Since the 5/5 on 464b584 the branch merged origin/main (ed3e4c4) and fixed one import-sort lint in tests/unit/caching/test_dual_cache.py (fedb248); no functional change to litellm/.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

…omprehension

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

@greptileai

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

bugbot run

1 similar comment
@yassin-berriai

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit a71e6ff. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants