Skip to content

feat(cache): select Rust caching through explicit cache objects - #43601

Merged
yujonglee-berri merged 14 commits into
mainfrom
litellm_cache_object_backend
Sep 29, 2026
Merged

yujonglee-berri merged 14 commits into
mainfrom
litellm_cache_object_backend

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Rust routes need an explicitly injected response-cache service
  • Cache selection should follow configured objects, without cache catalog rules

How it solves it:

  • Shares typed cache lookup, storage, and replay across Rust routes
  • Native cache keys include resolved credentials and callback changes
  • Scopes Redis flushes and honors existing cache opt-in mode
  • Separates execution facts from Python-owned accounting and coordination
  • Keeps experimental factories behind litellm._v2.cache
  • Uses configured Python caches while Messages inference stays native

Messages calls retain Rust inference when a legacy Python cache is configured. Rust delegates reads and writes to that cache through the Python host, with async operations awaited in the caller's task. Rust omits the cache_key override when delegating to Python, so Python owns namespace and semantic-cache key rules. The adapter stores ordinary Messages responses and Python-compatible stream entries, allowing both implementations to read the same entries. V2 caches use native Rust storage directly. Chat Completions and Responses still fall back to Python inference for legacy caches

The existing cache facade retains an optional backend injection parameter for the experimental factories and selects backends without cache catalog rules

The cache root exposes the API used by routes and module registration. Shared selection and inference composition live in selection.rs, and the Python-facing runtime lives in runtime.rs. Private native/ and python/ directories implement the adapters; python/host.rs handles Python calls during Rust inference. Neither adapter imports shared selection or selects the other adapter. Scoped instructions document these boundaries

Pre-Submission checklist

  • I have added meaningful tests
  • The focused test files pass locally
  • All required CI/CD checks pass
  • The scope is limited to response-cache foundations
  • Greptile confidence is at least 4/5

Caveats

High

  • Legacy hit headers report Rust keys, not Python storage keys

Medium

  • SigV4 requests bypass caching until signing identity is represented

Low

  • Existing Python nested cache-cost logging remains unchanged
  • Legacy cache delegation currently covers Messages only
  • Native V2 entries retain separate keys and typed envelopes
  • Legacy keys exclude credentials and Rust request callback changes

Validation

Before removing the Rust key override, both namespace cases and the semantic-cache scope regression failed with Python inference disabled. After removal they pass. The four cross-language response and stream cases failed before the format conversion and now pass, checking equal payloads and exactly one provider request

All 82 cases in tests/test_litellm_rust/cache/test_v2.py pass with a freshly rebuilt extension. Rust inference is required in the native test calls, with Python fallback disabled. Semantic coverage verifies the Python cache facade's key scope with injected memory storage, not embedding similarity

make check, workspace formatting, and Clippy for the core and Python bridge crates pass. Live external-provider validation has not been run. CI, coverage, and review bots must still be checked on the new tip before maintainer review

Type

Refactoring

Final Attestation

  • Reviewed edge cases and open findings are resolved

@devin-ai-integration

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[High risk] Refactors caching architecture across core inference and gateway.

The PR is not yet safe to merge because legacy-cache Messages calls can still reuse entries across requested namespaces and lose semantic matches across differently worded prompts.

Findings

  1. P1 Security Per-request namespaces are ignored ▶
  2. P1 Semantic matches become exact ▶

Summary

This PR injects response-cache services into Rust routes and selects native or Python-backed caches from configured objects. The latest changes separate shared selection and runtime code from the native and Python adapters without changing their selection logic.

  • Messages retains Python-cache delegation while using Rust inference.
  • The two unresolved findings about per-request namespaces and semantic matching remain outstanding.

Reviews (6) · Last reviewed commit: "refactor(cache): enforce shared composit..."

Comment thread litellm-rust/crates/python-bridge/src/cache/native/v2.rs
Comment thread litellm-rust/crates/core/src/caching.rs Outdated
Comment thread litellm-rust/crates/core/src/messages/route.rs Outdated
Comment thread litellm-rust/crates/python-bridge/src/cache/native/v2.rs
@yujonglee-berri
yujonglee-berri force-pushed the litellm_cache_object_backend branch from 39efa47 to 3ea68da Compare September 28, 2026 21:55
@yujonglee-berri

Copy link
Copy Markdown
Contributor

@greptileai Please review the latest commit, including cache admission, supported call types, reduced Python scope, and the four remaining open findings

yujonglee-berri and others added 2 commits September 28, 2026 22:06
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@yujonglee-berri

Copy link
Copy Markdown
Contributor

@greptileai Please review 9de79d8, which fixes all four findings with regression tests for Redis flush, callbacks, resolved identity, and opt-in mode

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 21.73913% with 54 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/_v2/cache/__init__.py 0.00% 46 Missing ⚠️
litellm/rust_bridge/public_call.py 0.00% 4 Missing ⚠️
litellm/_v2/__init__.py 0.00% 2 Missing ⚠️
litellm/caching/caching.py 75.00% 1 Missing ⚠️
litellm/rust_bridge/response_cache.py 0.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

_DURATION: Final[TypeAdapter[float | None]] = TypeAdapter(float | None)


class NativeBackend(BaseCache):
@codspeed

codspeed Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_cache_object_backend (f6d82a3) with main (eae8ed7)

Open in CodSpeed

@yujonglee-berri

Copy link
Copy Markdown
Contributor

@greptileai Please review 7456f93, especially Python cache delegation from Rust Messages, inline async operations, cancellation, and regression coverage

Comment thread litellm-rust/crates/python-bridge/src/cache/hosted.rs Outdated
Comment thread litellm-rust/crates/python-bridge/src/cache/hosted.rs Outdated
@yujonglee-berri

Copy link
Copy Markdown
Contributor

@greptileai Please review 7c4658f, which separates shared cache selection, native backend bindings, and Python delegation with scoped directory instructions

@yujonglee-berri

Copy link
Copy Markdown
Contributor

@greptileai Please review the latest cache restructuring, focusing on adapter boundaries, shared selection, and preservation of existing behavior

@yujonglee-berri
yujonglee-berri enabled auto-merge (squash) September 28, 2026 23:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants