Repository navigation
feat(cache): select Rust caching through explicit cache objects - #43601
Conversation
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
39efa47 to
3ea68da
Compare
|
@greptileai Please review the latest commit, including cache admission, supported call types, reduced Python scope, and the four remaining open findings |
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
@greptileai Please review 9de79d8, which fixes all four findings with regression tests for Redis flush, callbacks, resolved identity, and opt-in mode |
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
| _DURATION: Final[TypeAdapter[float | None]] = TypeAdapter(float | None) | ||
|
|
||
|
|
||
| class NativeBackend(BaseCache): |
|
@greptileai Please review 7456f93, especially Python cache delegation from Rust Messages, inline async operations, cancellation, and regression coverage |
|
@greptileai Please review 7c4658f, which separates shared cache selection, native backend bindings, and Python delegation with scoped directory instructions |
|
@greptileai Please review the latest cache restructuring, focusing on adapter boundaries, shared selection, and preservation of existing behavior |
TLDR
Problem this solves:
How it solves it:
litellm._v2.cacheMessages calls retain Rust inference when a legacy Python cache is configured. Rust delegates reads and writes to that cache through the Python host, with async operations awaited in the caller's task. Rust omits the
cache_keyoverride when delegating to Python, so Python owns namespace and semantic-cache key rules. The adapter stores ordinary Messages responses and Python-compatible stream entries, allowing both implementations to read the same entries. V2 caches use native Rust storage directly. Chat Completions and Responses still fall back to Python inference for legacy cachesThe existing cache facade retains an optional backend injection parameter for the experimental factories and selects backends without cache catalog rules
The cache root exposes the API used by routes and module registration. Shared selection and inference composition live in
selection.rs, and the Python-facing runtime lives inruntime.rs. Privatenative/andpython/directories implement the adapters;python/host.rshandles Python calls during Rust inference. Neither adapter imports shared selection or selects the other adapter. Scoped instructions document these boundariesPre-Submission checklist
Caveats
High
Medium
Low
Validation
Before removing the Rust key override, both namespace cases and the semantic-cache scope regression failed with Python inference disabled. After removal they pass. The four cross-language response and stream cases failed before the format conversion and now pass, checking equal payloads and exactly one provider request
All 82 cases in
tests/test_litellm_rust/cache/test_v2.pypass with a freshly rebuilt extension. Rust inference is required in the native test calls, with Python fallback disabled. Semantic coverage verifies the Python cache facade's key scope with injected memory storage, not embedding similaritymake check, workspace formatting, and Clippy for the core and Python bridge crates pass. Live external-provider validation has not been run. CI, coverage, and review bots must still be checked on the new tip before maintainer reviewType
Refactoring
Final Attestation