Skip to content

feat(cache): type prompt cache keys across chat APIs - #1247

Merged
tbille merged 2 commits into
mainfrom
feat/messages-prompt-cache-key
Aug 7, 2026
Merged

tbille merged 2 commits into
mainfrom
feat/messages-prompt-cache-key

Conversation

@HareeshBahuleyan

@HareeshBahuleyan HareeshBahuleyan commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Description

prompt_cache_key was available only through untyped provider kwargs on Chat Completions and Messages, so schema-derived gateways (Otari) could not expose it safely. This adds it to both typed parameter models and validates provider capability before dispatch.

OpenAI accepts the field, Otari and custom OpenAI-compatible endpoints pass it to the service that resolves support, and known unsupported providers raise UnsupportedParameterError before an SDK or HTTP call.

PR Type

  • 🆕 New Feature

Relevant issues

Fixes #1243

Follow-up: mozilla-ai/otari#514

Checklist

  • I understand the code I am submitting.
  • I have added unit tests that prove my fix/feature works
  • I have run this code locally and verified it fixes the issue.
  • New and existing tests pass locally
  • Documentation was updated where necessary
  • I have read and followed the contribution guidelines
  • AI Usage:
    • No AI was used.
    • AI was used for drafting/refactoring.
    • This is fully AI-generated.

AI Usage Information

  • AI Model used: GPT-5.6 Sol
  • AI Developer Tool used: Pi
  • Any other info you'd like to share: Implemented with test-driven development and reviewed interactively with the contributor.

When answering questions by the reviewer, please respond yourself, do not copy/paste the reviewer comments into an AI system and paste back its answer. We want to discuss with you, not your AI :)

  • I am an AI Agent filling out this form (check box if true)

Summary by CodeRabbit

  • New Features

    • Added optional prompt cache key support to completion and Messages APIs.
    • Synchronous and asynchronous requests now forward cache keys where supported.
    • Unsupported providers now return a clear validation error.
  • Tests

    • Expanded coverage for supported, passthrough, and unsupported provider scenarios.
    • Verified cache keys are correctly forwarded across completion and Messages requests.

Expose prompt_cache_key through both MessagesParams and CompletionParams so schema-derived gateways can accept it without relying on provider kwargs.

Validate provider capability centrally: OpenAI supports the field, Otari and custom OpenAI-compatible endpoints pass it through, and other providers reject it with UnsupportedParameterError.
@coderabbitai

coderabbitai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: ad0969c0-471d-432b-8f9c-3aabe6cec4ef

📥 Commits

Reviewing files that changed from the base of the PR and between c8bcf28 and c0f34de.

📒 Files selected for processing (2)
  • tests/unit/test_completion.py
  • tests/unit/test_messages.py

Walkthrough

The PR adds prompt_cache_key to completion and Messages APIs, parameter models, provider capability declarations, validation, provider forwarding, and compatibility conversion paths. Unsupported providers raise UnsupportedParameterError before provider execution.

Changes

Prompt cache key support

Layer / File(s) Summary
API contracts and validation
src/any_llm/any_llm.py, src/any_llm/api.py, src/any_llm/types/*, tests/unit/test_completion.py, tests/unit/test_messages.py
Completion and Messages interfaces accept prompt_cache_key. Parameter models expose the field. Unsupported providers reject it before client calls.
Provider capability and forwarding
src/any_llm/providers/openai/*, src/any_llm/providers/otari/otari.py, tests/unit/providers/*
Providers declare supported or passthrough behaviour. OpenAI and Otari tests verify forwarding. Anthropic tests verify early rejection.
Messages-to-completion forwarding
src/any_llm/utils/messages_compat.py, tests/unit/test_messages.py
The compatibility bridge includes a supplied prompt cache key in converted completion parameters. Tests verify the converted object and call arguments.

Possibly related issues

  • mozilla-ai/otari#514 — Relates to exposing and forwarding prompt_cache_key through Otari gateway schemas.

Possibly related PRs

Suggested labels: 1.24.0

Suggested reviewers: njbrake

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 22.22% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly describes the main change: typed prompt cache key support across chat APIs.
Description check ✅ Passed The description follows the repository template and includes the change, issue, checklist, tests, and AI usage details.
Linked Issues check ✅ Passed The changes satisfy [#1243] by adding typed support, forwarding the key, rejecting Anthropic usage early, and adding tests.
Out of Scope Changes check ✅ Passed The changes remain within scope and do not add the excluded prompt_cache_retention parameter.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/messages-prompt-cache-key

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Aug 6, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
src/any_llm/any_llm.py 78.27% <100.00%> (-0.85%) ⬇️
src/any_llm/api.py 93.85% <ø> (ø)
src/any_llm/providers/openai/custom.py 100.00% <100.00%> (ø)
src/any_llm/providers/openai/openai.py 100.00% <100.00%> (ø)
src/any_llm/providers/otari/otari.py 93.82% <100.00%> (+0.01%) ⬆️
src/any_llm/types/completion.py 97.60% <100.00%> (+0.03%) ⬆️
src/any_llm/types/messages.py 95.49% <100.00%> (+0.08%) ⬆️
src/any_llm/utils/messages_compat.py 100.00% <100.00%> (ø)

... and 32 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/unit/test_completion.py`:
- Around line 18-30: Extend test_completion_params_exposes_prompt_cache_key with
a synchronous mock-provider invocation of completion(), passing
prompt_cache_key="tenant-1". Assert the provider.completion() call receives and
forwards that value, while preserving the existing schema and signature
assertions.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 1d62df71-db69-480e-a043-42820f33f0c0

📥 Commits

Reviewing files that changed from the base of the PR and between b4bb154 and c8bcf28.

📒 Files selected for processing (14)
  • src/any_llm/any_llm.py
  • src/any_llm/api.py
  • src/any_llm/providers/openai/custom.py
  • src/any_llm/providers/openai/openai.py
  • src/any_llm/providers/otari/otari.py
  • src/any_llm/types/completion.py
  • src/any_llm/types/messages.py
  • src/any_llm/utils/messages_compat.py
  • tests/unit/providers/test_anthropic_messages.py
  • tests/unit/providers/test_openai_base_provider.py
  • tests/unit/providers/test_openai_compatible_provider.py
  • tests/unit/providers/test_otari_provider.py
  • tests/unit/test_completion.py
  • tests/unit/test_messages.py

Comment thread tests/unit/test_completion.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds first-class, typed prompt_cache_key support across the Chat Completions and Anthropic Messages-style APIs so schema-driven callers can safely provide it. The change introduces provider capability gating so unsupported providers fail fast with UnsupportedParameterError before any SDK or HTTP request is attempted.

Changes:

  • Add prompt_cache_key: str | None to CompletionParams and MessagesParams, and thread it through the public completion/acompletion/messages/amessages APIs.
  • Introduce AnyLLM.PROMPT_CACHE_KEY_SUPPORT plus centralized validation to reject unsupported providers pre-dispatch.
  • Update OpenAI, Otari, and OpenAI-compatible providers’ capability flags and expand unit test coverage for forwarding and rejection behavior.

Reviewed changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
tests/unit/test_messages.py Adds schema + forwarding assertions and a new unsupported-provider rejection test for Messages.
tests/unit/test_completion.py Adds schema + signature assertions and a new unsupported-provider rejection test for Completions.
tests/unit/providers/test_otari_provider.py Verifies Otari capability and asserts prompt_cache_key is forwarded on completion and messages paths.
tests/unit/providers/test_openai_compatible_provider.py Asserts OpenAI-compatible default capability is passthrough.
tests/unit/providers/test_openai_base_provider.py Adds capability-flag coverage and an HTTP transport test ensuring OpenAI sends prompt_cache_key.
tests/unit/providers/test_anthropic_messages.py Adds tests asserting Anthropic rejects prompt_cache_key before any HTTP call.
src/any_llm/utils/messages_compat.py Extends Messages→Completions bridge conversion to include prompt_cache_key when set.
src/any_llm/types/messages.py Adds prompt_cache_key to the typed MessagesParams model and schema.
src/any_llm/types/completion.py Adds prompt_cache_key to the typed CompletionParams model and schema.
src/any_llm/providers/otari/otari.py Marks Otari as prompt-cache-key passthrough.
src/any_llm/providers/openai/openai.py Marks OpenAI as prompt-cache-key supported.
src/any_llm/providers/openai/custom.py Marks OpenAI-compatible provider as prompt-cache-key passthrough.
src/any_llm/api.py Exposes prompt_cache_key as a typed argument on top-level API functions and forwards into parameter models.
src/any_llm/any_llm.py Introduces capability flag + validation and threads prompt_cache_key through provider entrypoints.
Suppressed comments (2)

tests/unit/test_messages.py:79

  • This test instantiates BedrockProvider, which can raise ImportError in environments without boto3 due to AnyLLM._verify_no_missing_packages. Since the purpose here is only to verify prompt_cache_key is rejected before any client call, use a lightweight AnyLLM subclass that does not require optional dependencies and still has PROMPT_CACHE_KEY_SUPPORT left as "unsupported".
    client = Mock()
    provider = BedrockProvider(client=client)

    with pytest.raises(UnsupportedParameterError, match="prompt_cache_key"):
        await provider.amessages(

tests/unit/test_completion.py:39

  • This test uses BedrockProvider to represent an unsupported provider, but BedrockProvider can raise ImportError if boto3 is not installed due to AnyLLM._verify_no_missing_packages. Since you only need to exercise AnyLLM's prompt_cache_key validation before any SDK call, use a tiny AnyLLM subclass with a mocked client instead of a provider with optional dependencies.
    client = Mock()
    provider = BedrockProvider(client=client)

    with pytest.raises(UnsupportedParameterError, match="prompt_cache_key"):
        await provider.acompletion(

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread tests/unit/test_messages.py Outdated
Comment thread tests/unit/test_completion.py Outdated
Cover synchronous completion forwarding and isolate optional Bedrock imports so general API test modules remain usable without boto3.
@tbille
tbille merged commit 430d968 into main Aug 7, 2026
14 checks passed
@tbille
tbille deleted the feat/messages-prompt-cache-key branch August 7, 2026 08:32
@github-actions github-actions Bot added the 1.25.0 Included in release 1.25.0 label Aug 11, 2026

This branch was previously deployed

1 inactive deployment
integration-tests — c0f34dee Deployed Aug 6, 2026 by HareeshBahuleyan via run-docs-tests #2325
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

1.25.0 Included in release 1.25.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Expose prompt_cache_key on MessagesParams

3 participants