Skip to content

feat(caching): capture and price cached tokens across providers - #1114

Merged
murdore merged 1 commit into
releasefrom
feat/generic-prompt-caching
Jul 8, 2026
Merged

murdore merged 1 commit into
releasefrom
feat/generic-prompt-caching

Conversation

@pdogra1299

@pdogra1299 pdogra1299 commented Jun 24, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Makes prompt-cache usage visible and correctly priced for every provider, not just Claude. Telemetry showed 0 cache reads across all non-Claude main-flow models even though Gemini/OpenAI cache automatically. Root cause was three stacked gaps; this is a generic, additive fix.

Builds on 9.79.1 (fix(vertex): enable prompt caching on native Claude paths), which added Anthropic cache_control breakpoints. This PR is the observability + cost half, generalized across providers.

Type of Change

  • Bug fix (cost mis-pricing of cached tokens)
  • New feature (cross-provider cached-token capture)
  • Performance improvement (cached tokens now billed at the cheaper rate)

Motivation and Context

Non-Claude models reported zero cache usage because:

  1. The extractor was Anthropic-only — tokenUtils.extractCacheReadTokens read only cacheReadInputTokens/cacheReadTokens, missing OpenAI's prompt_tokens_details.cached_tokens and Gemini's cachedContentTokenCount.
  2. Native Gemini paths never read cachedContentTokenCount (Vertex ×2 + AI Studio built usage by hand).
  3. No cache pricing for google/openai models — so even captured cached tokens couldn't be discounted.

For Gemini/OpenAI, caching already happens automatically; this surfaces and prices it.

⚠️ Token-overlap correctness (the core risk)

Conventions differ and getting it wrong double-counts cost:

  • Anthropic / Vertex-Claude report input excluding cache (non-overlapping) — left untouched.
  • Gemini / OpenAI report input including the cached subset (overlapping) — so cached tokens are subtracted from input before being surfaced as cacheReadTokens (guarded by cached <= input, total conserved).

calculateCost prices cacheRead at the discounted rate where one exists, else falls back to the input rate — so any provider that emits cached tokens but has no cacheRead rate bills identically to before (no silent undercharge), while priced providers get the discount.

Changes Made

  • tokenUtils.ts — new extractCachedInputTokensOverlapping; overlap-safe fallback in extractTokenUsage (subtract from input).
  • googleVertex.ts (Gemini generate+stream), googleAiStudio.ts, googleNativeGemini3.ts — extract cachedContentTokenCount, surface as cacheReadTokens with input adjustment. Vertex Claude path untouched.
  • pricing.ts — cacheRead rates (0.25× input) for openai + google models; calculateCost input-rate fallback. Vertex Gemini covered by the existing vertex→google fallback.
  • types/common.ts, types/providers.ts — additive optional fields.
  • test/continuous-test-suite-cache-breakpoints.ts — +13 assertions (overlap subtraction, total conservation, malformed cached>input guard, non-overlapping untouched, per-provider cacheRead rates, no-undercharge fallback). 33/33 pass.

Breaking Changes

  • No breaking changes — all fields additive; Anthropic/Claude behavior unchanged; unpriced providers' cost unchanged.

Testing

  • Unit tests added (pnpm run test:cache — 33/33)
  • tsc --noEmit --strict — 0 errors
  • pnpm run lint — clean
  • pnpm run build — 0 errors, publint clean

This change was developed via a multi-agent workflow (parallel per-provider mapping → design → implement → adversarial verify); the verify pass caught and fixed an undercharge edge case (the input-rate fallback above).

Commit Message Format

  • feat(caching): capture and price cached tokens across providers

Deployment Notes

  • No special steps. Caching for Gemini/OpenAI is provider-side and automatic; this only captures + prices it. Pairs with bumping the consuming app to this release.

Summary by CodeRabbit

  • New Features
    • Added end-to-end cached-token accounting for Gemini native streaming/generation (including Vertex) and overlapping cache formats, so reported usage totals correctly reflect cache reads.
    • Expanded pricing and billing to charge cache-read tokens with model-specific rates when available (with safe fallbacks).
  • Bug Fixes
    • Prevented double-counting of cached tokens in both token usage and cost calculations, including when cached tokens are embedded in nested usage fields.
  • Tests
    • Added coverage for overlapping cached-token extraction, edge cases (e.g., malformed inputs), and cache-read pricing behavior across providers/models.

@vercel

vercel Bot commented Jun 24, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
neurolink Ready Ready Preview, Comment Jul 8, 2026 3:08am

@coderabbitai

coderabbitai Bot commented Jun 24, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds cached-token tracking to Gemini 3 provider usage paths, overlapping cache extraction for OpenAI-compatible payloads, cache-read pricing rates and fallback billing, and tests covering extraction and cost accounting.

Changes

Cache Token Billing

Layer / File(s) Summary
Cache token type contracts
src/lib/types/common.ts, src/lib/types/providers.ts
RawUsageObject gains optional prompt_tokens_details.cached_tokens; CollectedChunkResult gains optional cacheReadTokens and cacheCreationTokens.
Gemini3 native chunk collectors
src/lib/providers/googleNativeGemini3.ts
collectStreamChunks and collectStreamChunksIncremental read usageMetadata.cachedContentTokenCount, track the maximum cached-token value, and return it as cacheReadTokens.
Google AI Studio cached-token accounting
src/lib/providers/googleAiStudio.ts
The native stream and generate loops accumulate totalCacheReadTokens, subtract it from billed input tokens in final usage objects, and include cacheReadTokens when present.
Google Vertex cached-token accounting
src/lib/providers/googleVertex.ts
The native stream and generate loops accumulate usageMetadata.cachedContentTokenCount, then compute adjusted input totals and return cacheReadTokens in StreamResult.usage and EnhancedGenerateResult.usage.
OpenAI overlapping cache token extraction
src/lib/utils/tokenUtils.ts
extractCachedInputTokensOverlapping reads prompt_tokens_details.cached_tokens, and extractTokenUsage falls back to it when no non-overlapping cache field is present.
Pricing tables and calculateCost fallback
src/lib/utils/pricing.ts
OpenAI and Gemini model pricing entries add cacheRead rates, and calculateCost() bills cached tokens at rates.cacheRead ?? rates.input.
Tests for cache extraction and pricing
test/continuous-test-suite-cache-breakpoints.ts
Adds testOverlappingCacheExtraction() for overlapping cache extraction, token subtraction, malformed inputs, and cache-read pricing fallback, then wires it into the suite runner.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant GoogleAiStudioProvider
  participant GeminiChunkCollector
  participant UsageBuilder
  GoogleAiStudioProvider->>GeminiChunkCollector: collect stream or generate step
  GeminiChunkCollector-->>GoogleAiStudioProvider: cacheReadTokens
  GoogleAiStudioProvider->>UsageBuilder: compute adjustedInputTokens and total
  UsageBuilder-->>GoogleAiStudioProvider: final usage with cacheReadTokens
Loading
sequenceDiagram
  participant TokenUtils
  participant Pricing
  participant Tests
  Tests->>TokenUtils: extract prompt_tokens_details.cached_tokens
  TokenUtils-->>Tests: cacheReadTokens and adjusted input
  Tests->>Pricing: calculateCost(cacheReadTokens)
  Pricing-->>Tests: billed cached tokens with fallback
Loading

Possibly related PRs

  • juspay/neurolink#1113: Also propagates cache-read token accounting into provider usage and cost computation paths.

Suggested labels: released

Suggested reviewers: Tara-ag

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 75.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: capturing and pricing cached tokens across providers.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/generic-prompt-caching

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Jun 24, 2026 •

Copy link
Copy Markdown
Contributor

✅ Single Commit Policy - COMPLIANT

Status: Policy requirements met • 1 commit • Valid format • Ready for merge

📊 View validation details

📝 Commit Details

  • Hash: cd5bb905839e3cc2b54b35e0b606f88d0db508a6
  • Message: feat(caching): capture and price cached tokens across providers
  • Author: Parth Dogra

✅ Validation Results

  • Single commit requirement met
  • No merge commits in branch
  • Semantic commit message format verified
  • Ready for squash merge to release branch

🤖 Automated validation by NeuroLink Single Commit Enforcement

@pdogra1299
pdogra1299 requested a review from Tara-ag June 24, 2026 06:37
@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Summary

Files reviewed: 8
New issues raised: 2 (both MINOR/SUGGESTION)
Blocking issues: 0

Assessment

This PR correctly implements cross-provider cached token capture and pricing. The implementation demonstrates solid engineering:

Strengths:

  1. Correct handling of overlapping vs non-overlapping conventions - The extractCachedInputTokensOverlapping function properly identifies OpenAI/Gemini's nested prompt_tokens_details.cached_tokens path and the subtraction logic in extractTokenUsage prevents double-counting.

  2. Safety guards - The cached <= input check prevents malformed responses from producing negative input or inflated totals.

  3. Graceful fallback - The calculateCost fallback to rates.input when cacheRead rate is missing ensures no silent undercharging for providers that report cached tokens but lack pricing.

  4. Comprehensive test coverage - 13 new assertions covering overlap subtraction, total conservation, malformed data guards, and pricing fallbacks.

  5. Additive changes only - No breaking changes to the public SDK API; all new fields are optional.

Minor observations (non-blocking):

  • Floating-point precision in pricing calculations (cosmetic, mitigated by final rounding)
  • Could benefit from one additional edge case test for cached_tokens: 0

Verification

  • ✅ No hardcoded secrets
  • ✅ No security vulnerabilities
  • ✅ No breaking API changes
  • ✅ All CLAUDE.md critical rules respected
  • ✅ Tests pass (33/33)

Recommendation: Approve for merge.

Comment thread src/lib/utils/pricing.ts
cost += (usage.output || 0) * rates.output;
if (usage.cacheReadTokens && rates.cacheRead) {
cost += usage.cacheReadTokens * rates.cacheRead;
if (usage.cacheReadTokens) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Minor: Floating-point precision consideration

The pricing calculations use division that could result in floating-point representation issues:

cacheRead: 0.0625 / 1_000_000, // = 6.25e-8

While the final rounding in calculateCost() mitigates display issues, consider documenting that these are pre-computed rates (0.25× input) so future maintainers know to preserve the mathematical relationship when updating prices.

Alternatively, consider computing cacheRead dynamically as input * 0.25 to ensure consistency:

const geminiRates = { input: 2.0 / 1_000_000, output: 12.0 / 1_000_000 };
return {
  ...geminiRates,
  cacheRead: geminiRates.input * 0.25, // Ensures 0.25× relationship always holds
};

@@ -247,10 +252,145 @@ function testEdgeCases(): void {
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Suggestion: Additional edge case test for zero cached tokens

Consider adding a test case for when cached_tokens: 0 is explicitly returned (vs. undefined/absent):

recordTest(
  "cached_tokens: 0 treated same as absent (no cacheReadTokens field)",
  extractTokenUsage({
    promptTokens: 1000,
    completionTokens: 200,
    prompt_tokens_details: { cached_tokens: 0 },
  }).cacheReadTokens === undefined
);

This ensures the > 0 check in extractCachedInputTokensOverlapping behaves consistently with downstream consumers that may treat 0 differently from undefined.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/providers/googleVertex.ts`:
- Around line 1663-1671: The usage aggregation in the Gemini stream and generate
loops is carrying the latest cache count forward instead of accumulating per
agentic step, which can misattribute cached tokens across tool-loop turns.
Update the logic around the step/usage handling in googleVertex.ts so each step
tracks its own latest prompt and cache values, clamps cache reads to that step’s
prompt, and then adds both input and cache totals into the cumulative counters
before constructing usage. Apply the same fix consistently in the affected
stream/generate paths so the final usage object reflects the correct per-step
billing split.

In `@src/lib/utils/pricing.ts`:
- Around line 739-748: Update the direct Anthropic model rate definitions so
`claude-3-sonnet` and `claude-3-haiku` include a `cacheRead` rate, matching the
existing pattern used by the Vertex Claude entries. This will ensure the
`pricing.ts` cost calculation path that uses `usage.cacheReadTokens` in the
pricing logic can apply the discounted cache-read price instead of falling back
to `rates.input` for these two models.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: a3f7a9ab-d0f8-4698-8d44-a72bbfa74b79

📥 Commits

Reviewing files that changed from the base of the PR and between 8f0fda3 and 7782bd6.

📒 Files selected for processing (8)
  • src/lib/providers/googleAiStudio.ts
  • src/lib/providers/googleNativeGemini3.ts
  • src/lib/providers/googleVertex.ts
  • src/lib/types/common.ts
  • src/lib/types/providers.ts
  • src/lib/utils/pricing.ts
  • src/lib/utils/tokenUtils.ts
  • test/continuous-test-suite-cache-breakpoints.ts

Comment thread src/lib/providers/googleVertex.ts Outdated
Comment thread src/lib/utils/pricing.ts
@pdogra1299
pdogra1299 force-pushed the feat/generic-prompt-caching branch from 7782bd6 to f3db31c Compare June 24, 2026 06:55
@pdogra1299

Copy link
Copy Markdown
Collaborator Author

Thanks — addressed in the amended commit.

🟠 Major — stale cache subtracted across agentic steps (googleVertex.ts)

Fixed. Root cause confirmed: in the two native Gemini loops (executeNativeGemini3Stream + executeNativeGemini3Generate), totalInputTokens updated on promptTokenCount > 0 while totalCacheReadTokens updated separately on cachedContentTokenCount > 0 — so a cached early step followed by an uncached final step left a stale cache value that got subtracted from a different step's prompt.

Fix: read cachedContentTokenCount from the same usageMetadata as promptTokenCount and clamp to it, inside the input-update block:

if (usageMetadata.promptTokenCount !== undefined && usageMetadata.promptTokenCount > 0) {
  totalInputTokens = usageMetadata.promptTokenCount;
  totalCacheReadTokens = Math.min(
    usageMetadata.cachedContentTokenCount ?? 0,
    usageMetadata.promptTokenCount,
  );
}

A later uncached step now resets cache to 0 alongside its input, so the split can never desync. I kept the existing "latest" input semantics rather than switching to per-step summing — summing pre-existing input accounting is out of scope for this caching PR and would change billed numbers for multi-step Gemini calls (each step re-sends the growing prefix, so summing would over-count it).

I also verified the other Gemini paths were already consistent, so no change was needed there:

  • googleAiStudio.ts accumulates both totalInputTokens += chunkResult.inputTokens and totalCacheReadTokens += chunkResult.cacheReadTokens (same convention → no desync).
  • googleNativeGemini3.ts collectors take Math.max(...) of both per response (co-reported in the final chunk → consistent).

🟡 Minor — Claude pricing cacheRead coverage (pricing.ts)

No change needed. Anthropic and Vertex Claude models already carry cacheRead rates in the anthropic/vertex tables, and Bedrock resolves via the bedrock → anthropic alias. Additionally, this PR's calculateCost change prices cacheRead at rates.cacheRead ?? rates.input — so even a model that populates cacheReadTokens without a discount rate bills those tokens at the input rate (never $0), i.e. no silent undercharge.

Verification: tsc --strict 0 errors · lint clean · test:cache 33/33 · build clean. Kept to a single commit (amended).

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Summary

PR #1114: feat(caching): capture and price cached tokens across providers

Files Reviewed: 8

  • src/lib/utils/pricing.ts - Added cacheRead rates for OpenAI and Google models
  • src/lib/utils/tokenUtils.ts - Added overlapping cache token extraction for OpenAI/DeepSeek
  • src/lib/providers/googleVertex.ts - Added cachedContentTokenCount extraction for Gemini
  • src/lib/providers/googleAiStudio.ts - Added cachedContentTokenCount extraction for Gemini
  • src/lib/providers/googleNativeGemini3.ts - Added cacheReadTokens to chunk collectors
  • src/lib/types/common.ts - Added prompt_tokens_details type
  • src/lib/types/providers.ts - Added cacheReadTokens to CollectedChunkResult
  • test/continuous-test-suite-cache-breakpoints.ts - Added 13 new test assertions

Issues Found: 1 MAJOR

⚠️ MAJOR: Missing cacheRead rates for legacy Claude 3 models

The direct Anthropic entries for claude-3-sonnet and claude-3-haiku are missing cacheRead and cacheCreation rates while all other Claude models have them. This means cached tokens on these models will be billed at the full input rate instead of the discounted 0.1x cache-read rate.

Existing Comments Acknowledged

  • CodeRabbit's finding about missing cacheRead rates for these models is valid and should be addressed
  • Tara-ag's suggestions about floating-point precision and additional edge case tests are minor and can be addressed in follow-up

Overall Assessment

The PR correctly implements the overlapping cache token convention for Gemini/OpenAI (where cached tokens are a subset of input) while preserving the non-overlapping convention for Anthropic/Claude. The token accounting logic properly subtracts cached tokens from input to avoid double-counting, and the fallback in calculateCost prevents silent undercharging for providers without cache rates.

Action Required: Add cacheRead and cacheCreation rates for claude-3-sonnet and claude-3-haiku in pricing.ts to ensure consistent cache pricing across all Claude models.

Comment thread src/lib/utils/pricing.ts

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ MAJOR: Missing cacheRead rates for legacy Claude 3 models

The direct Anthropic entries for claude-3-sonnet and claude-3-haiku are missing cacheRead rates while all other Claude models have them:

"claude-3-sonnet": { input: 3.0 / 1_000_000, output: 15.0 / 1_000_000 },
"claude-3-haiku": { input: 0.25 / 1_000_000, output: 1.25 / 1_000_000 },

Per the PR description, Anthropic/Vertex Claude models use the non-overlapping convention where cache_read_input_tokens is reported separately. If these models return cache read tokens, the calculateCost function will fall back to billing them at the full input rate instead of the discounted 0.1x cache-read rate.

Suggested fix: Add cacheRead rates consistent with other Claude models (0.1x input for cache reads, 1.25x for cache creation):

"claude-3-sonnet": { 
  input: 3.0 / 1_000_000, 
  output: 15.0 / 1_000_000,
  cacheRead: 0.3 / 1_000_000,
  cacheCreation: 3.75 / 1_000_000,
},
"claude-3-haiku": { 
  input: 0.25 / 1_000_000, 
  output: 1.25 / 1_000_000,
  cacheRead: 0.025 / 1_000_000,
  cacheCreation: 0.3125 / 1_000_000,
},

@pdogra1299
pdogra1299 force-pushed the feat/generic-prompt-caching branch from f3db31c to 675d55a Compare June 24, 2026 07:14
@pdogra1299

Copy link
Copy Markdown
Collaborator Author

Thanks @Tara-ag — addressed in the amended commit.

⚠️ MAJOR — missing cacheRead rates for claude-3-sonnet / claude-3-haiku

Fixed. Added cacheRead (0.1×) and cacheCreation (1.25×) to both entries in the anthropic table, matching the convention used by every other Claude model:

"claude-3-sonnet": { input: 3.0/1M, output: 15.0/1M, cacheRead: 0.3/1M, cacheCreation: 3.75/1M },
"claude-3-haiku":  { input: 0.25/1M, output: 1.25/1M, cacheRead: 0.025/1M, cacheCreation: 0.3125/1M },

I audited the rest of the anthropic + vertex Claude tables — these two were the only models missing cache rates; all others already have them. This also resolves CodeRabbit's pricing-coverage comment.

💡 Minor — cached_tokens: 0 edge case

Added a test. Asserts an explicit cached_tokens: 0 behaves exactly like absent — extractCachedInputTokensOverlapping returns undefined, and extractTokenUsage surfaces no cacheReadTokens and leaves input intact (the > 0 guard). test:cache is now 34/34.

💡 Minor — floating-point / rate documentation

The new OpenAI + Google cacheRead rates already carry an inline // cacheRead = 0.25× input note, and the Claude entries follow the table's established 0.1× / 1.25× convention (consistent with claude-3-5-haiku, claude-3-opus, etc.). The final Math.round(cost * 1e6) / 1e6 in calculateCost bounds any FP representation noise. No code change needed.

Verification: tsc --strict 0 errors · lint clean · test:cache 34/34. Single commit (amended).

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Summary

Files Reviewed: 8 files
New Issues Raised: 0
Status: ✅ APPROVE

Analysis

This PR implements cross-provider cached token capture and pricing correctly:

Architecture:

  • Properly distinguishes overlapping (Gemini/OpenAI) vs non-overlapping (Anthropic) token conventions
  • Correctly subtracts cached tokens from input for overlapping providers to avoid double-counting
  • Per-step cache accumulation with clamping in agentic loops (googleVertex.ts)

Pricing:

  • Cache read rates added for all OpenAI and Google models (0.25× input)
  • Legacy Claude 3 models (sonnet, haiku) now have cacheRead rates (0.1× input)
  • Safe fallback to input rate when cacheRead rate unavailable (prevents undercharge)

Safety:

  • Guard against malformed responses: cached <= input check prevents negative input
  • Zero cached tokens handled correctly (treated as absent)
  • Total token conservation verified in tests

Testing:

  • 33/33 tests passing
  • Coverage includes overlap subtraction, malformed input guards, non-overlapping preservation, and pricing fallbacks

Existing Comments

  • One minor suggestion from Tara-ag on floating-point precision remains open (non-blocking)
  • All other review comments have been addressed in the code

The PR is ready to merge.

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/utils/pricing.ts`:
- Around line 164-252: The cached-input pricing in the pricing table is stale
and still uses the old blanket multiplier. Update the model entries in the
pricing map in pricing.ts so each affected model has the correct cacheRead rate,
especially gpt-5.4, gpt-4o, o1, and o3-mini, using the existing model keys in
the pricing constants to locate the changes.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: a6b67727-66ac-4ab8-9195-84e938870b32

📥 Commits

Reviewing files that changed from the base of the PR and between f3db31c and cd5bb90.

📒 Files selected for processing (8)
  • src/lib/providers/googleAiStudio.ts
  • src/lib/providers/googleNativeGemini3.ts
  • src/lib/providers/googleVertex.ts
  • src/lib/types/common.ts
  • src/lib/types/providers.ts
  • src/lib/utils/pricing.ts
  • src/lib/utils/tokenUtils.ts
  • test/continuous-test-suite-cache-breakpoints.ts
💤 Files with no reviewable changes (1)
  • src/lib/providers/googleVertex.ts
🚧 Files skipped from review as they are similar to previous changes (6)
  • src/lib/types/providers.ts
  • src/lib/utils/tokenUtils.ts
  • src/lib/types/common.ts
  • src/lib/providers/googleAiStudio.ts
  • src/lib/providers/googleNativeGemini3.ts
  • test/continuous-test-suite-cache-breakpoints.ts

Comment thread src/lib/utils/pricing.ts
Comment on lines +164 to +252
// cacheRead = 0.25x input (cached input tokens; no separate cacheCreation).
"gpt-5.4": {
input: 2.5 / 1_000_000,
output: 15.0 / 1_000_000,
cacheRead: 0.625 / 1_000_000,
},
"gpt-5.2": {
input: 1.75 / 1_000_000,
output: 14.0 / 1_000_000,
cacheRead: 0.4375 / 1_000_000,
},
"gpt-5.1": {
input: 0.625 / 1_000_000,
output: 5.0 / 1_000_000,
cacheRead: 0.15625 / 1_000_000,
},
"gpt-5.1-codex": {
input: 1.25 / 1_000_000,
output: 10.0 / 1_000_000,
cacheRead: 0.3125 / 1_000_000,
},
"gpt-5": {
input: 1.25 / 1_000_000,
output: 10.0 / 1_000_000,
cacheRead: 0.3125 / 1_000_000,
},
"gpt-5-mini": {
input: 0.25 / 1_000_000,
output: 2.0 / 1_000_000,
cacheRead: 0.0625 / 1_000_000,
},
"gpt-5-nano": {
input: 0.05 / 1_000_000,
output: 0.4 / 1_000_000,
cacheRead: 0.0125 / 1_000_000,
},
// GPT-4.1 family
"gpt-4.1": { input: 2.0 / 1_000_000, output: 8.0 / 1_000_000 },
"gpt-4.1-mini": { input: 0.4 / 1_000_000, output: 1.6 / 1_000_000 },
"gpt-4.1-nano": { input: 0.1 / 1_000_000, output: 0.4 / 1_000_000 },
"gpt-4.1": {
input: 2.0 / 1_000_000,
output: 8.0 / 1_000_000,
cacheRead: 0.5 / 1_000_000,
},
"gpt-4.1-mini": {
input: 0.4 / 1_000_000,
output: 1.6 / 1_000_000,
cacheRead: 0.1 / 1_000_000,
},
"gpt-4.1-nano": {
input: 0.1 / 1_000_000,
output: 0.4 / 1_000_000,
cacheRead: 0.025 / 1_000_000,
},
// GPT-4o family
"gpt-4o": { input: 2.5 / 1_000_000, output: 10.0 / 1_000_000 },
"gpt-4o-mini": { input: 0.15 / 1_000_000, output: 0.6 / 1_000_000 },
"gpt-4o": {
input: 2.5 / 1_000_000,
output: 10.0 / 1_000_000,
cacheRead: 0.625 / 1_000_000,
},
"gpt-4o-mini": {
input: 0.15 / 1_000_000,
output: 0.6 / 1_000_000,
cacheRead: 0.0375 / 1_000_000,
},
// o-series reasoning
o3: { input: 2.0 / 1_000_000, output: 8.0 / 1_000_000 },
"o3-mini": { input: 1.1 / 1_000_000, output: 4.4 / 1_000_000 },
"o4-mini": { input: 1.1 / 1_000_000, output: 4.4 / 1_000_000 },
o1: { input: 15.0 / 1_000_000, output: 60.0 / 1_000_000 },
"o1-mini": { input: 0.55 / 1_000_000, output: 2.2 / 1_000_000 },
o3: {
input: 2.0 / 1_000_000,
output: 8.0 / 1_000_000,
cacheRead: 0.5 / 1_000_000,
},
"o3-mini": {
input: 1.1 / 1_000_000,
output: 4.4 / 1_000_000,
cacheRead: 0.275 / 1_000_000,
},
"o4-mini": {
input: 1.1 / 1_000_000,
output: 4.4 / 1_000_000,
cacheRead: 0.275 / 1_000_000,
},
o1: {
input: 15.0 / 1_000_000,
output: 60.0 / 1_000_000,
cacheRead: 3.75 / 1_000_000,
},
"o1-mini": {
input: 0.55 / 1_000_000,
output: 2.2 / 1_000_000,
cacheRead: 0.1375 / 1_000_000,
},

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🌐 Web query:

OpenAI API pricing cached input tokens 2025 2026 gpt-4o gpt-4.1 prompt caching discount

💡 Result:

As of July 8, 2026, OpenAI provides automatic prompt caching for its API, which reduces input token costs and latency for repetitive prompt prefixes [1][2]. Key features of OpenAI prompt caching include: Automatic Implementation: Caching is handled automatically by the API without requiring manual configuration or specific code changes [1][2]. It operates by detecting repeated prompt prefixes across requests [2]. Eligibility Requirements: To be eligible for caching, prompts must generally exceed 1,024 tokens [3]. Cached tokens are identified in the API response within the usage.prompt_tokens_details field [3]. Pricing and Discounts: - Discount Rates: Savings vary by model family. For GPT-4o, the cache discount is 50% (e.g., $1.25 per 1M tokens for cached input versus $2.50 for standard input) [4]. Newer models available in 2026, such as those in the GPT-5 series, offer significantly steeper discounts, ranging from 75% to over 90% [4][2]. - Cost to Write: Unlike some other providers, OpenAI does not charge extra for writing to the cache; the only cost incurred is the standard input token rate for processing the data initially [2]. Performance Benefits: Prompt caching can reduce input token costs by up to 90% depending on the model and hit rate, and it can lower latency by up to 80% for long prompts [1]. Optimization Recommendations: To maximize savings, developers are encouraged to place static content (such as system instructions, context, or examples) at the beginning of the prompt and variable content at the end, ensuring that the cached prefix remains identical across requests [3]. For the most current and specific pricing for models like gpt-4o and newer, users should always consult the official OpenAI API pricing page, as discount structures and rates are updated periodically as inference technology evolves [4][5].

Citations:


🌐 Web query:

OpenAI official pricing cached input tokens GPT-4o GPT-4.1 GPT-5 o3 o1 2026 pricing page

💡 Result:

As of July 8, 2026, OpenAI's API pricing for cached input tokens varies by model, with newer generations generally offering deeper discounts than earlier ones [1]. While standard input tokens are billed at full price, cached input tokens—which occur when the API recognizes a prefix that matches a recently processed prompt—receive a percentage discount automatically [2][1]. Key pricing details for major models as of July 2026: GPT-4o: Input: $2.50 per 1M tokens Cached Input: $1.25 per 1M tokens (50% discount) [3][1] GPT-4.1: Input: $2.00 per 1M tokens Cached Input: $0.50 per 1M tokens (75% discount) [4][1] o1 / o3-mini: Standard rates for o1 are $15.00 input / $60.00 output per 1M tokens, with cached input typically billed at $7.50 (50% discount) [5][2]. Models like o3-mini are priced at $1.10 input / $4.40 output per 1M tokens, with cached input at $0.55 (50% discount) [5][2]. GPT-5 Series (e.g., GPT-5.5, GPT-5.4): The GPT-5 generation has introduced more aggressive caching discounts, often reaching 90% off standard input rates for specific tiers [1][6]. For example, GPT-5.5 is priced at $5.00 input / $30.00 output per 1M tokens, with cached input at $0.50 (90% discount) [7][8]. GPT-5.4 is priced at $2.50 input / $15.00 output per 1M tokens, with cached input at $0.25 (90% discount) [9][8]. Prompt caching is enabled by default for all recent models (GPT-4o and newer) and requires no manual API changes to activate [1]. Users may also utilize the Batch API or Flex Processing (service_tier="flex") to achieve additional efficiency and cost savings [1]. For the most current and specific pricing per model snapshot, refer to the official OpenAI API pricing page at https://openai.com/api/pricing/ [7][2].

Citations:


Update the cached-input rates per model
OpenAI prompt-caching pricing is model-specific now, not a blanket 0.25× multiplier. gpt-4o, o1, and o3-mini should be at 0.5× input, and gpt-5.4 is 0.1× input; several constants here are stale.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/utils/pricing.ts` around lines 164 - 252, The cached-input pricing in
the pricing table is stale and still uses the old blanket multiplier. Update the
model entries in the pricing map in pricing.ts so each affected model has the
correct cacheRead rate, especially gpt-5.4, gpt-4o, o1, and o3-mini, using the
existing model keys in the pricing constants to locate the changes.

@murdore
murdore merged commit fbf931a into release Jul 8, 2026
17 checks passed
@murdore
murdore deleted the feat/generic-prompt-caching branch July 8, 2026 03:15
@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 9.83.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

This branch was successfully deployed

1 active deployment
Preview — cd5bb905 Deployed Jul 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants