feat(caching): capture and price cached tokens across providers - #1114
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
📝 WalkthroughWalkthroughAdds cached-token tracking to Gemini 3 provider usage paths, overlapping cache extraction for OpenAI-compatible payloads, cache-read pricing rates and fallback billing, and tests covering extraction and cost accounting. ChangesCache Token Billing
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant GoogleAiStudioProvider
participant GeminiChunkCollector
participant UsageBuilder
GoogleAiStudioProvider->>GeminiChunkCollector: collect stream or generate step
GeminiChunkCollector-->>GoogleAiStudioProvider: cacheReadTokens
GoogleAiStudioProvider->>UsageBuilder: compute adjustedInputTokens and total
UsageBuilder-->>GoogleAiStudioProvider: final usage with cacheReadTokens
sequenceDiagram
participant TokenUtils
participant Pricing
participant Tests
Tests->>TokenUtils: extract prompt_tokens_details.cached_tokens
TokenUtils-->>Tests: cacheReadTokens and adjusted input
Tests->>Pricing: calculateCost(cacheReadTokens)
Pricing-->>Tests: billed cached tokens with fallback
Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Single Commit Policy - COMPLIANTStatus: Policy requirements met • 1 commit • Valid format • Ready for merge 📊 View validation details📝 Commit Details
✅ Validation Results
🤖 Automated validation by NeuroLink Single Commit Enforcement |
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
Tara-ag
left a comment
There was a problem hiding this comment.
Review Summary
Files reviewed: 8
New issues raised: 2 (both MINOR/SUGGESTION)
Blocking issues: 0
Assessment
This PR correctly implements cross-provider cached token capture and pricing. The implementation demonstrates solid engineering:
Strengths:
-
Correct handling of overlapping vs non-overlapping conventions - The
extractCachedInputTokensOverlappingfunction properly identifies OpenAI/Gemini's nestedprompt_tokens_details.cached_tokenspath and the subtraction logic inextractTokenUsageprevents double-counting. -
Safety guards - The
cached <= inputcheck prevents malformed responses from producing negative input or inflated totals. -
Graceful fallback - The
calculateCostfallback torates.inputwhencacheReadrate is missing ensures no silent undercharging for providers that report cached tokens but lack pricing. -
Comprehensive test coverage - 13 new assertions covering overlap subtraction, total conservation, malformed data guards, and pricing fallbacks.
-
Additive changes only - No breaking changes to the public SDK API; all new fields are optional.
Minor observations (non-blocking):
- Floating-point precision in pricing calculations (cosmetic, mitigated by final rounding)
- Could benefit from one additional edge case test for
cached_tokens: 0
Verification
- ✅ No hardcoded secrets
- ✅ No security vulnerabilities
- ✅ No breaking API changes
- ✅ All CLAUDE.md critical rules respected
- ✅ Tests pass (33/33)
Recommendation: Approve for merge.
| cost += (usage.output || 0) * rates.output; | ||
| if (usage.cacheReadTokens && rates.cacheRead) { | ||
| cost += usage.cacheReadTokens * rates.cacheRead; | ||
| if (usage.cacheReadTokens) { |
There was a problem hiding this comment.
💡 Minor: Floating-point precision consideration
The pricing calculations use division that could result in floating-point representation issues:
cacheRead: 0.0625 / 1_000_000, // = 6.25e-8While the final rounding in calculateCost() mitigates display issues, consider documenting that these are pre-computed rates (0.25× input) so future maintainers know to preserve the mathematical relationship when updating prices.
Alternatively, consider computing cacheRead dynamically as input * 0.25 to ensure consistency:
const geminiRates = { input: 2.0 / 1_000_000, output: 12.0 / 1_000_000 };
return {
...geminiRates,
cacheRead: geminiRates.input * 0.25, // Ensures 0.25× relationship always holds
};| @@ -247,10 +252,145 @@ function testEdgeCases(): void { | |||
| } | |||
There was a problem hiding this comment.
💡 Suggestion: Additional edge case test for zero cached tokens
Consider adding a test case for when cached_tokens: 0 is explicitly returned (vs. undefined/absent):
recordTest(
"cached_tokens: 0 treated same as absent (no cacheReadTokens field)",
extractTokenUsage({
promptTokens: 1000,
completionTokens: 200,
prompt_tokens_details: { cached_tokens: 0 },
}).cacheReadTokens === undefined
);This ensures the > 0 check in extractCachedInputTokensOverlapping behaves consistently with downstream consumers that may treat 0 differently from undefined.
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/providers/googleVertex.ts`:
- Around line 1663-1671: The usage aggregation in the Gemini stream and generate
loops is carrying the latest cache count forward instead of accumulating per
agentic step, which can misattribute cached tokens across tool-loop turns.
Update the logic around the step/usage handling in googleVertex.ts so each step
tracks its own latest prompt and cache values, clamps cache reads to that step’s
prompt, and then adds both input and cache totals into the cumulative counters
before constructing usage. Apply the same fix consistently in the affected
stream/generate paths so the final usage object reflects the correct per-step
billing split.
In `@src/lib/utils/pricing.ts`:
- Around line 739-748: Update the direct Anthropic model rate definitions so
`claude-3-sonnet` and `claude-3-haiku` include a `cacheRead` rate, matching the
existing pattern used by the Vertex Claude entries. This will ensure the
`pricing.ts` cost calculation path that uses `usage.cacheReadTokens` in the
pricing logic can apply the discounted cache-read price instead of falling back
to `rates.input` for these two models.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro
Run ID: a3f7a9ab-d0f8-4698-8d44-a72bbfa74b79
📒 Files selected for processing (8)
src/lib/providers/googleAiStudio.tssrc/lib/providers/googleNativeGemini3.tssrc/lib/providers/googleVertex.tssrc/lib/types/common.tssrc/lib/types/providers.tssrc/lib/utils/pricing.tssrc/lib/utils/tokenUtils.tstest/continuous-test-suite-cache-breakpoints.ts
7782bd6 to
f3db31c
Compare
|
Thanks — addressed in the amended commit. 🟠 Major — stale cache subtracted across agentic steps (
|
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
Tara-ag
left a comment
There was a problem hiding this comment.
Review Summary
PR #1114: feat(caching): capture and price cached tokens across providers
Files Reviewed: 8
src/lib/utils/pricing.ts- Added cacheRead rates for OpenAI and Google modelssrc/lib/utils/tokenUtils.ts- Added overlapping cache token extraction for OpenAI/DeepSeeksrc/lib/providers/googleVertex.ts- Added cachedContentTokenCount extraction for Geminisrc/lib/providers/googleAiStudio.ts- Added cachedContentTokenCount extraction for Geminisrc/lib/providers/googleNativeGemini3.ts- Added cacheReadTokens to chunk collectorssrc/lib/types/common.ts- Added prompt_tokens_details typesrc/lib/types/providers.ts- Added cacheReadTokens to CollectedChunkResulttest/continuous-test-suite-cache-breakpoints.ts- Added 13 new test assertions
Issues Found: 1 MAJOR
The direct Anthropic entries for claude-3-sonnet and claude-3-haiku are missing cacheRead and cacheCreation rates while all other Claude models have them. This means cached tokens on these models will be billed at the full input rate instead of the discounted 0.1x cache-read rate.
Existing Comments Acknowledged
- CodeRabbit's finding about missing cacheRead rates for these models is valid and should be addressed
- Tara-ag's suggestions about floating-point precision and additional edge case tests are minor and can be addressed in follow-up
Overall Assessment
The PR correctly implements the overlapping cache token convention for Gemini/OpenAI (where cached tokens are a subset of input) while preserving the non-overlapping convention for Anthropic/Claude. The token accounting logic properly subtracts cached tokens from input to avoid double-counting, and the fallback in calculateCost prevents silent undercharging for providers without cache rates.
Action Required: Add cacheRead and cacheCreation rates for claude-3-sonnet and claude-3-haiku in pricing.ts to ensure consistent cache pricing across all Claude models.
There was a problem hiding this comment.
The direct Anthropic entries for claude-3-sonnet and claude-3-haiku are missing cacheRead rates while all other Claude models have them:
"claude-3-sonnet": { input: 3.0 / 1_000_000, output: 15.0 / 1_000_000 },
"claude-3-haiku": { input: 0.25 / 1_000_000, output: 1.25 / 1_000_000 },Per the PR description, Anthropic/Vertex Claude models use the non-overlapping convention where cache_read_input_tokens is reported separately. If these models return cache read tokens, the calculateCost function will fall back to billing them at the full input rate instead of the discounted 0.1x cache-read rate.
Suggested fix: Add cacheRead rates consistent with other Claude models (0.1x input for cache reads, 1.25x for cache creation):
"claude-3-sonnet": {
input: 3.0 / 1_000_000,
output: 15.0 / 1_000_000,
cacheRead: 0.3 / 1_000_000,
cacheCreation: 3.75 / 1_000_000,
},
"claude-3-haiku": {
input: 0.25 / 1_000_000,
output: 1.25 / 1_000_000,
cacheRead: 0.025 / 1_000_000,
cacheCreation: 0.3125 / 1_000_000,
},f3db31c to
675d55a
Compare
|
Thanks @Tara-ag — addressed in the amended commit.
|
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
Tara-ag
left a comment
There was a problem hiding this comment.
Review Summary
Files Reviewed: 8 files
New Issues Raised: 0
Status: ✅ APPROVE
Analysis
This PR implements cross-provider cached token capture and pricing correctly:
Architecture:
- Properly distinguishes overlapping (Gemini/OpenAI) vs non-overlapping (Anthropic) token conventions
- Correctly subtracts cached tokens from input for overlapping providers to avoid double-counting
- Per-step cache accumulation with clamping in agentic loops (googleVertex.ts)
Pricing:
- Cache read rates added for all OpenAI and Google models (0.25× input)
- Legacy Claude 3 models (sonnet, haiku) now have cacheRead rates (0.1× input)
- Safe fallback to input rate when cacheRead rate unavailable (prevents undercharge)
Safety:
- Guard against malformed responses:
cached <= inputcheck prevents negative input - Zero cached tokens handled correctly (treated as absent)
- Total token conservation verified in tests
Testing:
- 33/33 tests passing
- Coverage includes overlap subtraction, malformed input guards, non-overlapping preservation, and pricing fallbacks
Existing Comments
- One minor suggestion from Tara-ag on floating-point precision remains open (non-blocking)
- All other review comments have been addressed in the code
The PR is ready to merge.
675d55a to
cd5bb90
Compare
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/utils/pricing.ts`:
- Around line 164-252: The cached-input pricing in the pricing table is stale
and still uses the old blanket multiplier. Update the model entries in the
pricing map in pricing.ts so each affected model has the correct cacheRead rate,
especially gpt-5.4, gpt-4o, o1, and o3-mini, using the existing model keys in
the pricing constants to locate the changes.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro
Run ID: a6b67727-66ac-4ab8-9195-84e938870b32
📒 Files selected for processing (8)
src/lib/providers/googleAiStudio.tssrc/lib/providers/googleNativeGemini3.tssrc/lib/providers/googleVertex.tssrc/lib/types/common.tssrc/lib/types/providers.tssrc/lib/utils/pricing.tssrc/lib/utils/tokenUtils.tstest/continuous-test-suite-cache-breakpoints.ts
💤 Files with no reviewable changes (1)
- src/lib/providers/googleVertex.ts
🚧 Files skipped from review as they are similar to previous changes (6)
- src/lib/types/providers.ts
- src/lib/utils/tokenUtils.ts
- src/lib/types/common.ts
- src/lib/providers/googleAiStudio.ts
- src/lib/providers/googleNativeGemini3.ts
- test/continuous-test-suite-cache-breakpoints.ts
| // cacheRead = 0.25x input (cached input tokens; no separate cacheCreation). | ||
| "gpt-5.4": { | ||
| input: 2.5 / 1_000_000, | ||
| output: 15.0 / 1_000_000, | ||
| cacheRead: 0.625 / 1_000_000, | ||
| }, | ||
| "gpt-5.2": { | ||
| input: 1.75 / 1_000_000, | ||
| output: 14.0 / 1_000_000, | ||
| cacheRead: 0.4375 / 1_000_000, | ||
| }, | ||
| "gpt-5.1": { | ||
| input: 0.625 / 1_000_000, | ||
| output: 5.0 / 1_000_000, | ||
| cacheRead: 0.15625 / 1_000_000, | ||
| }, | ||
| "gpt-5.1-codex": { | ||
| input: 1.25 / 1_000_000, | ||
| output: 10.0 / 1_000_000, | ||
| cacheRead: 0.3125 / 1_000_000, | ||
| }, | ||
| "gpt-5": { | ||
| input: 1.25 / 1_000_000, | ||
| output: 10.0 / 1_000_000, | ||
| cacheRead: 0.3125 / 1_000_000, | ||
| }, | ||
| "gpt-5-mini": { | ||
| input: 0.25 / 1_000_000, | ||
| output: 2.0 / 1_000_000, | ||
| cacheRead: 0.0625 / 1_000_000, | ||
| }, | ||
| "gpt-5-nano": { | ||
| input: 0.05 / 1_000_000, | ||
| output: 0.4 / 1_000_000, | ||
| cacheRead: 0.0125 / 1_000_000, | ||
| }, | ||
| // GPT-4.1 family | ||
| "gpt-4.1": { input: 2.0 / 1_000_000, output: 8.0 / 1_000_000 }, | ||
| "gpt-4.1-mini": { input: 0.4 / 1_000_000, output: 1.6 / 1_000_000 }, | ||
| "gpt-4.1-nano": { input: 0.1 / 1_000_000, output: 0.4 / 1_000_000 }, | ||
| "gpt-4.1": { | ||
| input: 2.0 / 1_000_000, | ||
| output: 8.0 / 1_000_000, | ||
| cacheRead: 0.5 / 1_000_000, | ||
| }, | ||
| "gpt-4.1-mini": { | ||
| input: 0.4 / 1_000_000, | ||
| output: 1.6 / 1_000_000, | ||
| cacheRead: 0.1 / 1_000_000, | ||
| }, | ||
| "gpt-4.1-nano": { | ||
| input: 0.1 / 1_000_000, | ||
| output: 0.4 / 1_000_000, | ||
| cacheRead: 0.025 / 1_000_000, | ||
| }, | ||
| // GPT-4o family | ||
| "gpt-4o": { input: 2.5 / 1_000_000, output: 10.0 / 1_000_000 }, | ||
| "gpt-4o-mini": { input: 0.15 / 1_000_000, output: 0.6 / 1_000_000 }, | ||
| "gpt-4o": { | ||
| input: 2.5 / 1_000_000, | ||
| output: 10.0 / 1_000_000, | ||
| cacheRead: 0.625 / 1_000_000, | ||
| }, | ||
| "gpt-4o-mini": { | ||
| input: 0.15 / 1_000_000, | ||
| output: 0.6 / 1_000_000, | ||
| cacheRead: 0.0375 / 1_000_000, | ||
| }, | ||
| // o-series reasoning | ||
| o3: { input: 2.0 / 1_000_000, output: 8.0 / 1_000_000 }, | ||
| "o3-mini": { input: 1.1 / 1_000_000, output: 4.4 / 1_000_000 }, | ||
| "o4-mini": { input: 1.1 / 1_000_000, output: 4.4 / 1_000_000 }, | ||
| o1: { input: 15.0 / 1_000_000, output: 60.0 / 1_000_000 }, | ||
| "o1-mini": { input: 0.55 / 1_000_000, output: 2.2 / 1_000_000 }, | ||
| o3: { | ||
| input: 2.0 / 1_000_000, | ||
| output: 8.0 / 1_000_000, | ||
| cacheRead: 0.5 / 1_000_000, | ||
| }, | ||
| "o3-mini": { | ||
| input: 1.1 / 1_000_000, | ||
| output: 4.4 / 1_000_000, | ||
| cacheRead: 0.275 / 1_000_000, | ||
| }, | ||
| "o4-mini": { | ||
| input: 1.1 / 1_000_000, | ||
| output: 4.4 / 1_000_000, | ||
| cacheRead: 0.275 / 1_000_000, | ||
| }, | ||
| o1: { | ||
| input: 15.0 / 1_000_000, | ||
| output: 60.0 / 1_000_000, | ||
| cacheRead: 3.75 / 1_000_000, | ||
| }, | ||
| "o1-mini": { | ||
| input: 0.55 / 1_000_000, | ||
| output: 2.2 / 1_000_000, | ||
| cacheRead: 0.1375 / 1_000_000, | ||
| }, |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🌐 Web query:
OpenAI API pricing cached input tokens 2025 2026 gpt-4o gpt-4.1 prompt caching discount
💡 Result:
As of July 8, 2026, OpenAI provides automatic prompt caching for its API, which reduces input token costs and latency for repetitive prompt prefixes [1][2]. Key features of OpenAI prompt caching include: Automatic Implementation: Caching is handled automatically by the API without requiring manual configuration or specific code changes [1][2]. It operates by detecting repeated prompt prefixes across requests [2]. Eligibility Requirements: To be eligible for caching, prompts must generally exceed 1,024 tokens [3]. Cached tokens are identified in the API response within the usage.prompt_tokens_details field [3]. Pricing and Discounts: - Discount Rates: Savings vary by model family. For GPT-4o, the cache discount is 50% (e.g., $1.25 per 1M tokens for cached input versus $2.50 for standard input) [4]. Newer models available in 2026, such as those in the GPT-5 series, offer significantly steeper discounts, ranging from 75% to over 90% [4][2]. - Cost to Write: Unlike some other providers, OpenAI does not charge extra for writing to the cache; the only cost incurred is the standard input token rate for processing the data initially [2]. Performance Benefits: Prompt caching can reduce input token costs by up to 90% depending on the model and hit rate, and it can lower latency by up to 80% for long prompts [1]. Optimization Recommendations: To maximize savings, developers are encouraged to place static content (such as system instructions, context, or examples) at the beginning of the prompt and variable content at the end, ensuring that the cached prefix remains identical across requests [3]. For the most current and specific pricing for models like gpt-4o and newer, users should always consult the official OpenAI API pricing page, as discount structures and rates are updated periodically as inference technology evolves [4][5].
Citations:
- 1: https://developers.openai.com/api/docs/guides/prompt-caching
- 2: https://aicostcalc.net/blog/openai-prompt-caching-when-worth-it
- 3: https://developers.openai.com/cookbook/examples/prompt_caching101
- 4: https://developers.openai.com/cookbook/examples/prompt_caching_201
- 5: https://openai.com/api/pricing/
🌐 Web query:
OpenAI official pricing cached input tokens GPT-4o GPT-4.1 GPT-5 o3 o1 2026 pricing page
💡 Result:
As of July 8, 2026, OpenAI's API pricing for cached input tokens varies by model, with newer generations generally offering deeper discounts than earlier ones [1]. While standard input tokens are billed at full price, cached input tokens—which occur when the API recognizes a prefix that matches a recently processed prompt—receive a percentage discount automatically [2][1]. Key pricing details for major models as of July 2026: GPT-4o: Input: $2.50 per 1M tokens Cached Input: $1.25 per 1M tokens (50% discount) [3][1] GPT-4.1: Input: $2.00 per 1M tokens Cached Input: $0.50 per 1M tokens (75% discount) [4][1] o1 / o3-mini: Standard rates for o1 are $15.00 input / $60.00 output per 1M tokens, with cached input typically billed at $7.50 (50% discount) [5][2]. Models like o3-mini are priced at $1.10 input / $4.40 output per 1M tokens, with cached input at $0.55 (50% discount) [5][2]. GPT-5 Series (e.g., GPT-5.5, GPT-5.4): The GPT-5 generation has introduced more aggressive caching discounts, often reaching 90% off standard input rates for specific tiers [1][6]. For example, GPT-5.5 is priced at $5.00 input / $30.00 output per 1M tokens, with cached input at $0.50 (90% discount) [7][8]. GPT-5.4 is priced at $2.50 input / $15.00 output per 1M tokens, with cached input at $0.25 (90% discount) [9][8]. Prompt caching is enabled by default for all recent models (GPT-4o and newer) and requires no manual API changes to activate [1]. Users may also utilize the Batch API or Flex Processing (service_tier="flex") to achieve additional efficiency and cost savings [1]. For the most current and specific pricing per model snapshot, refer to the official OpenAI API pricing page at https://openai.com/api/pricing/ [7][2].
Citations:
- 1: https://developers.openai.com/cookbook/examples/prompt_caching_201
- 2: https://www.promptstoproduct.com/openai-api-pricing-explained
- 3: https://developers.openai.com/api/docs/models/gpt-4o
- 4: https://developers.openai.com/api/docs/models/gpt-4.1
- 5: https://developers.openai.com/api/docs/models/o1
- 6: https://aicost.ai/tools/prompt-cache-roi
- 7: https://openai.com/api/pricing/
- 8: https://llm.okamomedia.tokyo/en/pricing/
- 9: https://developers.openai.com/api/docs/pricing
Update the cached-input rates per model
OpenAI prompt-caching pricing is model-specific now, not a blanket 0.25× multiplier. gpt-4o, o1, and o3-mini should be at 0.5× input, and gpt-5.4 is 0.1× input; several constants here are stale.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/lib/utils/pricing.ts` around lines 164 - 252, The cached-input pricing in
the pricing table is stale and still uses the old blanket multiplier. Update the
model entries in the pricing map in pricing.ts so each affected model has the
correct cacheRead rate, especially gpt-5.4, gpt-4o, o1, and o3-mini, using the
existing model keys in the pricing constants to locate the changes.
|
🎉 This PR is included in version 9.83.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
Description
Makes prompt-cache usage visible and correctly priced for every provider, not just Claude. Telemetry showed
0cache reads across all non-Claude main-flow models even though Gemini/OpenAI cache automatically. Root cause was three stacked gaps; this is a generic, additive fix.Builds on
9.79.1(fix(vertex): enable prompt caching on native Claude paths), which added Anthropiccache_controlbreakpoints. This PR is the observability + cost half, generalized across providers.Type of Change
Motivation and Context
Non-Claude models reported zero cache usage because:
tokenUtils.extractCacheReadTokensread onlycacheReadInputTokens/cacheReadTokens, missing OpenAI'sprompt_tokens_details.cached_tokensand Gemini'scachedContentTokenCount.cachedContentTokenCount(Vertex ×2 + AI Studio builtusageby hand).google/openaimodels — so even captured cached tokens couldn't be discounted.For Gemini/OpenAI, caching already happens automatically; this surfaces and prices it.
Conventions differ and getting it wrong double-counts cost:
inputbefore being surfaced ascacheReadTokens(guarded bycached <= input, total conserved).calculateCostpricescacheReadat the discounted rate where one exists, else falls back to the input rate — so any provider that emits cached tokens but has no cacheRead rate bills identically to before (no silent undercharge), while priced providers get the discount.Changes Made
tokenUtils.ts— newextractCachedInputTokensOverlapping; overlap-safe fallback inextractTokenUsage(subtract from input).googleVertex.ts(Gemini generate+stream),googleAiStudio.ts,googleNativeGemini3.ts— extractcachedContentTokenCount, surface ascacheReadTokenswith input adjustment. Vertex Claude path untouched.pricing.ts—cacheReadrates (0.25× input) foropenai+googlemodels;calculateCostinput-rate fallback. Vertex Gemini covered by the existing vertex→google fallback.types/common.ts,types/providers.ts— additive optional fields.test/continuous-test-suite-cache-breakpoints.ts— +13 assertions (overlap subtraction, total conservation, malformedcached>inputguard, non-overlapping untouched, per-provider cacheRead rates, no-undercharge fallback). 33/33 pass.Breaking Changes
Testing
pnpm run test:cache— 33/33)tsc --noEmit --strict— 0 errorspnpm run lint— cleanpnpm run build— 0 errors, publint cleanThis change was developed via a multi-agent workflow (parallel per-provider mapping → design → implement → adversarial verify); the verify pass caught and fixed an undercharge edge case (the input-rate fallback above).
Commit Message Format
feat(caching): capture and price cached tokens across providersDeployment Notes
Summary by CodeRabbit