Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
fix(anthropic): prompt-cache breakpoints + accounting parity for the direct-Anthropic path #1141
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Uh oh!
There was an error while loading. Please reload this page.
fix(anthropic): prompt-cache breakpoints + accounting parity for the direct-Anthropic path #1141
Changes from all commits
5c05322File filter
Filter by extension
Conversations
Uh oh!
There was an error while loading. Please reload this page.
Jump to
Uh oh!
There was an error while loading. Please reload this page.
There are no files selected for viewing
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
💡 The comment claims
conversationitself is never mutated, but the loop later callsconversation.push(...)twice per tool step. The breakpoint helper is pure (it clones the array), but the statement aboutconversationis misleading. Consider rephrasing to clarify that the helper’s purity is what keeps re-counting stable, e.g.:There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
💡 The new
DeferredUsagetype is a good consolidation. Just confirming:createDeferredAnalyticsis also consumed byopenaiChatCompletionsBase.ts(via this module). The base class currently resolves usage with onlypromptTokens/completionTokens/totalTokens. SinceDeferredUsagenow carries optional cache fields, consider whetheropenaiChatCompletionsBase.runStreamLoopshould forward cache values when providers populate them. Not blocking for this PR, but worth a follow-up so OpenAI-compatible cache reads aren't dropped at the resolve boundary.There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
Repository: juspay/neurolink
Length of output: 26702
🏁 Script executed:
Repository: juspay/neurolink
Length of output: 21836
🌐 Web query:
AI SDK 6 Anthropic inputTokens inputTokenDetails cacheReadTokens cacheWriteTokens inclusive exclusive semantics💡 Result:
In the Vercel AI SDK (version 6 and later), token usage reporting for Anthropic has been standardized to use the provider-agnostic
usageobject, moving away from provider-specific metadata fields [1]. Key components of theusageobject related to Anthropic prompt caching are found withinresult.usage.inputTokenDetails[2][1][3][4]: - cacheReadTokens: Represents the number of input tokens served from the cache (cache hit) [1][5][3]. - cacheWriteTokens: Represents the number of input tokens written to the cache (cache creation/miss) [1][5][3]. - noCacheTokens: Represents the number of input tokens that were not cached [3]. These values are accessible directly from theusageobject returned bygenerateTextandstreamText[1][4]. The AI SDK previously exposed cache creation tokens viaproviderMetadata.anthropic(ascacheCreationInputTokens), but this field was removed to unify the interface across all providers [1]. Regarding semantics: - TheinputTokenstotal in theusageobject generally reflects the total number of tokens processed [6]. - Cache read and cache write tokens are distinct; the first turn often involvescacheWriteTokens > 0(andcacheReadTokens = 0), while subsequent turns with a byte-stable prefix will showcacheReadTokens > 0[5]. - "Inclusive" or "exclusive" phrasing is not standard terminology for these fields in the SDK; rather, these fields are additive components that categorize the totalinputTokens(i.e.,inputTokens≈noCacheTokens+cacheReadTokens+cacheWriteTokens) [6][3]. For implementation, you set cache breakpoints usingproviderOptionsin your message content [4]: { providerOptions: { anthropic: { cacheControl: { type: 'ephemeral' } } } } [4]Citations:
🏁 Script executed:
Repository: juspay/neurolink
Length of output: 6499
🏁 Script executed:
Repository: juspay/neurolink
Length of output: 14236
🏁 Script executed:
Repository: juspay/neurolink
Length of output: 715
🏁 Script executed:
Repository: juspay/neurolink
Length of output: 715
Treat ai@6 cache tokens as overlapping with
inputTokensinputTokensis the inclusive total for the direct-AnthropicgenerateTextshape, socachedInputTokens,inputTokenDetails.cacheReadTokens, andinputTokenDetails.cacheWriteTokensneed to be removed frominputbefore pricing. As written, they’re handled like additive buckets andcalculateCostwill bill the cached/write portion twice.🤖 Prompt for AI Agents
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.