Skip to content

fix(vertex): enable prompt caching on native Claude paths - #1113

Merged
murdore merged 1 commit into
releasefrom
fix/vertex-claude-prompt-caching
Jun 23, 2026
Merged

murdore merged 1 commit into
releasefrom
fix/vertex-claude-prompt-caching

Conversation

@pdogra1299

@pdogra1299 pdogra1299 commented Jun 23, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Enables Anthropic prompt caching on the native Vertex+Claude request paths. The conversation history was never given a cache_control breakpoint, so every turn re-sent the full (growing, tool-result-heavy) prompt as fresh input — cache_read_input_tokens was 0 on every turn, including 600K+ token turns. This restores caching parity with the Claude Agent SDK on the same Vertex project.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • Performance improvement

Motivation and Context

Vertex does not support automatic prompt caching — caching only activates from explicit cache_control breakpoints in the request. The native executeNativeAnthropicGenerate / executeNativeAnthropicStream paths set none, so the conversation prefix fell after the last breakpoint and was billed at full input price every turn (cost scaled with conversation length). The cost math (pricing.ts) was already correct — this was a missing-breakpoint bug, not a mispricing.

Changes Made

  • New applyVertexAnthropicCacheBreakpoints() (src/lib/utils/anthropicCacheBreakpoints.ts) — pure helper placing ≤4 cache_control breakpoints: one on the system block (caches tools + system, since system renders after tools), plus a rolling breakpoint on the trailing history messages (resilient to Anthropic's 20-block lookback on tool-heavy turns). Falls back to the last tool when there is no system prompt.
  • Wired into both native Claude paths in googleVertex.ts, re-applied per agentic step so the stable prefix stays byte-identical while the history breakpoint rolls forward.
  • Cache metrics surfaced: both paths now extract cache_read_input_tokens / cache_creation_input_tokens and expose them as cacheReadTokens / cacheCreationTokens on result.usage, so analytics can see caching working and calculateCost prices the ~0.1x read / ~1.25x write tiers.
  • Types: cache_control added to the Vertex-Anthropic message/tool/system types (canonical src/lib/types/).
  • Tests: test/continuous-test-suite-cache-breakpoints.ts + pnpm run test:cache.

Breaking Changes

  • No breaking changes

Testing

  • Unit tests added (pnpm run test:cache — 14/14 pass, no API key)
  • Existing tests pass
  • Built the project successfully with pnpm build (0 errors, publint clean)

Verification after deploy: on consecutive main-flow turns, usage.cacheReadTokens goes from 0 → nonzero (turn 1 writes the cache, turns 2+ read at 0.1x); input_tokens/call drops toward the newest-turn size; the cache-read SKU appears in the GCP billing export.

Code Quality

  • ESLint passes / Prettier applied
  • TypeScript strict mode compliance
  • No hardcoded secrets

Commit Message Format

  • fix(vertex): enable prompt caching on native Claude paths

Deployment Notes

  • No special deployment steps (no Cloud Console / env changes — caching is request-driven)

Additional Notes

Scope here is the caching fix (Vertex+Claude). The complementary history-growth track (summarize large tool results at ingest + a history token budget) is intentionally out of scope for this PR and will follow separately; caching makes the large history cheap to re-send, summarization makes it smaller.

Summary by CodeRabbit

Release Notes

  • New Features
    • Enabled explicit prompt caching for the Vertex Anthropic provider by adding prompt-cache breakpoints to stable prefixes and recent conversation context, improving token efficiency for repeat-like requests.
  • Bug Fixes
    • Updated usage and cost tracking to include cache read and cache creation token metrics for more accurate reporting.
  • Tests
    • Added a continuous test suite validating cache-breakpoint injection, caps, edge cases, and input immutability.
    • Added a new test:cache script for running the cache-breakpoint test suite.

@vercel

vercel Bot commented Jun 23, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
neurolink Ready Ready Preview, Comment Jun 23, 2026 11:27am

@github-actions

github-actions Bot commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

✅ Single Commit Policy - COMPLIANT

Status: Policy requirements met • 1 commit • Valid format • Ready for merge

📊 View validation details

📝 Commit Details

  • Hash: 42b9287d9df00968cad6134cd149c66210e810f4
  • Message: fix(vertex): enable prompt caching on native Claude paths
  • Author: Parth Dogra

✅ Validation Results

  • Single commit requirement met
  • No merge commits in branch
  • Semantic commit message format verified
  • Ready for squash merge to release branch

🤖 Automated validation by NeuroLink Single Commit Enforcement

@coderabbitai

coderabbitai Bot commented Jun 23, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: a8d5a8bc-9be4-413b-8eb9-cac4b1b8e231

📥 Commits

Reviewing files that changed from the base of the PR and between d4e7241 and c765c0d.

📒 Files selected for processing (5)
  • package.json
  • src/lib/providers/googleVertex.ts
  • src/lib/types/providers.ts
  • src/lib/utils/anthropicCacheBreakpoints.ts
  • test/continuous-test-suite-cache-breakpoints.ts

📝 Walkthrough

Walkthrough

Adds Anthropic prompt-cache breakpoint support for Vertex+Claude calls. New types define cache_control shapes; a new utility applyVertexAnthropicCacheBreakpoints places up to four ephemeral markers across system/tools/messages; both stream and generate paths in googleVertex.ts apply these breakpoints and surface accumulated cache token counts in usage results.

Changes

Vertex Anthropic Prompt-Cache Breakpoints

Layer / File(s) Summary
Cache control type contracts
src/lib/types/providers.ts
Adds VertexAnthropicCacheControl, VertexAnthropicSystemBlock, VertexAnthropicCacheInput, and VertexAnthropicCacheOutput; augments VertexAnthropicMessage content block variants and VertexAnthropicTool with optional cache_control.
applyVertexAnthropicCacheBreakpoints utility
src/lib/utils/anthropicCacheBreakpoints.ts
Implements breakpoint allocation: one to the stable prefix (last system block or last tool), the remainder to a backward-walking message tail via markLastContentBlock, respecting MAX_BREAKPOINTS=4 and maxHistoryBreakpoints.
Stream and generate path integration
src/lib/providers/googleVertex.ts
Imports the utility; extends usage with cacheReadTokens/cacheCreationTokens; calls applyVertexAnthropicCacheBreakpoints per iteration in both executeNativeAnthropicStream and executeNativeAnthropicGenerate; accumulates and surfaces cache token totals on the returned result. Updates attachUsageAndCostAttributes to emit cache token span attributes and pass cache counts to cost calculation.
Unit tests and npm script
test/continuous-test-suite-cache-breakpoints.ts, package.json
Adds comprehensive unit tests covering system-present placement, no-system tool-only annotation, input immutability, maxHistoryBreakpoints capping, and edge cases (empty messages, skipped content, output identity guarantees). Registers the test:cache npm script to run the test suite.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • juspay/neurolink#1033: Modifies the same Anthropic Vertex execution paths to thread per-call timeouts through executeNativeAnthropicStream and executeNativeAnthropicGenerate.
  • juspay/neurolink#1084: Modifies native Anthropic paths in src/lib/providers/googleVertex.ts for model-id normalization and tool-schema dialect conversion.

Suggested reviewers

  • Tara-ag
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly describes the main change: enabling prompt caching on native Claude paths in Vertex, which is the core objective of this PR.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/vertex-claude-prompt-caching

Warning

There were issues while running some tools. Please review the errors and either fix the tool's configuration or disable the tool if it's a critical failure.

🔧 ESLint

If the error stems from missing dependencies, add them to the package.json file. For unrecoverable errors (e.g., due to private dependencies), disable the tool in the CodeRabbit configuration.

ESLint install failed: private package registry requires authentication. Disable ESLint in CodeRabbit settings or use public packages.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/providers/googleVertex.ts`:
- Around line 3403-3411: The cache usage tokens (cacheReadTokens and
cacheCreationTokens) are being added to result.usage but are being stripped out
before cost calculation. Ensure that when attachUsageAndCostAttributes() is
called, it includes the full usage object with the cache token fields intact,
and verify that calculateCost() receives these cache usage values so it can
properly calculate and include cache read and write charges in the final cost
calculation instead of treating cached content as full input rate billing.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 9abfb3db-4313-46b4-b30e-a9adae8769fc

📥 Commits

Reviewing files that changed from the base of the PR and between 781ebd8 and e4c7e72.

📒 Files selected for processing (5)
  • package.json
  • src/lib/providers/googleVertex.ts
  • src/lib/types/providers.ts
  • src/lib/utils/anthropicCacheBreakpoints.ts
  • test/continuous-test-suite-cache-breakpoints.ts

Comment thread src/lib/providers/googleVertex.ts Outdated

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Summary

Files reviewed: 5
New issues raised: 5

Issues by Severity

Severity Count Description
🔒 CRITICAL 0 None
⚠️ MAJOR 1 Cache metrics accumulation bug in streaming path
💡 MINOR 1 Shallow clone documentation
💬 SUGGESTION 3 Type safety, test coverage, naming

Blocking Issue (MAJOR)

Cache metrics incorrectly handled in streaming path (googleVertex.ts)

In executeNativeAnthropicStream, the cache metrics are being assigned directly to usage.cacheReadTokens and usage.cacheCreationTokens inside the agentic loop. While turnCacheUsage correctly accumulates across steps, the assignment pattern overwrites these values each iteration. This means only the last step's cache metrics survive, not the running total.

Fix required: Accumulate into usage using += pattern (like the generate path does), or assign final turnCacheUsage values after the loop completes.

Non-blocking Observations

  1. Type definitions are well-structured and maintain backward compatibility
  2. Test coverage is good for the main scenarios; edge cases could be expanded
  3. Documentation clearly explains the Anthropic 4-breakpoint limit and allocation strategy
  4. Code organization follows existing patterns with the new utility properly isolated

Next Steps

Please address the MAJOR issue with cache metric accumulation in the streaming path. Once fixed, this PR should be ready for merge - the implementation correctly enables prompt caching on Vertex+Claude paths as described.

turnCacheUsage.creation5m += stepCacheCreation5m;
turnCacheUsage.creation1h += stepCacheCreation1h;

usage.input += response.usage?.input_tokens || 0;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ MAJOR: Cache metrics incorrectly overwritten in streaming path

In the streaming path (executeNativeAnthropicStream), the cache metrics are being overwritten on each loop iteration instead of accumulated:

// Lines ~3400-3409 (in the loop)
if (turnCacheUsage.read > 0) {
  usage.cacheReadTokens = turnCacheUsage.read;  // Overwrites!
}
if (turnCacheUsage.creation > 0) {
  usage.cacheCreationTokens = turnCacheUsage.creation;  // Overwrites!
}

Problem: Unlike the generate path which uses += to accumulate, the streaming path assigns directly. While turnCacheUsage correctly aggregates across steps, the assignment pattern overwrites usage.cacheReadTokens each iteration, so only the last step's values survive.

Compare with generate path (correct):

totalCacheReadTokens += response.usage?.cache_read_input_tokens || 0;
totalCacheCreationTokens += response.usage?.cache_creation_input_tokens || 0;

Fix: Either accumulate directly into usage:

usage.cacheReadTokens = (usage.cacheReadTokens || 0) + stepCacheRead;
usage.cacheCreationTokens = (usage.cacheCreationTokens || 0) + stepCacheCreation;

Or move the assignment outside the loop to use final turnCacheUsage values.

return false;
}
messages[i] = {
...message,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 MINOR: Shallow clone may not preserve deep structure integrity

The markLastContentBlock function performs shallow clones of content blocks:

const content = message.content.map((block) => ({ ...block }));

While this is fine for the current type definitions (which don't have nested objects in blocks), if future content block types add nested properties (e.g., source in image/document blocks with nested metadata), this shallow clone could lead to unintended mutations of the original input.

Suggestion: Consider adding a comment documenting this assumption, or use a deep clone utility if the type definitions evolve to include deeply nested structures.

This is currently safe given the type definitions in providers.ts, but worth noting for maintainability.

* block / tool / system block to make the rendered prefix up to that point a
* cache breakpoint. Vertex has NO automatic caching, so these explicit markers
* are the only way the conversation prefix is cached across turns.
*/

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Good: Proper type definitions for cache control

The type definitions correctly extend the Anthropic message/content block types with optional cache_control fields. This maintains backward compatibility while enabling the new caching functionality.

One minor observation: The VertexAnthropicCacheControl type is defined as { type: "ephemeral" } which matches Anthropic's API specification. Consider exporting this as a const assertion if the API guarantees this is the only valid value:

export const EPHEMERAL_CACHE_CONTROL = { type: "ephemeral" } as const;
export type VertexAnthropicCacheControl = typeof EPHEMERAL_CACHE_CONTROL;

This would provide stronger type safety and prevent accidental mutation of the constant.

testNoSystemMarksLastTool();
testPurity();
testHistoryBreakpointCap();
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 SUGGESTION: Missing edge case tests

The test suite covers the main scenarios well, but consider adding tests for:

  1. Empty messages array - What happens when messages: []?
  2. Messages with empty content arrays - content: [] should skip that message
  3. Very long message history (>4 messages) - Verify only the last 3 get history breakpoints when system is present
  4. Mixed content types - Test with image/document blocks to ensure cache_control is placed correctly

Also, the purity test verifies inputs aren't mutated, but doesn't verify that the output is a new object (not a reference to input). Consider adding:

recordTest("output is new object", out.messages !== inputMessages);
recordTest("output messages are new objects", out.messages[0] !== inputMessages[0]);

Comment thread package.json
"test": "npx tsx test/continuous-test-suite.ts",
"test:client": "npx tsx test/continuous-test-suite-client.ts",
"test:context": "npx tsx test/continuous-test-suite-context.ts",
"test:cache": "npx tsx test/continuous-test-suite-cache-breakpoints.ts",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Good: Test script properly added

The new test:cache script follows the existing naming convention and correctly references the new test file. This maintains consistency with the existing test suite structure.

@pdogra1299

Copy link
Copy Markdown
Collaborator Author

Thanks for the thorough reviews — addressed in the latest commit. Per finding:

⚠️ MAJOR — cache tokens dropped before cost calc (CodeRabbit + Tara-ag)

Fixed. attachUsageAndCostAttributes() narrowed usage to input/output/total and called calculateCost() without the cache fields, so neurolink.cost billed cached content at the full input rate. It now accepts cacheReadTokens/cacheCreationTokens, sets the matching gen_ai.usage.cache_* span attributes, and passes them into calculateCost() (which already prices the ~0.1× read / ~1.25× write tiers). Verified both native call sites — stream (result?.usage) and generate (result?.usage) — flow the cache-bearing usage through.

⚠️ MAJOR — stream cache metrics "overwritten" each iteration (Tara-ag)

Clarified + relocated. To be precise: there was no data loss — turnCacheUsage.read is a cumulative running total (+= on every step), so the per-iteration usage.cacheReadTokens = turnCacheUsage.read already ended at the correct final value. But it read like an overwrite, so I moved the assignment out of the loop: it's now set once, post-loop, from the final turnCacheUsage totals (right next to the turn-level span that already reads them). Same result, unambiguous code. Note the suggested usage.cacheReadTokens = (usage.cacheReadTokens||0) + stepCacheRead would have double-counted against the existing turnCacheUsage accumulation — the post-loop single assignment avoids that.

💡 MINOR — shallow clone in markLastContentBlock (Tara-ag)

Documented. Added a comment: the shallow clone is safe because we only ever add a top-level cache_control and never mutate nested members (e.g. an image's source), which stay shared by reference but are never written. Flagged deep-clone as the path if a future block shape needs nested mutation.

💬 SUGGESTION — edge-case tests (Tara-ag)

Added (test:cache now 20 assertions): empty messages array; empty-content message skipped so the breakpoint falls through to the prior message; >4 messages capped at 3 history breakpoints with only the newest 3 marked; and output identity (new array + new message objects).

💬 SUGGESTION — EPHEMERAL_CACHE_CONTROL const assertion (Tara-ag)

Skipped, with reason. VertexAnthropicCacheControl is already a single-member literal type ({ type: "ephemeral" }), so values are constrained at the type level, and the helper uses one local EPHEMERAL const. Exporting a runtime const from the types barrel adds surface (and mildly crosses the "types-only" convention in src/lib/types/) for no extra safety. Happy to add it if you'd prefer it for discoverability.

Verification: pnpm run check → 0 errors / 4833 files · pnpm run lint → clean · pnpm run test:cache → 20/20 · pnpm run build → clean.

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review Summary

I've completed a thorough review of this PR that enables Anthropic prompt caching on native Vertex+Claude request paths.

Files Reviewed

  • package.json - Test script addition ✅
  • src/lib/providers/googleVertex.ts - Cache metrics integration ✅
  • src/lib/types/providers.ts - Cache control type definitions ✅
  • src/lib/utils/anthropicCacheBreakpoints.ts - New utility (pure, well-documented) ✅
  • test/continuous-test-suite-cache-breakpoints.ts - Comprehensive unit tests ✅

Verification of Previous Review Comments

All 6 existing review threads have been addressed:

  1. Cache metrics accumulation - Confirmed fixed: streaming path assigns cumulative values once after the agentic loop (not per-iteration)
  2. Cost path integration - Confirmed fixed: attachUsageAndCostAttributes() passes cache tokens to calculateCost()
  3. Type definitions - Properly extend Anthropic types with optional cache_control fields
  4. Test coverage - 14 tests covering system/tools breakpoints, rolling history, purity, caps, and edge cases

Key Implementation Highlights

  • Correct accumulation pattern: Generate path uses +=, streaming path accumulates in turnCacheUsage then assigns once
  • Cost calculation wired: Cache read (~0.1x) and creation (~1.25x) tiers properly passed to pricing
  • Pure utility: applyVertexAnthropicCacheBreakpoints clones inputs, never mutates
  • Budget-aware: Respects Anthropic's 4-breakpoint limit with smart allocation (system/tools + rolling history)

Compliance Check

  • ✅ No hardcoded secrets
  • ✅ No breaking changes to public SDK API
  • ✅ Follows CLAUDE.md architectural rules
  • ✅ Type-safe with proper TypeScript strict mode compliance
  • ✅ Error handling via ErrorFactory patterns
  • ✅ Observability: cache metrics exposed on spans for analytics

Approved - This is a well-implemented performance fix that restores caching parity with the Claude Agent SDK on Vertex.

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Summary

I've completed a comprehensive review of PR #1113 which enables Anthropic prompt caching on the native Vertex+Claude request paths.

Files Reviewed (5)

  1. src/lib/utils/anthropicCacheBreakpoints.ts (new - 125 lines)
  2. src/lib/types/providers.ts (modifications - 71 additions)
  3. src/lib/providers/googleVertex.ts (modifications - 105 additions, 4 deletions)
  4. test/continuous-test-suite-cache-breakpoints.ts (new - 256 lines)
  5. package.json (1 new test script)

Verification Against Project Standards

Criteria Status
Security - No hardcoded secrets ✅ Pass
CLAUDE.md Rule 1 - Dynamic imports only in registry ✅ Pass
CLAUDE.md Rule 6 - formatProviderError returns, never throws ✅ Pass
Type Safety - No any papering ✅ Pass
Backward Compatibility - Public API unchanged ✅ Pass
Cache Metrics Accumulation - Correct += pattern ✅ Pass

Key Implementation Details Verified

  1. Cache Breakpoint Logic: The applyVertexAnthropicCacheBreakpoints() function correctly implements Anthropic's 4-breakpoint limit with proper budget allocation (1 for system/tools prefix, up to 3 for rolling history).

  2. Immutability: Inputs are cloned (shallow copy sufficient for current type definitions) and never mutated - verified in both implementation and tests.

  3. Cache Metrics Flow:

    • Streaming path: turnCacheUsage accumulates via += during the agentic loop, assigned to usage once after completion
    • Generate path: totalCacheReadTokens/totalCacheCreationTokens accumulate via += across steps
    • Both paths pass cache tokens to calculateCost() for proper pricing (~0.1x read / ~1.25x write tiers)
  4. Test Coverage: 14 tests covering system+history breakpoints, no-system fallback to tools, purity, history cap, and edge cases (empty arrays, overflow, output identity).

Existing Review Comments

All blocking issues previously raised have been resolved:

  • Cache metrics overwriting issue (Tara-ag) - ✅ Fixed
  • Cost path integration (CodeRabbit) - ✅ Fixed

Non-blocking suggestions (shallow clone documentation, additional edge case tests) are minor and can be addressed in follow-ups.

Decision

APPROVE - This PR is ready for merge. The implementation correctly fixes the Vertex+Claude prompt caching regression, follows all architectural rules, maintains backward compatibility, and includes adequate test coverage.

@murdore
murdore force-pushed the fix/vertex-claude-prompt-caching branch from 42b9287 to c765c0d Compare June 23, 2026 11:26
@murdore
murdore merged commit d68c3a9 into release Jun 23, 2026
11 of 14 checks passed
@murdore
murdore deleted the fix/vertex-claude-prompt-caching branch June 23, 2026 11:26
@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 9.79.1 🎉

The release is available on:

Your semantic-release bot 📦🚀

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Complete ✅

Files reviewed: 5 (package.json, googleVertex.ts, providers.ts, anthropicCacheBreakpoints.ts, continuous-test-suite-cache-breakpoints.ts)

New issues found: 0

Assessment

This PR correctly implements prompt caching for the native Vertex+Claude paths:

  1. Cache metrics accumulation properly uses += pattern across agentic loop iterations, with final values assigned once after loop completion (lines 3694-3705)
  2. Cache breakpoints are correctly placed on system/tools (stable prefix) and rolling history messages
  3. Type safety maintained with backward-compatible optional fields
  4. No security issues - no hardcoded secrets or injection risks
  5. CLAUDE.md compliance verified - no architectural rule violations

Existing Open Comments Acknowledged

Two minor items from prior review remain open but non-blocking:

  • Thread 2 (MINOR): Shallow clone documentation - acceptable with existing comment
  • Thread 4 (SUGGESTION): Additional edge case tests - nice to have, not required

The PR is approved for merge (already merged). The cache fix will significantly reduce token costs for Vertex Claude users by enabling the ~0.1x cache-read pricing tier.

This branch was successfully deployed

1 active deployment
Preview — c765c0d7 Deployed Jun 23, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants