Skip to content

feat(llm): explicit Anthropic cache_control breakpoints on both transports - #6997

Merged
serrrfirat merged 5 commits into
mainfrom
feat/anthropic-cache-breakpoints
Aug 11, 2026
Merged

serrrfirat merged 5 commits into
mainfrom
feat/anthropic-cache-breakpoints

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

Closes #6984 — P0 #1 of the pi-harness adoption program (docs/research/pi-agent-deep-dive.md §7.3, PR #6991).

Summary

Both Anthropic transports now place explicit cache_control breakpoints instead of relying on automatic caching alone (rig/API-key path) or emitting nothing (OAuth path — previously no cache markers at all, so subscription-auth users got zero prompt caching).

OAuth transport (anthropic_oauth.rs): new apply_cache_breakpoints marks the three pi-style boundaries — system prompt block, last tool definition, last content block of the last message (including the tool_result tail common in agent loops) — each carrying the retention TTL. Retention none preserves the legacy wire shape exactly (plain-string system, no markers).

API-key transport (rig_adapter.rs + create_anthropic_from_registry):

  • keeps the request-level automatic-caching marker (moving breakpoint for the growing conversation),
  • adds an explicit marker on the last tool so the tool prefix stays cached when later content churns — rig's typed ToolDefinition can't carry cache_control, so the last tool moves into rig's raw additional_params.tools (appended after typed tools, order preserved, Anthropic-native input_schema shape),
  • short retention additionally enables rig's typed system/last-message breakpoints (CompletionModel::prompt_caching).

long retention deliberately does not enable rig's typed breakpoints: they're always plain 5m ephemeral, and per Anthropic's rules a 5m block marker beside a 1h automatic marker on the last block is a 400 (TTL conflict), and 1h entries must precede 5m entries. All markers within a request therefore share one TTL. Verified against the current prompt-caching docs (request-level + block-level compatibility, 4-breakpoint budget, longer-TTL-first ordering).

Unsupported models (claude-2 era) downgrade to none on both paths via supports_prompt_cache.

Tests (written first, red → green)

  • Wire-level capture tests in both files: a loopback server captures the exact request JSON; six tests pin the breakpoint layout for short/long/none on each transport (system block marker, last-tool-only marker with order preservation, last-message block marker incl. tool_result, TTL values, and the no-caching legacy shape).
  • Seam tests on build_rig_request for the tool move (marker + input_schema key + typed-tools untouched when caching is off).
  • Full ironclaw_llm suite: 893 passed; crate clippy -D warnings clean.

Note on the issue's second acceptance item (sustained cache reads visible in prompt_cache_activity): cache-read counts come from the live API and can't be asserted hermetically; this PR pins the emitting side. The realized hit rate also depends on the rest of the P0 program (#6985 prefix stability, #6986 tool-array stability, #6987 regression test).

Spec

crates/ironclaw_llm/CLAUDE.md gains an "Anthropic Prompt Caching" section documenting the layout and the TTL-ordering constraint.

🤖 Generated with Claude Code

…ports

Closes #6984 (P0 of the pi-harness adoption program, docs/research/
pi-agent-deep-dive.md §7.3).

The rig transport previously relied solely on Anthropic automatic
caching via a top-level cache_control field, and the OAuth transport
emitted no cache markers at all. Now both place explicit breakpoints
so the tool/system prefix and the growing conversation cache
independently:

- OAuth transport: apply_cache_breakpoints marks the system prompt
  block, the last tool definition, and the last content block of the
  last message, all carrying the retention TTL. Retention None keeps
  the legacy wire shape (plain-string system, no markers).
- rig transport: build_rig_request marks the last tool by moving it
  into rig's raw additional_params.tools (appended after typed tools,
  order preserved, Anthropic-native input_schema shape) and keeps the
  top-level automatic marker; Short retention additionally enables
  rig's typed system/last-message breakpoints. Long must not enable
  the typed breakpoints: rig markers cannot carry a TTL and a 5m
  block marker beside a 1h automatic marker is an API error.

All markers in a request share one TTL, satisfying Anthropic's
longer-TTL-first ordering rule. Unsupported models downgrade to None
via supports_prompt_cache on both paths.

Wire shape is pinned by loopback capture-server tests in both files
(three per transport: short, long, none), plus build_rig_request seam
tests for the tool move.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 1, 2026 05:05
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6997 August 1, 2026 05:05 Destroyed
@railway-app

railway-app Bot commented Aug 1, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6997 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 1, 2026 at 7:00 am

@github-actions github-actions Bot added scope: docs Documentation size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Aug 1, 2026
@coderabbitai

coderabbitai Bot commented Aug 1, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features
    • Added configurable Anthropic prompt caching for OAuth and API-key connections.
    • Supports short-term, long-term, and disabled cache retention settings.
    • Applies caching to system prompts, tools, messages, images, and tool results.
    • Automatically adjusts caching for models that do not support the selected retention mode.
    • Preserves request behavior and tool ordering when caching is disabled.

Walkthrough

Anthropic OAuth and rig requests now support configurable prompt caching. Retention markers apply to system prompts, tools, and final message blocks. Unsupported models downgrade to uncached requests. Dedicated tests cover wire formats, conversions, errors, tokens, and streaming.

Changes

Anthropic prompt-cache retention

Layer / File(s) Summary
Retention capability and provider wiring
crates/domains/ironclaw_llm/src/config.rs, crates/domains/ironclaw_llm/src/rig_adapter.rs, crates/domains/ironclaw_llm/src/anthropic_oauth.rs, crates/domains/ironclaw_llm/src/lib.rs, crates/domains/ironclaw_llm/CONTRACT.md
Retention modes now map to Anthropic markers. Unsupported models use none. Provider construction and documentation cover short, long, and disabled retention.
OAuth request cache markers
crates/domains/ironclaw_llm/src/anthropic_oauth.rs
OAuth request types support cache-control metadata. Requests mark system prompts, the final tool, and the final message block when retention is enabled.
Rig request integration
crates/domains/ironclaw_llm/src/rig_adapter.rs
Rig requests add top-level cache markers and move the final tool into Anthropic-shaped parameters when caching is enabled.
OAuth transport and conversion validation
crates/domains/ironclaw_llm/src/anthropic_oauth/tests.rs, tests/integration/coverage-exemptions.toml
The external test module covers request capture, cache serialization, conversions, error mapping, token persistence, and streaming. The test-only module receives a coverage exemption.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Provider
  participant RigAdapter
  participant OAuthTransport
  participant AnthropicAPI
  Provider->>RigAdapter: effective cache retention
  RigAdapter->>RigAdapter: build cache markers and tool parameters
  Provider->>OAuthTransport: cache retention
  OAuthTransport->>OAuthTransport: apply system, tool, and message markers
  RigAdapter->>AnthropicAPI: serialized rig request
  OAuthTransport->>AnthropicAPI: serialized OAuth request
Loading

Possibly related PRs

Suggested reviewers: benkurrek

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description explains the implementation and tests, but it omits most required template sections, including change type, test strategy, security, database, rollback, and review details. Complete the required template sections and explicitly record applicability or results for each validation, security, trust-boundary, database, blast-radius, rollback, and review item.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title uses Conventional Commits style and accurately describes explicit Anthropic cache_control breakpoints added to both transports.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_llm/src/anthropic_oauth.rs`:
- Around line 706-712: Update the system-prompt conversion around request.system
to match through a mutable borrow, such as request.system.as_mut(), instead of
calling take(). Convert only the AnthropicSystem::Text variant into cached
blocks and leave AnthropicSystem::Blocks unchanged, removing any unreachable
handling that is no longer needed.

In `@crates/ironclaw_llm/src/lib.rs`:
- Around line 505-514: Centralize model-specific retention resolution by adding
CacheRetention::for_model(&self, model: &str) -> CacheRetention, reusing the
existing prompt-cache capability check. In crates/ironclaw_llm/src/lib.rs lines
505-514 and crates/ironclaw_llm/src/anthropic_oauth.rs lines 133-139, replace
the duplicated conditional logic with this resolver before configuring prompt
caching; update RigAdapter::with_cache_retention to use the same method while
preserving its warning behavior.

In `@crates/ironclaw_llm/src/rig_adapter.rs`:
- Line 1086: Above the #[allow(clippy::too_many_arguments)] associated with
build_rig_request, add an immediately preceding // arch-exempt: comment naming
the missing request-shape aggregation and referencing the applicable plan
number. Keep the existing attribute and function behavior unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: e700c3a2-d9b7-4bdc-b94c-a7943aa3fa67

📥 Commits

Reviewing files that changed from the base of the PR and between a50ad06 and 1227a74.

📒 Files selected for processing (5)
  • crates/ironclaw_llm/CLAUDE.md
  • crates/ironclaw_llm/src/anthropic_oauth.rs
  • crates/ironclaw_llm/src/config.rs
  • crates/ironclaw_llm/src/lib.rs
  • crates/ironclaw_llm/src/rig_adapter.rs

Comment thread crates/ironclaw_llm/src/anthropic_oauth.rs Outdated
Comment thread crates/ironclaw_llm/src/lib.rs Outdated
/// only: rig's typed markers are always plain 5m ephemeral, and a 5m block
/// marker combined with a 1h automatic marker is rejected by the API (TTL
/// conflict on the last block), so `Long` relies on the automatic marker for
/// the conversation tail.
#[allow(clippy::too_many_arguments)]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

Add the required // arch-exempt: line above this #[allow].

Repo policy requires every #[allow(clippy::too_many_arguments)] to carry an immediately preceding exemption comment naming the missing aggregation and a plan link. This call site is a good candidate for a small request-shape struct, because build_rig_request now takes seven inputs including cache_retention.

♻️ Proposed annotation
+// arch-exempt: too_many_args, no RigRequestSpec aggregation for preamble/history/tools/tool_choice/sampling/cache_retention, plan `#6984`
 #[allow(clippy::too_many_arguments)]
 fn build_rig_request(

Based on learnings, #[allow(clippy::too_many_arguments)] requires an architecture-exemption comment explaining the justification, and as per coding guidelines "Do not add #[allow(clippy::too_many_arguments)] without an immediately preceding // arch-exempt: too_many_args, <specific missing aggregation>, plan #NNNN`` comment."

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
#[allow(clippy::too_many_arguments)]
// arch-exempt: too_many_args, no RigRequestSpec aggregation for preamble/history/tools/tool_choice/sampling/cache_retention, plan `#6984`
#[allow(clippy::too_many_arguments)]
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_llm/src/rig_adapter.rs` at line 1086, Above the
#[allow(clippy::too_many_arguments)] associated with build_rig_request, add an
immediately preceding // arch-exempt: comment naming the missing request-shape
aggregation and referencing the applicable plan number. Keep the existing
attribute and function behavior unchanged.

Sources: Coding guidelines, Learnings

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the ironclaw_llm Anthropic integrations to emit explicit cache_control prompt-caching breakpoints on both transports (rig/API-key and OAuth), aligning with the pi-harness caching layout and ensuring subscription-auth users also benefit from prompt caching.

Changes:

  • Add explicit Anthropic cache breakpoints: system prompt, last tool definition, and the last content block of the last message.
  • Extend the rig adapter to move the last tool into additional_params.tools (Anthropic-native shape) so it can carry cache_control, while retaining the request-level automatic caching marker.
  • Add capture-server wire-shape tests for short/long/none retention and document the caching layout in crates/ironclaw_llm/CLAUDE.md.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
crates/ironclaw_llm/src/rig_adapter.rs Emits cache markers via request root + moved last-tool entry; adds seam + wire-capture tests.
crates/ironclaw_llm/src/lib.rs Aligns rig model prompt_caching with retention and pre-downgrades unsupported models.
crates/ironclaw_llm/src/config.rs Adds CacheRetention::cache_control_json() helper for consistent marker emission.
crates/ironclaw_llm/src/anthropic_oauth.rs Adds apply_cache_breakpoints and updates request encoding to support per-block cache_control; adds wire-capture tests.
crates/ironclaw_llm/CLAUDE.md Documents Anthropic prompt caching breakpoint layout and TTL constraints.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread crates/ironclaw_llm/src/lib.rs Outdated
Comment on lines +508 to +514
let cache_retention = if config.cache_retention != CacheRetention::None
&& !rig_adapter::supports_prompt_cache(&config.model)
{
CacheRetention::None
} else {
config.cache_retention
};
Comment on lines +133 to +139
let cache_retention = if config.cache_retention != crate::config::CacheRetention::None
&& !crate::rig_adapter::supports_prompt_cache(&config.model)
{
crate::config::CacheRetention::None
} else {
config.cache_retention
};
@ironloopai

ironloopai Bot commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6997

🟢 Completed · Review submitted

1 actionable findings →

Reviewed the complete trusted base-to-head comparison. The breakpoint layouts are well covered, but cache eligibility is determined from the configured model rather than the effective request model, allowing unsupported model overrides to receive invalid cache markers.

Automatic · PR opened · attempt 1 of 3 · completed in 1m 30s

Run details
  • Repository: nearai/ironclaw
  • Base: main at a50ad06
  • Head: feat/anthropic-cache-breakpoints at 1227a74
  • Created: Aug 1, 2026, 5:10 AM UTC
  • Updated: Aug 1, 2026, 5:11 AM UTC
  • Run: 9b17c485-107a-404f-81e8-6608e205abfa
  • Latest attempt: 1 · Completed · 1d8a5a4d-8cc0-499b-9ab1-90a4cb79fcad

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #6997

⚠️ 1 finding · 1 blocking

Reviewed the complete trusted base-to-head comparison. The breakpoint layouts are well covered, but cache eligibility is determined from the configured model rather than the effective request model, allowing unsupported model overrides to receive invalid cache markers.

Findings

  1. 🟠 Medium · Revalidate cache support after selecting the effective model — crates/ironclaw_llm/src/anthropic_oauth.rs:352
    Details are attached to the relevant diff.
Validation and technical details
  • Inspected the complete diff from refs/ironloop/base (a50ad06) to refs/ironloop/head (1227a74), including all five changed files and surrounding provider/request code.
  • Traced cache retention through provider construction, OAuth complete/complete_with_tools/set_model, rig complete and streaming variants, request model overrides, tool conversion, and additional-parameter merging.
  • git diff --check refs/ironloop/base..refs/ironloop/head completed without errors.
  • Focused Rust tests could not be executed because cargo is unavailable in the review environment (/bin/bash: cargo: command not found).
  • Base: main
  • Head: feat/anthropic-cache-breakpoints at 1227a74
  • Run: 9b17c485-107a-404f-81e8-6608e205abfa

max_tokens,
temperature: req.temperature,
tools: None,
tool_choice: None,
};

apply_cache_breakpoints(&mut request, self.cache_retention);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Medium · Revalidate cache support after selecting the effective model

cache_retention is fixed when the provider is constructed, but this method first selects a per-request model override (and OAuth also supports changing active_model through set_model) and then applies the fixed retention unconditionally. Thus a provider configured for a cache-capable Claude model can later send cache_control markers to an unsupported model such as claude-2, despite the documented downgrade, causing Anthropic to reject the request. The rig path has the same issue because build_rig_request emits markers before the typed model override is injected. Determine effective retention from the selected model for every request, and add override/set-model wire tests for unsupported models.

@github-actions

github-actions Bot commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.95% (323586 / 376495 lines)
  floor:    85.11% (tolerance 0.5pp -> effective floor 84.61%)
  denominator: 376495 lines now vs 375097 at floor capture (+1398 lines, +0.37%) — not a material change

RATCHET PASS: ironclaw_runner
  observed: 85.93% (14917 / 17359 lines)
  floor:    85.55% (tolerance 0.5pp -> effective floor 85.05%)
  floor_covered_lines: 14658 (tolerance 20 lines -> effective floor 14638)
  denominator: 17359 lines now vs 17133 at floor capture (+226 lines, +1.32%) — not a material change

RATCHET PASS: ironclaw_processes
  observed: 88.76% (5889 / 6635 lines)
  floor:    88.07% (tolerance 0.5pp -> effective floor 87.57%)
  floor_covered_lines: 5839 (tolerance 20 lines -> effective floor 5819)
  denominator: 6635 lines now vs 6630 at floor capture (+5 lines, +0.08%) — not a material change

RATCHET PASS: ironclaw_turns
  observed: 88.46% (3709 / 4193 lines)
  floor:    85.11% (tolerance 0.5pp -> effective floor 84.61%)

RATCHET PASS: ironclaw_authorization
  observed: 86.59% (723 / 835 lines)
  floor:    62.51% (tolerance 0.5pp -> effective floor 62.01%)
  floor_covered_lines: 612 (tolerance 20 lines -> effective floor 592)
  denominator: 835 lines now vs 979 at floor capture (-144 lines, -14.71%) — material change (>5%)

RATCHET PASS: ironclaw_approvals
  observed: 91.05% (1820 / 1999 lines)
  floor:    85.86% (tolerance 0.5pp -> effective floor 85.36%)
  floor_covered_lines: 1822 (tolerance 20 lines -> effective floor 1802)
  denominator: 1999 lines now vs 2122 at floor capture (-123 lines, -5.8%) — material change (>5%)

RATCHET PASS: ironclaw_secrets
  observed: 85.81% (2896 / 3375 lines)
  floor:    84.01% (tolerance 0.5pp -> effective floor 83.51%)
  floor_covered_lines: 2795 (tolerance 20 lines -> effective floor 2775)
  denominator: 3375 lines now vs 3327 at floor capture (+48 lines, +1.44%) — not a material change

RATCHET PASS: ironclaw_filesystem
  observed: 77.04% (5911 / 7673 lines)
  floor:    75.93% (tolerance 0.5pp -> effective floor 75.43%)
  floor_covered_lines: 5826 (tolerance 20 lines -> effective floor 5806)
  denominator: 7673 lines now vs 7673 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_llm
  observed: 79.63% (21065 / 26455 lines)
  floor:    79.22% (tolerance 0.5pp -> effective floor 78.72%)
  floor_covered_lines: 20885 (tolerance 20 lines -> effective floor 20865)
  denominator: 26455 lines now vs 26364 at floor capture (+91 lines, +0.35%) — not a material change

RATCHET PASS: ironclaw_triggers
  observed: 94.68% (3134 / 3310 lines)
  floor:    86.04% (tolerance 0.5pp -> effective floor 85.54%)
  floor_covered_lines: 2804 (tolerance 20 lines -> effective floor 2784)
  denominator: 3310 lines now vs 3259 at floor capture (+51 lines, +1.56%) — not a material change

RATCHET PASS: ironclaw_product
  observed: 87.38% (22831 / 26129 lines)
  floor:    86.94% (tolerance 0.5pp -> effective floor 86.44%)
  floor_covered_lines: 21367 (tolerance 20 lines -> effective floor 21347)
  denominator: 26129 lines now vs 24576 at floor capture (+1553 lines, +6.32%) — material change (>5%)

RATCHET PASS: ironclaw_outbound
  observed: 94.68% (4271 / 4511 lines)
  floor:    93.49% (tolerance 0.5pp -> effective floor 92.99%)
  floor_covered_lines: 4105 (tolerance 20 lines -> effective floor 4085)
  denominator: 4511 lines now vs 4391 at floor capture (+120 lines, +2.73%) — not a material change

RATCHET PASS: ironclaw_extension_host
  observed: 84.99% (24343 / 28641 lines)
  floor:    83.82% (tolerance 0.5pp -> effective floor 83.32%)
  floor_covered_lines: 22271 (tolerance 20 lines -> effective floor 22251)
  denominator: 28641 lines now vs 26569 at floor capture (+2072 lines, +7.8%) — material change (>5%)

RATCHET PASS: ironclaw_events
  observed: 80.55% (1197 / 1486 lines)
  floor:    80.55% (tolerance 0.5pp -> effective floor 80.05%)
  floor_covered_lines: 1197 (tolerance 20 lines -> effective floor 1177)
  denominator: 1486 lines now vs 1486 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_safety
  observed: 92.75% (4468 / 4817 lines)
  floor:    92.44% (tolerance 0.5pp -> effective floor 91.94%)
  floor_covered_lines: 3973 (tolerance 20 lines -> effective floor 3953)
  denominator: 4817 lines now vs 4298 at floor capture (+519 lines, +12.08%) — material change (>5%)

RATCHET PASS: ironclaw_host_runtime
  observed: 88.41% (21338 / 24135 lines)
  floor:    88.23% (tolerance 0.5pp -> effective floor 87.73%)
  floor_covered_lines: 20538 (tolerance 20 lines -> effective floor 20518)
  denominator: 24135 lines now vs 23277 at floor capture (+858 lines, +3.69%) — not a material change

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.95% — 323586 / 376495 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_host_ingress 42.5% 17 / 40
ironclaw_memory 53.48% 630 / 1178
ironclaw_projects 72.36% 233 / 322
ironclaw_capabilities 74.59% 2876 / 3856
ironclaw_trust 75.79% 748 / 987
ironclaw_extractors 75.88% 538 / 709
ironclaw_reborn_cli 76.1% 11084 / 14566
ironclaw_observability 76.19% 32 / 42
ironclaw_filesystem 77.04% 5911 / 7673
ironclaw_wasm 78.84% 704 / 893
ironclaw_product_contracts 79.28% 2066 / 2606
ironclaw_llm 79.63% 21065 / 26455
ironclaw_events 80.55% 1197 / 1486
ironclaw_loop_contracts 82.4% 5637 / 6841
ironclaw_first_party_extensions 82.57% 6784 / 8216
ironclaw_memory_native 82.85% 2850 / 3440
ironclaw_libsql_runtime 83.3% 384 / 461
ironclaw_auth 83.95% 6699 / 7980
ironclaw_operator 84.47% 5309 / 6285
ironclaw_hooks 84.57% 9896 / 11702
ironclaw_event_projections 84.81% 854 / 1007
ironclaw_reborn_event_store 84.93% 1206 / 1420
ironclaw_extension_host 84.99% 24343 / 28641
ironclaw_reborn_config 85.29% 2110 / 2474
ironclaw_network 85.31% 894 / 1048
ironclaw_reborn_composition 85.51% 21794 / 25488
ironclaw_extension_contracts 85.64% 2546 / 2973
ironclaw_secrets 85.81% 2896 / 3375
ironclaw_runner 85.93% 14917 / 17359
ironclaw_host_api 86.05% 6367 / 7399
ironclaw_authorization 86.59% 723 / 835
ironclaw_webui 86.95% 11937 / 13729
ironclaw_wasm_limiter 87.06% 74 / 85
ironclaw_common 87.07% 1152 / 1323
ironclaw_product 87.38% 22831 / 26129
ironclaw_reborn_traces 87.61% 11720 / 13377
ironclaw_scripts 87.87% 420 / 478
ironclaw_threads 88.14% 5189 / 5887
ironclaw_host_runtime 88.41% 21338 / 24135
ironclaw_turns 88.46% 3709 / 4193
ironclaw_telegram_extension 88.55% 588 / 664
ironclaw_skills 88.58% 2784 / 3143
ironclaw_process_sandbox 88.64% 281 / 317
ironclaw_processes 88.76% 5889 / 6635
ironclaw_reborn_openai_compat 89.4% 3644 / 4076
ironclaw_telegram_v2_adapter 89.47% 1580 / 1766
ironclaw_extensions 89.55% 6249 / 6978
ironclaw_loop_host 90.47% 18043 / 19944
ironclaw_resources 90.76% 4084 / 4500
ironclaw_approvals 91.05% 1820 / 1999
ironclaw_reborn_identity 91.3% 451 / 494
ironclaw_mcp 92% 1426 / 1550
ironclaw_conversations 92.08% 2383 / 2588
ironclaw_event_streams 92.5% 1048 / 1133
ironclaw_safety 92.75% 4468 / 4817
ironclaw_agent_loop 93.52% 10430 / 11153
ironclaw_slack_extension 93.95% 3697 / 3935
ironclaw_first_party_extension_ports 94.66% 3758 / 3970
ironclaw_outbound 94.68% 4271 / 4511
ironclaw_triggers 94.68% 3134 / 3310
ironclaw_prompt_envelope 97.46% 192 / 197
ironclaw_runtime_policy 97.6% 855 / 876
ironclaw_attachments 98.23% 831 / 846

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (19 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657
crates/ironclaw_attachments/src/lib.rs Declarative crate facade: module declarations, constants, and re-exports only; executable attachment modules remain covered. #6524
crates/ironclaw_extension_host/src/ingress/mod.rs Declarative ingress module facade and documentation only; executable router modules remain covered. #6524
crates/ironclaw_host_api/src/lib.rs Declarative crate facade: module declarations and re-exports only; executable host API modules remain covered. #6524
crates/ironclaw_host_api/src/product_adapter/mod.rs Declarative product-adapter facade: module declarations and re-exports only; executable adapter modules remain covered. #6524
crates/ironclaw_llm/src/anthropic_oauth/tests.rs Test-only module stored under src/ for private transport access; cargo-llvm-cov omits test harness source from production LCOV while the exercised anthropic_oauth.rs production lines remain coverage-gated. #6984
crates/ironclaw_llm/src/rig_adapter/tests/finish_reason_tests.rs Test-only module stored under src/ for private adapter access; cargo-llvm-cov omits test harness source from production LCOV while the exercised rig_adapter.rs production lines remain coverage-gated. #6284
crates/ironclaw_loop_contracts/src/lib.rs Declaration-only public facade with no executable Rust statements; rustc emits no LCOV source record. Executable loop-contract behavior remains covered in the owned implementation modules. #6524
crates/ironclaw_outbound/src/error.rs Declarative error vocabulary only; variants have no LLVM-instrumentable production statements. #6524
crates/ironclaw_outbound/src/lib.rs Declarative crate facade: module declarations and re-exports only; executable outbound modules remain covered. #6524
crates/ironclaw_product/src/lib.rs Declaration-only public facade with no executable Rust statements; rustc emits no LCOV source record. Executable product behavior remains covered in the owned implementation modules. #6524
crates/ironclaw_product/src/lib.rs Declarative crate facade: module declarations and re-exports only; executable product modules remain covered. #6524
crates/ironclaw_product/src/scoped_fs/mod.rs Declarative scoped-filesystem facade and documentation only; executable scoped filesystem modules remain covered. #6524
crates/ironclaw_reborn_composition/src/support/fs/mod.rs Declarative composition support facade: module declarations and re-exports only; executable filesystem adapters remain covered. #6524
crates/ironclaw_slack_extension/src/lib.rs Declarative Slack crate facade: module declarations and re-exports only; executable Slack modules remain covered. #6524
crates/ironclaw_telegram_extension/src/lib.rs Declarative Telegram crate facade: module declarations and re-exports only; executable Telegram modules remain covered. #6524
crates/ironclaw_threads/src/lib.rs Declaration-only public facade with no executable Rust statements; rustc emits no LCOV source record. Executable thread behavior remains covered in the owned implementation modules. #6524
crates/ironclaw_webui/src/webui_v2/mod.rs Declaration-only WebUI v2 facade with no executable Rust statements; rustc emits no LCOV source record. Executable route behavior remains covered in the owned implementation modules. #6524

The Reborn integration-tier changed-coverage gate flagged uncovered
branches in #6997: the unsupported-model retention downgrade on both
transports, the Image/ToolUse marker arms, the empty-text guard, and
the create_anthropic_from_registry wiring.

- Extract the duplicated downgrade logic into
  rig_adapter::effective_cache_retention, shared by lib.rs and the
  OAuth constructor, with a direct unit test over all branches.
- Wire test: an unsupported model (claude-2.1) with Short retention
  keeps the legacy no-caching shape end-to-end.
- Direct apply_cache_breakpoints tests: tool_use tail without
  system/tools, image tail, empty-text tail, empty transcript.
- Construction test driving create_anthropic_from_registry across all
  retention modes including the downgrade path.
- Move the OAuth transport test suite to src/anthropic_oauth/tests.rs
  (same idiom as rig_adapter/tests/) to stay inside the file-size
  budget, with the matching coverage exemption entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings August 1, 2026 05:38
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6997 August 1, 2026 05:38 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 7 out of 7 changed files in this pull request and generated no new comments.

Suppressed comments (2)

crates/ironclaw_llm/src/anthropic_oauth.rs:434

  • Same as complete(): cache breakpoint stamping should be downgraded based on the actual request.model used after take_model_override(), otherwise a model override to a non-caching model can still emit cache_control markers and fail the request.
        apply_cache_breakpoints(&mut request, self.cache_retention);

crates/ironclaw_llm/src/anthropic_oauth.rs:347

  • apply_cache_breakpoints is always driven by self.cache_retention, which was computed from the configured model at construction. If a per-request model override selects an unsupported model (e.g. claude-2), we can still emit cache markers and risk an API 400. Consider downgrading retention at request time based on the actual request.model being sent.

This issue also appears on line 434 of the same file.

        apply_cache_breakpoints(&mut request, self.cache_retention);

The changed-coverage gate flagged three remnants: the Text arm of
set_cache_control, the empty-blocks tail, and the non-Text side of the
system take-and-rebuild — which was also a latent drop: a system value
already in block form was taken and never restored. Restore it
untouched and pin all three paths with direct tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings August 1, 2026 06:12
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6997 August 1, 2026 06:12 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_llm/src/anthropic_oauth.rs (1)

732-735: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Do not mark an empty trailing text block.

Line 733 applies cache_control to Text { text: "" }. This contradicts Lines 722-724 and sends a request that Anthropic rejects when a multimodal message ends with an empty text block. Skip empty trailing text blocks. Add a captured-wire regression through complete() or complete_with_tools().

As per path instructions, “Test through the caller: when a helper gates a side effect, require a test driving the real call site.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_llm/src/anthropic_oauth.rs` around lines 732 - 735, Update
the AnthropicContent::Blocks handling in the cache-control marker helper to
leave a trailing Text block with empty text unmarked, while preserving marking
for non-empty trailing blocks. Add a captured-wire regression that drives the
real complete() or complete_with_tools() caller and verifies the empty trailing
text block is not sent with cache_control.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_llm/src/anthropic_oauth.rs`:
- Around line 732-735: Update the AnthropicContent::Blocks handling in the
cache-control marker helper to leave a trailing Text block with empty text
unmarked, while preserving marking for non-empty trailing blocks. Add a
captured-wire regression that drives the real complete() or
complete_with_tools() caller and verifies the empty trailing text block is not
sent with cache_control.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 71043475-d160-43e8-812b-1ae24cd6303d

📥 Commits

Reviewing files that changed from the base of the PR and between 6f7047e and 33dd617.

📒 Files selected for processing (2)
  • crates/ironclaw_llm/src/anthropic_oauth.rs
  • crates/ironclaw_llm/src/anthropic_oauth/tests.rs

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 7 out of 7 changed files in this pull request and generated no new comments.

The last uncovered branch on the changed-coverage gate: Some(tools)
with an empty vec (only constructible directly — complete_with_tools
maps empty to None). Extend the empty-cases test to pin the no-op.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings August 1, 2026 06:52
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6997 August 1, 2026 06:52 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 7 out of 7 changed files in this pull request and generated no new comments.

Suppressed comments (3)

crates/ironclaw_llm/src/anthropic_oauth.rs:434

  • Same as complete(): the resolved model can differ from the construction-time model, but breakpoints are applied using self.cache_retention without re-validating against the actual model. This can emit cache_control for unsupported claude-2-era models when model_override/set_model is used.
        apply_cache_breakpoints(&mut request, self.cache_retention);

crates/ironclaw_llm/src/lib.rs:509

  • effective_cache_retention downgrades unsupported models to None, but this path no longer emits the warning that RigAdapter::with_cache_retention would have produced (because the adapter now receives the already-downgraded value). That makes a misconfigured ANTHROPIC_CACHE_RETENTION silent in production logs.
    // Downgrade retention up front for models without prompt-cache support so
    // the rig `prompt_caching` flag below agrees with the adapter's own
    // `with_cache_retention` validation.
    let cache_retention =
        rig_adapter::effective_cache_retention(config.cache_retention, &config.model);

crates/ironclaw_llm/src/anthropic_oauth.rs:347

  • apply_cache_breakpoints is driven by self.cache_retention computed at construction, but the request’s model can differ via take_model_override() / set_model(). If the active/override model is a claude-2-era model, this can still emit cache_control markers for an unsupported model (likely 400). Downgrade retention against the resolved model before applying breakpoints.

This issue also appears on line 434 of the same file.

        apply_cache_breakpoints(&mut request, self.cache_retention);

@abbyshekit

Copy link
Copy Markdown
Contributor

Reviewed alongside #7001, #6992 and #5981. One thing I'd treat as blocking, then some test-coverage notes. I've also included one concern I chased down and refuted, so nobody else spends time on it.

Blocking: retention is frozen at construction, but the model is resolved per request

AnthropicOAuthProvider::new computes cache_retention once from config.model (line ~133). But complete() resolves the model at request time:

let model = req.take_model_override().unwrap_or_else(|| self.active_model_name());   // 329-331
...
apply_cache_breakpoints(&mut request, self.cache_retention);                          // 347

self.cache_retention never re-derives from model, and set_model() (line ~481) mutates active_model behind the RwLock without touching it. The rig path has the same shape — build_rig_request stamps markers from the construction-time retention, then inject_model_override can send a different model on that same request.

Two failure modes:

  • construct on a cache-capable model, switch to claude-2.1 → cache_control emitted on a model that rejects it, 400 on every subsequent turn;
  • construct on an unsupported model, switch to a supported one → effective_cache_retention already downgraded to None, so caching stays off silently forever.

What makes this a regression rather than a pre-existing wart: before this PR the OAuth path emitted no markers at all, so a stale retention was harmless. Now it's load-bearing. Suggest evaluating effective_cache_retention against the resolved per-request model rather than snapshotting it in the constructor.

Should fix

Empty tool_result can carry a marker. convert_messages guards empty text blocks everywhere (if !msg.content.is_empty()), and apply_cache_breakpoints has an explicit guard for the AnthropicContent::Text case with the comment "Empty text blocks cannot carry cache_control (API rejects them)". But ToolResult { content: msg.content, .. } (~line 814) has no empty guard, and set_cache_control stamps whichever block is last — which is exactly the tool_result tail this PR advertises handling.

long on the API-key path gets no system-prompt breakpoint. model.prompt_caching is set only for Short, and rig's apply_cache_control — the thing that places the system marker — runs only under prompt_caching. So the summary's "both Anthropic transports now place explicit cache_control breakpoints" holds for OAuth on both retentions, but on the rig path only for short. If that asymmetry is intended, worth stating in the CLAUDE.md section, since long is where the saving is largest.

Test coverage

  • The rig short wire test sets model.prompt_caching = true by hand with a comment that it "mirrors" the factory, so it never exercises create_anthropic_from_registry. The new anthropic_registry_provider_builds_for_every_cache_retention asserts only provider.model_name(). Between them, deleting every line of cache wiring in the factory leaves the suite green.
  • oauth_long_retention_uses_1h_ttl_markers calls complete() with no tools, so the last-tool 1h branch is never hit on the OAuth path.
  • The single-tool case is untested — with exactly one tool, tools.pop() leaves req.tools empty and moves the sole definition into additional_params.tools, a distinct serialization path from the two-tool tests.
  • Several apply_cache_breakpoints unit tests assert only cache_control.is_some() without checking type/ttl.

Refuted — no action needed

I suspected the additional_params.tools move would flatten into a duplicate tools key alongside rig's typed tools field and clobber every other tool on the wire. It doesn't: rig-core 0.33 extract_tools_from_additional_params (providers/anthropic/completion.rs:1177) removes the key and appends to the typed list, so the "appended after typed tools, order preserved" claim in the description is accurate. Noting it because the seam tests assert on RigRequest rather than the serialized body, so the property isn't pinned anywhere — it just happens to hold.

Cross-PR

#7001 makes the final message a host block carrying the minute-precision clock. This PR places its third breakpoint on "the last content block of the last message". Landed together, that breakpoint sits on host boilerplate that turns over on a clock rather than on conversation content — worth checking prompt_cache_activity on the pair rather than on either alone.

@serrrfirat serrrfirat left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Add explicit Anthropic prompt-caching breakpoints across OAuth and API-key transports while preserving compatible behavior for retention modes and unsupported models.

Mode: normal — preflight classified this as an XL but contained 7-file handwritten change, with no generated/vendor or stacked-PR routing.

Coverage: diff source github; production 5, tests 1, docs 1, generated/vendor 0, CI 0, config 0; diff_truncated=true. The reviewer prompt carried the mandated 40,000-character diff prefix, while every reviewer had the exact detached head worktree for full-file and sibling-code inspection. The codebase graph was unavailable, so review used crate guidance and targeted live-code search.

Stats: 4 findings (from 6 raw, 4 after dedup) across 3 files. Reviewers run: security, bugs, performance/concurrency, tests, conventions, local patterns, maintainability, approach. Reviewers failed: none. Body-only: 0.

Security

  1. Medium Prompt cache is not partitioned by Ironclaw tenant (crates/ironclaw_llm/src/anthropic_oauth.rs:720-729, confidence 78) — anchor: crates/ironclaw_llm/src/anthropic_oauth.rs:720
    The OAuth path now caches the message tail inside the shared Anthropic workspace, but this provider has no tenant namespace. With a shared OAuth credential, an exact-prefix probe can use returned cache-read counts to infer that a guessed sensitive prompt occurred during the TTL. Anthropic documents cache sharing/isolation at workspace granularity, not at the application-tenant boundary.

Bugs

  1. Medium Revalidate cache retention against the request's actual model (crates/ironclaw_llm/src/anthropic_oauth.rs:347, confidence 98) — anchor: crates/ironclaw_llm/src/anthropic_oauth.rs:347
    Retention support is frozen from the startup model even though both request paths support per-request or mutable active-model selection. Switching model families can therefore emit unsupported markers or leave caching disabled. Also flagged by Approach/Medium.

Tests

  1. Medium Factory cache-mode wiring is not exercised on the wire (crates/ironclaw_llm/src/lib.rs:508-519, confidence 100) — anchor: crates/ironclaw_llm/src/lib.rs:519
    Wire tests manually reproduce the factory's prompt_caching assignment, while the factory test asserts only construction and model name. A regression in the production Short/Long assignment can leave all current wire tests green. Also flagged by Maintainability/Medium.

Local Patterns

  1. Medium Unsupported cache requests are now silently downgraded (crates/ironclaw_llm/src/rig_adapter.rs:923-924, confidence 98) — anchor: crates/ironclaw_llm/src/rig_adapter.rs:923
    Pre-normalizing retention to None bypasses the adapter's established construction-time warning, so an explicitly requested cache setting can be ignored silently.

max_tokens,
temperature: req.temperature,
tools: None,
tool_choice: None,
};

apply_cache_breakpoints(&mut request, self.cache_retention);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Revalidate cache retention against the request's actual model.

cache_retention is frozen from config.model at construction, but complete() and complete_with_tools() resolve the actual model per request via an override or mutable active_model. Switching from a cache-capable model to a Claude 2-era model still emits cache_control markers and can make Anthropic reject the request; switching the other direction silently leaves caching disabled. This contradicts the stated unsupported-model downgrade guarantee.

Fix: Store the configured retention unchanged, then compute effective_cache_retention(self.cache_retention, &model) after resolving the request model in both completion paths before calling apply_cache_breakpoints.

Also flagged by: approach/Medium

// retention must NOT set this — rig's markers cannot carry a TTL, and a
// 5m block marker alongside a 1h automatic marker is an API error
// (TTL conflict on the last block). See issue #6984.
model.prompt_caching = cache_retention == CacheRetention::Short;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Factory cache-mode wiring is not exercised on the wire.

The wire tests construct the rig model directly and manually set model.prompt_caching, while the only create_anthropic_from_registry test checks construction and model_name(). Removing or inverting the production factory assignment can therefore leave every current cache-layout test green even though Short loses its system/message breakpoints or Long emits conflicting 5m markers beside the 1h marker.

Fix: Add tests::anthropic_registry_factory_emits_retention_specific_breakpoints covering Short, Long, None, and an unsupported model by capturing request JSON from a provider built through create_anthropic_from_registry.

Also flagged by: maintainability/Medium

last.cache_control = Some(marker.clone());
}

if let Some(last_message) = request.messages.last_mut() {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Prompt cache is not partitioned by Ironclaw tenant.

The OAuth path now caches the last message, including user text and tool results, under the shared Anthropic workspace. Anthropic isolates caches by workspace rather than by this application's tenant, while this provider receives no actor/tenant namespace. In a deployment sharing one OAuth credential, another tenant can submit a guessed identical prefix and use returned cache-read token counts to infer that the sensitive prompt occurred during the cache TTL. OAuth requests previously emitted no cache markers, so this exposure is introduced here.

Fix: Partition cached prefixes with an opaque tenant-scoped namespace or per-tenant provider credentials; until tenant scope reaches this layer, disable message-tail caching for shared multi-tenant OAuth providers.

retention: CacheRetention,
model_name: &str,
) -> CacheRetention {
if retention != CacheRetention::None && !supports_prompt_cache(model_name) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Medium — Unsupported cache requests are now silently downgraded.

effective_cache_retention converts an explicitly requested retention to None before either factory reaches with_cache_retention. The established construction-time warning is therefore bypassed on the API-key path, and OAuth also downgrades silently, so a user's cache setting can be ignored without the diagnostic promised by the adapter's method documentation.

Fix: Emit the existing model-scoped warning when effective_cache_retention performs the downgrade, or preserve the requested retention for validation while using only the effective value for marker configuration, ensuring each transport warns exactly once.

Resolve conflicts from the crates/ family reorganization (#7037 etc.):
- anthropic_oauth.rs: keep the PR's cache-breakpoint logic alongside
  main's stream field, convert_anthropic_tools/convert_anthropic_tool_choice
  helpers, and streaming paths (now typed to AnthropicSystem).
- Move the PR's tests to crates/domains/ironclaw_llm/src/anthropic_oauth/
  and merge main's streaming tests into the suite.
- Register anthropic_oauth/tests.rs in CONTRACT.md sub-owner map and
  repoint the coverage exemption at the new crate path.
Copilot AI review requested due to automatic review settings August 11, 2026 07:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review is ineligible. To be eligible to request a review, you need a paid Copilot license, or your organization must enable Copilot code review.

@coderabbitai

coderabbitai Bot commented Aug 11, 2026

Copy link
Copy Markdown

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/domains/ironclaw_llm/src/anthropic_oauth.rs (1)

528-539: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Apply cache breakpoints on both OAuth streaming paths.

Both streaming paths build AnthropicRequest and send it without apply_cache_breakpoints. Configured caching therefore works only for buffered OAuth calls. Streaming requests emit the legacy uncached wire shape.

  • crates/domains/ironclaw_llm/src/anthropic_oauth.rs#L528-L539: apply the effective retention after request construction and before send_streaming_request.
  • crates/domains/ironclaw_llm/src/anthropic_oauth.rs#L633-L648: apply the effective retention after tool request construction and before send_streaming_request.

Add streaming caller-level wire tests for Short, Long, and None.

As per path instructions, “Test through the caller” requires tests through the streaming provider entry points.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/domains/ironclaw_llm/src/anthropic_oauth.rs` around lines 528 - 539,
Update both streaming request paths in
crates/domains/ironclaw_llm/src/anthropic_oauth.rs:528-539 and :633-648 to apply
the effective cache retention to each constructed AnthropicRequest before
send_streaming_request. Add caller-level wire tests through the streaming
provider entry points covering Short, Long, and None retention, ensuring each
path emits the configured cache breakpoints.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/domains/ironclaw_llm/src/anthropic_oauth.rs`:
- Around line 1137-1152: Update the trailing-block handling in convert_messages
so AnthropicContentBlock::ToolResult receives no cache marker when its content
is empty, while preserving existing marking behavior for non-empty blocks and
other content types. Add a regression test covering a request whose final
message is an empty tool result.

---

Outside diff comments:
In `@crates/domains/ironclaw_llm/src/anthropic_oauth.rs`:
- Around line 528-539: Update both streaming request paths in
crates/domains/ironclaw_llm/src/anthropic_oauth.rs:528-539 and :633-648 to apply
the effective cache retention to each constructed AnthropicRequest before
send_streaming_request. Add caller-level wire tests through the streaming
provider entry points covering Short, Long, and None retention, ensuring each
path emits the configured cache breakpoints.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: dc1f4e97-5169-44e4-a3a1-fa198470f130

📥 Commits

Reviewing files that changed from the base of the PR and between ed8ac27 and fd36526.

📒 Files selected for processing (7)
  • crates/domains/ironclaw_llm/CONTRACT.md
  • crates/domains/ironclaw_llm/src/anthropic_oauth.rs
  • crates/domains/ironclaw_llm/src/anthropic_oauth/tests.rs
  • crates/domains/ironclaw_llm/src/config.rs
  • crates/domains/ironclaw_llm/src/lib.rs
  • crates/domains/ironclaw_llm/src/rig_adapter.rs
  • tests/integration/coverage-exemptions.toml

Comment on lines +1137 to +1152
if let Some(last_message) = request.messages.last_mut() {
match &mut last_message.content {
// Empty text blocks cannot carry cache_control (API rejects
// them), so an empty trailing message keeps the string form.
AnthropicContent::Text(text) if !text.is_empty() => {
last_message.content =
AnthropicContent::Blocks(vec![AnthropicContentBlock::Text {
text: std::mem::take(text),
cache_control: Some(marker),
}]);
}
AnthropicContent::Text(_) => {}
AnthropicContent::Blocks(blocks) => {
if let Some(last_block) = blocks.last_mut() {
last_block.set_cache_control(marker);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Skip cache markers on empty tool_result blocks.

convert_messages preserves empty tool-result content. This branch then marks that empty tail block. Anthropic rejects cache markers on empty content blocks, so a valid empty tool result can fail the request.

Do not mark an empty AnthropicContentBlock::ToolResult. Add a regression test for an empty final tool result.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/domains/ironclaw_llm/src/anthropic_oauth.rs` around lines 1137 - 1152,
Update the trailing-block handling in convert_messages so
AnthropicContentBlock::ToolResult receives no cache marker when its content is
empty, while preserving existing marking behavior for non-empty blocks and other
content types. Add a regression test covering a request whose final message is
an empty tool result.

@serrrfirat
serrrfirat enabled auto-merge August 11, 2026 19:32

@serrrfirat serrrfirat left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Merge conflicts resolved and full CI green (clippy, tests, buckets, coverage). Approving to unblock the merge queue.

@serrrfirat
serrrfirat added this pull request to the merge queue Aug 11, 2026
Merged via the queue into main with commit aa86b50 Aug 11, 2026
44 checks passed
@serrrfirat
serrrfirat deleted the feat/anthropic-cache-breakpoints branch August 11, 2026 19:50
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…ports (nearai#6997)

* feat(llm): explicit Anthropic cache_control breakpoints on both transports

Closes nearai#6984 (P0 of the pi-harness adoption program, docs/research/
pi-agent-deep-dive.md §7.3).

The rig transport previously relied solely on Anthropic automatic
caching via a top-level cache_control field, and the OAuth transport
emitted no cache markers at all. Now both place explicit breakpoints
so the tool/system prefix and the growing conversation cache
independently:

- OAuth transport: apply_cache_breakpoints marks the system prompt
  block, the last tool definition, and the last content block of the
  last message, all carrying the retention TTL. Retention None keeps
  the legacy wire shape (plain-string system, no markers).
- rig transport: build_rig_request marks the last tool by moving it
  into rig's raw additional_params.tools (appended after typed tools,
  order preserved, Anthropic-native input_schema shape) and keeps the
  top-level automatic marker; Short retention additionally enables
  rig's typed system/last-message breakpoints. Long must not enable
  the typed breakpoints: rig markers cannot carry a TTL and a 5m
  block marker beside a 1h automatic marker is an API error.

All markers in a request share one TTL, satisfying Anthropic's
longer-TTL-first ordering rule. Unsupported models downgrade to None
via supports_prompt_cache on both paths.

Wire shape is pinned by loopback capture-server tests in both files
(three per transport: short, long, none), plus build_rig_request seam
tests for the tool move.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(llm): close changed-line coverage gaps on cache breakpoints

The Reborn integration-tier changed-coverage gate flagged uncovered
branches in nearai#6997: the unsupported-model retention downgrade on both
transports, the Image/ToolUse marker arms, the empty-text guard, and
the create_anthropic_from_registry wiring.

- Extract the duplicated downgrade logic into
  rig_adapter::effective_cache_retention, shared by lib.rs and the
  OAuth constructor, with a direct unit test over all branches.
- Wire test: an unsupported model (claude-2.1) with Short retention
  keeps the legacy no-caching shape end-to-end.
- Direct apply_cache_breakpoints tests: tool_use tail without
  system/tools, image tail, empty-text tail, empty transcript.
- Construction test driving create_anthropic_from_registry across all
  retention modes including the downgrade path.
- Move the OAuth transport test suite to src/anthropic_oauth/tests.rs
  (same idiom as rig_adapter/tests/) to stay inside the file-size
  budget, with the matching coverage exemption entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(llm): cover remaining cache-breakpoint branch arms

The changed-coverage gate flagged three remnants: the Text arm of
set_cache_control, the empty-blocks tail, and the non-Text side of the
system take-and-rebuild — which was also a latent drop: a system value
already in block form was taken and never restored. Restore it
untouched and pin all three paths with direct tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(llm): cover the empty-tools arm of apply_cache_breakpoints

The last uncovered branch on the changed-coverage gate: Some(tools)
with an empty vec (only constructible directly — complete_with_tools
maps empty to None). Extend the empty-cases test to pin the no-op.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: serrrfirat <f@nuff.tech>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cache: place explicit Anthropic cache_control breakpoints (rig adapter + OAuth transport)

4 participants