Skip to content

feat(events): durable, queryable tool-disclosure rollout metrics - #7385

Closed
serrrfirat wants to merge 1 commit into
mainfrom
feat/disclosure-metrics-events
Closed

serrrfirat wants to merge 1 commit into
mainfrom
feat/disclosure-metrics-events

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

The problem in plain terms. Progressive tool disclosure — the tool_search / tool_describe / tool_call bridge that stops us shipping 48 tool schemas to the model on every call — has been running default-on in internal production. It already computes every number an operator would need to judge that rollout: how big the catalog was, how much of it we advertised, how many tokens that saved, how often the model searched, how often a search came back empty, which result rank it eventually acted on. All of those numbers went straight into debug! shadow logs (target ironclaw::reborn::context_shadow) and process-local ModelCallDiagnostic entries. Both die with the process. So today nobody can answer "did the wide catalogs regress?" about a run that already happened.

What this changes. Every model call now leaves one typed, durable measurement in the runtime event log, and there is a read path that groups those measurements by the three dimensions #7166 §5 asks for: model, run profile, and catalog-size bucket.

  • ironclaw_loop_contracts: new disclosure_metrics vocabulary (ToolDisclosureCallMetrics, CatalogSizeBucket), a defaulted LoopCapabilityPort::tool_disclosure_metrics() accessor, and a LoopHostMilestoneKind::ModelCallMetricsRecorded milestone carrying ModelCallMetricsRecord.
  • ironclaw_loop_host: run-scoped counters on the disclosure port, incremented at the exact sites that already emit the shadow logs (no recomputation), plus emission of the metrics milestone from ThreadBackedLoopModelPort on both the success and failure paths.
  • ironclaw_event_log: RuntimeEventKind::ModelCallMetricsRecorded with a typed nested ModelCallMetrics payload, re-sanitized on every wire crossing like error_kind already is.
  • ironclaw_event_projections: a model_call_metrics read path returning per-call entries plus ModelCallMetricsAggregate totals grouped by (run profile, model, catalog bucket).

Change Type

  • New feature
  • CI/Infrastructure (architecture-test size ceiling raised, see below)

Linked Issue

Related #7166 (acceptance criteria §5, "queryable rollout metrics"). Not closing: §5 also covers the disclosure-on/off evaluation matrix, weakest-model validation, and canary evidence, which this PR does not do.

Design notes

Why the milestone → runtime event path rather than a new store. DurableLoopHostMilestoneSink already projects loop milestones into RuntimeEvents, and its docs say counter-carrying milestones deliberately stay in the milestone substrate rather than being "collapsed into lossy RuntimeEvent rows". This record is not collapsed — it is a typed nested payload with named fields — and rollout evidence is worthless if it dies with the process, which is exactly the tradeoff that comment was balancing. That reasoning is written into the projection arm so the next reader sees why this one crossed.

Why a separate event kind rather than extending ModelCompleted. ModelCompleted is a lifecycle transition and only fires on success. Rollout evidence has to include failed calls — that is where timeouts, throttling, and invalid-output loops show up. One event per model call also makes "model calls per completed task" a plain count of events under a run whose LoopCompleted is present.

Why counters are cumulative per run. A dropped or best-effort-skipped record then costs precision, not correctness. The projection compensates by taking a per-run high-water mark instead of summing, which is pinned by a test — summing cumulative counters would report one search as three the moment a run makes three model calls.

One wiring bug found and fixed along the way. The production ThreadBackedLoopModelPort is built by ThreadResolvingLoopModelGateway with no milestone sink; the ModelStarted/ModelCompleted milestones come from the outer HostManagedLoopModelPort. Reusing milestone_sink on the inner port to reach the run's sink would have double-published every lifecycle transition and corrupted the run-status projection. The metrics record therefore gets its own single-purpose model_call_metrics_sink field, threaded through ThreadResolvingLoopModelGatewayParts from RebornLoopDriverHost.

Delegation hazard. tool_disclosure_metrics() has a None default, so a decorator that forgets to delegate erases the evidence silently with no error. All eight production wrapping ports delegate explicitly, each with a comment saying why: SurfaceDisclosure*, both capability_surface_filter ports, SyntheticCapability*, SubagentSpawn*, HookGateInvocationScopePort, the hooks middleware port, and both loop_driver_host ports.

Honest coverage gaps

The issue lists "retries, timeouts". What is recorded is fallback_index (a nonzero value means the call ran on a fallback route — the observable retry signal) and failure_kind (timeout-class failures land in the unavailable / rate_limited labels). The model gateway's internal repairable-tool-output retry (model_gateway.rs, "retrying after repairable provider tool output") is not counted — threading it out would require changing HostManagedModelResponse, which every construction site would have to be updated for. Flagging rather than claiming it.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — clean
  • cargo build (via clippy/test builds)
  • Relevant tests pass: see Test Strategy
  • cargo test -p <owning-crate> --features integration — Not applicable: no database-backed or runtime-integration behavior changed; the durable log rides the existing filesystem substrate with no new table or index
  • Manual testing: none beyond the automated tiers; the durable path is exercised by the crate-tier projection contract against a real InMemoryDurableEventLog

Test Strategy

User behavior: an operator asking "how did progressive tool disclosure behave for wide catalogs on model X last week" can get an answer from durable data instead of from logs that no longer exist.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence — a new durable event kind and payload field
  • Security or permissions
  • External provider
  • Cross-component behavior — loop host → milestone sink → durable event log → projection

Tests added or updated:

  • Unit or contract:
    • ironclaw_loop_contracts (disclosure_metrics.rs): catalog buckets split on the disclosure-cap boundary (32 vs 33); bucket labels round-trip and reject unknown cohorts; schema-token reduction reports None when there was nothing to reduce.
    • ironclaw_event_log (runtime_event.rs): metrics round-trip through the durable wire; events written before this field still decode; directly-assigned model labels are sanitized on the way out.
    • ironclaw_turn_runner (milestone_events.rs): the milestone → RuntimeEvent projection preserves every query dimension; a failed model call records its failure classification and distinguishes "no disclosure" from zeroed counters.
    • ironclaw_event_projections (new model_call_metrics_projection_contract.rs): totals group by model/run-profile/bucket; cumulative counters are not multiplied across a run; model_calls_per_completed_run stays None until a run actually completes; a payload-less metrics event is skipped rather than folded in as a zero-latency call.
  • Reborn integration: two new tests in tests/integration/tool_disclosure.rs, driven through RebornIntegrationHarness with the real ToolDisclosureCapabilityDecorator wiring:
    • deferred_bridge_flow_records_disclosure_metrics_for_every_model_call — one record per model call, disclosure numbers on every call, full > advertised for both tool count and schema tokens, Wide bucket, and the exact search/promotion/selected-rank counters.
    • empty_search_and_outside_surface_attempts_are_counted_separately — a no-match search and a call to a nonexistent tool are counted as distinct signals and stay recoverable.
    • Assertion goes through a named helper (assert_model_call_metrics_recorded_since) at the milestone seam, which is precisely what DurableLoopHostMilestoneSink projects into the durable log.
  • Recorded fixture: Not applicable: no provider wire shape changed.
  • Browser E2E: Not applicable: no UI surface.
  • Backend or runtime: Not applicable: no schema, migration, or backend-specific behavior; the projection contract runs against the real DurableEventLog trait.
  • Live canary: Not applicable to this PR: exact-head canary evidence is a separate §5 acceptance item covering the disclosure-on/off comparison, which this PR does not attempt.

What the tests prove: that the production loop-host path emits one measurement per model call with the disclosure numbers intact, that those measurements survive the durable wire without losing a field or leaking an unsanitized label, that pre-existing history still decodes, and that the read path answers the three grouping questions the issue names without inflating cumulative counters.

Commands run:

cargo fmt
cargo clippy --all --benches --tests --examples --all-features -- -D warnings          # clean
cargo test -p ironclaw_event_log -p ironclaw_event_projections \
           -p ironclaw_event_streams -p ironclaw_turn_runner --lib --tests             # pass
cargo test -p ironclaw_loop_host                                                       # pass
cargo test -p ironclaw_integration_tests --test reborn_integration_tool_disclosure     # 32 passed
cargo test -p ironclaw_architecture_tests                                              # pass

One pre-existing, unrelated failure: ironclaw_turn_runner's trace_capture::tests::capture_skips_when_policy_missing_or_disabled fails identically on a clean origin/main checkout in this environment (verified by stashing the branch and re-running). Not touched by this PR.

Security Impact

The durable event log is redaction-bound, so the payload is numbers and closed-vocabulary labels only — no tool names, search queries, schemas, or descriptions cross into it.

Model identity labels needed a wider vocabulary than the existing lower_snake_case telemetry guard (anthropic/claude-opus-4.5 would have been collapsed to unclassified), so a new sanitize_model_label allows ASCII alphanumerics plus _ - . : / and nothing else, bounded at 128 bytes. Whitespace, quotes, control characters, and anything non-ASCII are still rejected, which is what stops a label built from untrusted text carrying prose, a path fragment, or a token-shaped secret. Like error_kind, the guard runs on every wire crossing — serialize, deserialize, and the typed constructor — so a direct pub field assignment cannot bypass it. Pinned by a test.

Metrics emission is best-effort: a failed publish is logged at debug and swallowed, so telemetry can never end a run that otherwise succeeded.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: none added. ModelCallMetrics carries no authority and is never read back as an input to any decision.
  • Untrusted content enters prompts only through an envelope/escaping primitive: unchanged; nothing here touches prompt construction.
  • Hashes: none added.
  • New/changed status, exit, policy, runtime, or error variants: RuntimeEventKind::ModelCallMetricsRecorded added. All match sites audited — cargo check --workspace --all-features --tests surfaces every non-wildcard match, and the four that exist were updated deliberately: two in runtime_projection.rs (the metrics event must never move capability-activity or run status — it observes the run, it does not advance it), TimelineEntryKind::from, and ironclaw_hooks::is_lifecycle_kind (non-lifecycle; it carries no hook identity).
  • Security/durability serde(default) fields fail closed: model_call_metrics is Option, #[serde(default, skip_serializing_if = "Option::is_none")]. Absent decodes as "no metrics", which the projection treats as "skip", never as a zero-valued call. Migration test included.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: disclosure counters use fetch_update with saturating_add (a wrapped counter would read lower and understate a pathological run); the aggregate uses saturating_add throughout; the projection inherits MAX_PROJECTION_PAGE_LIMIT and the existing STATE_REPLAY_MAX_EVENTS rebase guard.
  • Driver/operator-visible errors have stable class semantics: the projection's default model_call_metrics returns ProjectionError::InvalidRequest rather than an empty page, so a service that cannot answer says so — "no metrics" and "zero model calls" are conclusions an operator would act on very differently.
  • Sandbox/native/host names accurately describe trust boundary: N/A.

Database Impact

None. The durable event log stores RuntimeEvent as an append-only JSON blob through the filesystem substrate (FilesystemDurableEventLog over ScopedFilesystem::append), so this needs no migration and no schema change on either PostgreSQL or libSQL.

Compatibility (durable event schema):

  • Backward — old rows have no model_call_metrics key; the field is #[serde(default)] on both RuntimeEventWire and TrustedRuntimeEventWire, so pre-existing history decodes unchanged. Pinned by runtime_events_written_before_model_call_metrics_still_decode.
  • Forward — a new row carries an extra key plus a new kind value. Neither wire struct uses deny_unknown_fields, so an older binary reading a new row ignores the extra field. It will fail to deserialize the new kind enum value; during a mixed-version window an older reader will error on new metrics rows rather than skip them. Mitigation: this event is additive telemetry that no product read path consumes, so a rollback is a pure code revert with no data migration and no data loss — the rows simply stop being written and any already written become unreadable-but-harmless log entries.
  • No existing event data semantics changed, deleted, or mutated. Every existing kind, field, and sanitizer behaves exactly as before.

Blast Radius

Touches ironclaw_loop_contracts, ironclaw_loop_host, ironclaw_turn_runner, ironclaw_hooks, ironclaw_event_log, ironclaw_event_projections, and their tests. Risk concentrates in two places:

  1. The new model-port sink. If it had reused milestone_sink it would have double-published lifecycle milestones; it does not, and the run-status projection tests plus the full disclosure integration suite pass unchanged.
  2. The counters on the hot disclosure path. They are relaxed atomics on an already-Arc'd port, taken outside the turn-state lock, so they cannot contend with or deadlock against catalog construction. Counters deliberately live outside ToolDisclosureTurnState because that state is rebuilt on a mid-turn surface-fingerprint change, which would have silently reset them.

The ironclaw_loop_contracts architecture size ceiling was raised 14,479 → 14,530 with the rationale in-place at the ceiling table: the growth is declarations only (the metrics DTO, the bucket enum, the milestone record); the counters are computed in ironclaw_loop_host and projected in ironclaw_turn_runner / ironclaw_event_projections.

Rollback Plan

Revert the commit. No migration, no data backfill, no config change. The event kind stops being written; existing rows are inert (see the forward-compatibility note above). REBORN_TOOL_DISCLOSURE=off is unaffected and remains the disclosure rollback.

Review Follow-Through

Reviewer judgment would help most on:

  1. Crossing the "counters stay in the milestone substrate" line. milestone_events.rs documents a deliberate rule that counter-carrying milestones should not become RuntimeEvent rows. I argue the rule was about lossy collapse rather than about durability, and that rollout evidence has to outlive the process — but that is a judgment call on someone else's stated design, and I would rather have it challenged than assumed.
  2. Cumulative vs per-call delta counters. Cumulative degrades gracefully under dropped records; per-call deltas would make the projection trivial but lose data silently. The high-water aggregation exists to compensate and is tested, but it is the subtlest part of the read path.
  3. The sanitize_model_label vocabulary. Allowing / and . is required for real provider model ids; if the durable-log redaction bar wants something narrower (a fixed allowlist of known model ids, say), this is where to say so.
  4. Whether the inspector store should also carry these numbers. It is process-local, so it adds no durability, and I left it alone to keep the diff scoped. Easy follow-up if the live inspector UI wants them.

Review track: C (durable event schema)

🤖 Generated with Claude Code

Progressive tool disclosure has been running default-on in internal
production while its measurements existed only as ephemeral
`debug!` shadow logs and process-local inspector entries. Nobody could
answer "did the wide catalogs regress?" after the fact.

Record one typed durable measurement per model call through the existing
milestone -> durable runtime event -> projection path, queryable by
model, run profile, and catalog-size bucket.

- `ironclaw_loop_contracts`: `ToolDisclosureCallMetrics` + `CatalogSizeBucket`
  vocabulary, a defaulted `LoopCapabilityPort::tool_disclosure_metrics`
  accessor, and the `ModelCallMetricsRecorded` milestone.
- `ironclaw_loop_host`: run-scoped disclosure counters on the disclosure
  port (searches, empty searches, selected rank, promotions, recoveries,
  outside-surface attempts) read at the point the shadow logs already
  computed them, plus a dedicated metrics milestone sink on the model port
  so the record reaches the durable log without double-publishing the
  lifecycle milestones the outer port owns.
- `ironclaw_event_log`: `RuntimeEventKind::ModelCallMetricsRecorded` with a
  typed, redaction-guarded `ModelCallMetrics` payload and a model-identity
  label sanitizer.
- `ironclaw_event_projections`: `model_call_metrics` read path returning
  per-call entries plus totals grouped by the three query dimensions.

Refs #7166

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7385 August 7, 2026 23:10 Destroyed
@railway-app

railway-app Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7385 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 7, 2026 at 11:19 pm

@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added durable model-call metrics, including outcomes, latency, token usage, fallback details, and tool-disclosure activity.
    • Added reporting for catalog size, searches, recoveries, promotions, rankings, schema-token reductions, and surface attempts.
    • Added projections and aggregation for reviewing metrics by model and catalog size.
    • Added sanitization and backward-compatible event handling for recorded metrics.
  • Bug Fixes

    • Preserved disclosure metrics through capability wrappers and model-routing paths.
    • Ensured metrics events do not alter lifecycle or run-status projections.

Walkthrough

Adds model-call and progressive-tool-disclosure metrics across loop contracts, runtime capture, durable events, replay projections, and integration tests. Metrics include model routing, outcomes, token usage, disclosure counters, catalog buckets, and schema-token reduction.

Changes

Model-call disclosure telemetry

Layer / File(s) Summary
Metrics contracts and milestones
crates/contracts/ironclaw_loop_contracts/src/disclosure_metrics.rs, crates/contracts/ironclaw_loop_contracts/src/host/capability.rs, crates/contracts/ironclaw_loop_contracts/src/milestones.rs, crates/contracts/ironclaw_loop_contracts/src/lib.rs
Defines catalog-size buckets, per-call disclosure metrics, model-call records, and the ModelCallMetricsRecorded milestone.
Runtime metrics capture
crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs, crates/loop/ironclaw_loop_host/src/lib.rs, crates/loop/ironclaw_loop_host/src/thread_resolving_model_gateway.rs, crates/loop/ironclaw_hooks/..., crates/loop/ironclaw_turn_runner/...
Collects disclosure counters, propagates metrics through capability wrappers, and emits best-effort model-call milestones for resolved and fallback gateways.
Durable event representation
crates/events/ironclaw_event_log/src/runtime_event.rs, crates/events/ironclaw_event_log/src/lib.rs, crates/loop/ironclaw_turn_runner/src/milestone_events.rs, crates/events/ironclaw_event_log/tests/durable_log_contract.rs
Adds serialized model-call metrics events, label sanitization, backward-compatible omission defaults, and milestone-to-event conversion.
Metrics replay and aggregation
crates/events/ironclaw_event_projections/src/..., crates/events/ironclaw_event_projections/tests/...
Projects metrics events into grouped aggregates by requested model, effective model, and catalog bucket. Per-run disclosure counters use high-water marks.
End-to-end telemetry validation
tests/integration/support/assertions.rs, tests/integration/tool_disclosure.rs
Validates durable records for deferred bridge flows, empty searches, outside-surface attempts, ranking, promotion, recovery, and schema-token reduction.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant RebornLoopDriverHost
  participant ThreadResolvingLoopModelGateway
  participant ThreadBackedLoopModelPort
  participant ToolDisclosurePort
  participant LoopHostMilestoneSink
  participant RuntimeEvent
  participant EventProjectionService
  RebornLoopDriverHost->>ThreadResolvingLoopModelGateway: attach metrics sink
  ThreadResolvingLoopModelGateway->>ThreadBackedLoopModelPort: configure sink
  ThreadBackedLoopModelPort->>ToolDisclosurePort: collect disclosure metrics
  ThreadBackedLoopModelPort->>LoopHostMilestoneSink: publish model-call record
  LoopHostMilestoneSink->>RuntimeEvent: write ModelCallMetricsRecorded
  EventProjectionService->>RuntimeEvent: read scoped metrics page
  EventProjectionService->>EventProjectionService: aggregate by model and catalog bucket
Loading

Possibly related PRs

Suggested reviewers: benkurrek

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title uses Conventional Commits syntax and accurately describes the durable rollout metrics feature.
Description check ✅ Passed The description covers the required sections, linked issue, validation, tests, security, compatibility, blast radius, rollback, and review follow-through.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
⚔️ Resolve merge conflicts 💡
  • Resolve merge conflict in branch feat/disclosure-metrics-events

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Aug 7, 2026
@ironloopai

ironloopai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

🧭 IronLoop Run · Review

This comment updates in place as the Run moves through its stages.

🟥 Final result · Could not complete

🟨 Queued → 🟦 Working → 🟥 Could not complete

Automatic trigger · attempt 1 of 3 · failed after 8s

IronLoop could not complete the review for this Run.

Failure details

  • Failure reference: ccb6a62d-e766-48fa-99c8-fd5d8698292e
Run details

Run: e67111e1-3780-44f2-b72a-59206a0cde2c
Base: main at 254483d
Head: feat/disclosure-metrics-events at b8f9fdf
Created: 2026-08-07 23:11 UTC
Updated: 2026-08-07 23:11 UTC

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 8

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/contracts/ironclaw_loop_contracts/src/host/capability.rs`:
- Around line 552-561: Update tool_disclosure_metrics and its decorators so
metrics describe the final capability surface: transparent decorators must
delegate unchanged, while capability_surface_filter.rs lines 78-85 and 184-191
must preserve inner full-surface values and derive advertised values after
filtering and policy resolution; subagent_spawn_port.rs lines 1161-1168 must
include spawn_subagent in full and advertised measurements. Add caller-level
tests through the built model host covering filtered and subagent-decorated
surfaces and asserting the durable metrics payload.

In `@crates/events/ironclaw_event_projections/src/model_call_metrics.rs`:
- Around line 282-289: Re-sanitize metrics from event.model_call_metrics before
cloning them into ModelCallMetricsEntry, applying the existing event-log
sanitizers for model, failure, and catalog labels so custom DurableEventLog data
is treated as untrusted. Add a regression test using a custom backend that
supplies a path- or secret-shaped model label and verify the projected entry
contains the sanitized value.
- Around line 215-231: Update the disclosure aggregation around
per_run_disclosure to retain prior cumulative counter values by InvocationId
across the ordered page, compute each call’s nonnegative counter delta, and
attribute only that delta to the current model/catalog group instead of
re-adding the run’s high-water totals. Preserve high-water tracking while
preventing duplicate attribution when a run changes groups, and add a regression
test covering a single run switching model or bucket between calls.
- Around line 255-266: The model_calls_per_completed_run calculation currently
includes calls from in-flight runs; track model-call counts by run ID and
separately accumulate calls belonging to completed runs, then use that
completed-call count as the ratio numerator while preserving None when
completed_runs is zero. Add coverage for a group containing one completed run
and one in-flight run, verifying only the completed run’s calls are included.

In `@crates/events/ironclaw_event_projections/src/runtime_projection.rs`:
- Around line 406-409: The status-preservation behavior for
RuntimeEventKind::ModelCallMetricsRecorded lacks caller-level coverage. Add
tests for ReplayEventProjectionService::snapshot that verify metrics after
LoopCompleted leave the run Completed, while a metrics-only run remains Running;
cover both relevant projection paths without changing production behavior.

In `@crates/loop/ironclaw_hooks/src/dispatch/lifecycle_owner.rs`:
- Around line 85-88: Add RuntimeEventKind::ModelCallMetricsRecorded to the
non_lifecycle array in is_lifecycle_kind_classifies_every_variant, keeping the
exhaustive classification test aligned with the match in lifecycle
classification.

In `@crates/loop/ironclaw_loop_host/src/lib.rs`:
- Around line 1715-1764: Add caller-level coverage in
thread_loop_host_contract.rs for the stream_model path configured via
with_model_call_metrics_sink, asserting a ModelCallMetricsRecorded milestone is
emitted with the expected data. Also configure a failing milestone sink and
verify emit_model_call_metrics swallows the publish error without failing the
loop operation.

In `@crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs`:
- Around line 269-299: Add a caller-level unit test in the SpyPort suite that
invokes the counter-recording methods record_search, record_promotion,
record_recovery, record_outside_surface_attempt, and record_selected_rank, then
calls tool_disclosure_metrics() and asserts tool_search_count,
empty_search_count, promotions, recoveries, selected_result_rank, and
outside_surface_attempts. Keep the assertion at this seam to verify the durable
DisclosureRunCounters telemetry path.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: bcadc261-881f-419e-b404-c9eb5560b660

📥 Commits

Reviewing files that changed from the base of the PR and between 254483d and b8f9fdf.

📒 Files selected for processing (27)
  • crates/app/ironclaw_architecture_tests/tests/reborn_dependency_boundaries.rs
  • crates/contracts/ironclaw_loop_contracts/src/disclosure_metrics.rs
  • crates/contracts/ironclaw_loop_contracts/src/host/capability.rs
  • crates/contracts/ironclaw_loop_contracts/src/lib.rs
  • crates/contracts/ironclaw_loop_contracts/src/milestones.rs
  • crates/events/ironclaw_event_log/src/lib.rs
  • crates/events/ironclaw_event_log/src/runtime_event.rs
  • crates/events/ironclaw_event_log/tests/durable_log_contract.rs
  • crates/events/ironclaw_event_projections/src/lib.rs
  • crates/events/ironclaw_event_projections/src/model_call_metrics.rs
  • crates/events/ironclaw_event_projections/src/runtime_projection.rs
  • crates/events/ironclaw_event_projections/tests/model_call_metrics_projection_contract.rs
  • crates/events/ironclaw_event_projections/tests/replay_projection_contract.rs
  • crates/loop/ironclaw_hooks/src/dispatch/lifecycle_owner.rs
  • crates/loop/ironclaw_hooks/src/middleware/capability_port.rs
  • crates/loop/ironclaw_loop_host/src/capability_surface_filter.rs
  • crates/loop/ironclaw_loop_host/src/lib.rs
  • crates/loop/ironclaw_loop_host/src/subagent_spawn_port.rs
  • crates/loop/ironclaw_loop_host/src/surface_disclosure.rs
  • crates/loop/ironclaw_loop_host/src/synthetic_capability.rs
  • crates/loop/ironclaw_loop_host/src/thread_resolving_model_gateway.rs
  • crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs
  • crates/loop/ironclaw_turn_runner/src/hook_gate_refs.rs
  • crates/loop/ironclaw_turn_runner/src/loop_driver_host.rs
  • crates/loop/ironclaw_turn_runner/src/milestone_events.rs
  • tests/integration/support/assertions.rs
  • tests/integration/tool_disclosure.rs

Comment on lines +552 to +561
/// Progressive-tool-disclosure measurements for the surface this port is
/// currently presenting, or `None` when disclosure is not in play.
///
/// Read-only observation of numbers the port already computed — never a
/// recomputation and never an authority. Any port that wraps another MUST
/// delegate this: a decorator that silently keeps the default `None`
/// erases the rollout evidence for every run that goes through it, and
/// does so without any error to notice.
fn tool_disclosure_metrics(&self) -> Option<ToolDisclosureCallMetrics> {
None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Record metrics for the final capability surface.

The contract requires unchanged delegation, but these decorators change the represented surface. The durable record can report incorrect advertised counts and schema tokens. The policy and subagent decorators can also report an incorrect full count and catalog bucket.

  • crates/contracts/ironclaw_loop_contracts/src/host/capability.rs#L552-L561: require transparent decorators to delegate, but require surface-transforming decorators to produce metrics for their transformed surface.
  • crates/loop/ironclaw_loop_host/src/capability_surface_filter.rs#L78-L85: preserve the inner full-surface values, but derive advertised values from the model-visible filtered surface.
  • crates/loop/ironclaw_loop_host/src/capability_surface_filter.rs#L184-L191: derive metrics after the policy-resolved filter applies.
  • crates/loop/ironclaw_loop_host/src/subagent_spawn_port.rs#L1161-L1168: include spawn_subagent in the relevant full and advertised measurements.

Add a caller-level test through the built model host. Test a filtered surface and a subagent-decorated surface. Assert the durable metrics payload. This follows the Test through the caller invariant.

📍 Affects 3 files
  • crates/contracts/ironclaw_loop_contracts/src/host/capability.rs#L552-L561 (this comment)
  • crates/loop/ironclaw_loop_host/src/capability_surface_filter.rs#L78-L85
  • crates/loop/ironclaw_loop_host/src/capability_surface_filter.rs#L184-L191
  • crates/loop/ironclaw_loop_host/src/subagent_spawn_port.rs#L1161-L1168
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/contracts/ironclaw_loop_contracts/src/host/capability.rs` around lines
552 - 561, Update tool_disclosure_metrics and its decorators so metrics describe
the final capability surface: transparent decorators must delegate unchanged,
while capability_surface_filter.rs lines 78-85 and 184-191 must preserve inner
full-surface values and derive advertised values after filtering and policy
resolution; subagent_spawn_port.rs lines 1161-1168 must include spawn_subagent
in full and advertised measurements. Add caller-level tests through the built
model host covering filtered and subagent-decorated surfaces and asserting the
durable metrics payload.

Sources: Coding guidelines, Path instructions

Comment on lines +215 to +231
// Cumulative counters: keep the per-run high-water mark, then
// re-derive the group total. Summing them per call would count a
// single search once for every later call in the run.
let high_water = self
.per_run_disclosure
.entry(entry.invocation_id)
.or_default();
high_water.tool_searches = high_water.tool_searches.max(disclosure.tool_search_count);
high_water.empty_tool_searches = high_water
.empty_tool_searches
.max(disclosure.empty_search_count);
high_water.promotions = high_water.promotions.max(disclosure.promotions);
high_water.recoveries = high_water.recoveries.max(disclosure.recoveries);
high_water.outside_surface_attempts = high_water
.outside_surface_attempts
.max(disclosure.outside_surface_attempts);
self.recompute_disclosure_totals();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Deduplicate cumulative disclosure counters before grouping.

Each ModelCallMetricsAggregate owns a separate per_run_disclosure map. If one run changes effective model or catalog bucket, its cumulative counters are added to each group it enters.

Track the prior counter values by InvocationId across the ordered page. Attribute only the nonnegative delta to the current group. Add a regression test with one run that changes model or bucket between calls.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/events/ironclaw_event_projections/src/model_call_metrics.rs` around
lines 215 - 231, Update the disclosure aggregation around per_run_disclosure to
retain prior cumulative counter values by InvocationId across the ordered page,
compute each call’s nonnegative counter delta, and attribute only that delta to
the current model/catalog group instead of re-adding the run’s high-water
totals. Preserve high-water tracking while preventing duplicate attribution when
a run changes groups, and add a regression test covering a single run switching
model or bucket between calls.

Comment on lines +255 to +266
/// Model calls per completed task, or `None` when no run in this group
/// completed within the observed window — an honest "not yet answerable"
/// rather than a ratio over an empty denominator.
pub fn model_calls_per_completed_run(&self) -> Option<f64> {
if self.completed_runs == 0 {
return None;
}
// Only calls belonging to completed runs may enter the numerator;
// counting in-flight runs' calls would understate the true cost per
// finished task while runs are still open.
Some(self.model_calls as f64 / self.completed_runs as f64)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Exclude in-flight calls from the completed-run ratio.

model_calls_per_completed_run divides self.model_calls by completed_runs. self.model_calls also includes calls from unfinished runs in the same group.

Track calls for completed run IDs separately. Use that count as the numerator. Add a test with one completed run and one in-flight run in the same group.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/events/ironclaw_event_projections/src/model_call_metrics.rs` around
lines 255 - 266, The model_calls_per_completed_run calculation currently
includes calls from in-flight runs; track model-call counts by run ID and
separately accumulate calls belonging to completed runs, then use that
completed-call count as the ratio numerator while preserving None when
completed_runs is zero. Add coverage for a group containing one completed run
and one in-flight run, verifying only the completed run’s calls are included.

Comment on lines +282 to +289
if let Some(record) = event.model_call_metrics.as_ref() {
metrics.push(ModelCallMetricsEntry {
cursor: entry.cursor,
timestamp: event.timestamp,
invocation_id: event.scope.invocation_id,
thread_id: event.scope.thread_id.clone(),
metrics: record.clone(),
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Re-sanitize metrics from custom durable backends.

This path clones event.model_call_metrics into a serializable projection without a wire crossing. A custom DurableEventLog can therefore expose raw model labels, failure labels, or catalog labels through ModelCallMetricsEntry.

Re-apply the event-log label sanitizers before constructing the entry. Add a custom-backend regression test with a path- or secret-shaped model label.

As per path instructions, treat external services as untrusted until a typed boundary establishes trust.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/events/ironclaw_event_projections/src/model_call_metrics.rs` around
lines 282 - 289, Re-sanitize metrics from event.model_call_metrics before
cloning them into ModelCallMetricsEntry, applying the existing event-log
sanitizers for model, failure, and catalog labels so custom DurableEventLog data
is treated as untrusted. Add a regression test using a custom backend that
supplies a path- or secret-shaped model label and verify the projected entry
contains the sanitized value.

Source: Path instructions

Comment on lines +406 to +409
| RuntimeEventKind::FailureRecovered
// Pure measurement. A metrics record must never move a capability
// activity's status, or observability would rewrite run state.
| RuntimeEventKind::ModelCallMetricsRecorded => None,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Add a caller-level status-preservation test.

Test ReplayEventProjectionService::snapshot with ModelCallMetricsRecorded after LoopCompleted. Assert that the run remains Completed. Also test a metrics-only run and assert Running.

As per coding guidelines, “For new or changed production-wired behavior, add a caller-level test at the nearest meaningful seam.”

Also applies to: 450-452

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/events/ironclaw_event_projections/src/runtime_projection.rs` around
lines 406 - 409, The status-preservation behavior for
RuntimeEventKind::ModelCallMetricsRecorded lacks caller-level coverage. Add
tests for ReplayEventProjectionService::snapshot that verify metrics after
LoopCompleted leave the run Completed, while a metrics-only run remains Running;
cover both relevant projection paths without changing production behavior.

Source: Coding guidelines

Comment on lines +85 to +88
| RuntimeEventKind::FailureRecovered
// Loop-host measurement. It carries no hook identity, so its owner is
// never registry-resolved.
| RuntimeEventKind::ModelCallMetricsRecorded => false,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Keep the exhaustive classification test complete.

Line 88 adds ModelCallMetricsRecorded, but is_lifecycle_kind_classifies_every_variant does not include it in non_lifecycle. The test no longer verifies this classification. Add the variant to that array.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/loop/ironclaw_hooks/src/dispatch/lifecycle_owner.rs` around lines 85 -
88, Add RuntimeEventKind::ModelCallMetricsRecorded to the non_lifecycle array in
is_lifecycle_kind_classifies_every_variant, keeping the exhaustive
classification test aligned with the match in lifecycle classification.

Comment on lines +1715 to +1764
/// Publish one model call's cost/latency/disclosure measurements.
///
/// Best-effort observability, exactly like the lifecycle milestones next
/// to it: a failed publish is logged at debug and swallowed. Losing a
/// measurement is an evidence gap; failing the run over one would be a
/// self-inflicted outage caused by telemetry.
async fn emit_model_call_metrics(
&self,
iteration: u32,
requested_model: ModelProfileId,
effective_model: Option<String>,
fallback_index: u32,
result: &Result<LoopModelResponse, AgentLoopHostError>,
duration_ms: u64,
) {
let Some(milestone_sink) = &self.model_call_metrics_sink else {
return;
};
let (failure_kind, usage) = match result {
Ok(response) => (None, response.usage),
Err(error) => (Some(error.kind), error.usage),
};
// Read the disclosure numbers the port already computed. A port that
// is not doing disclosure returns `None`, which is recorded as "no
// disclosure on this call" rather than as zeroed counters — the two
// are different findings for a rollout comparison.
let disclosure = self
.capabilities
.as_ref()
.and_then(|capabilities| capabilities.tool_disclosure_metrics());
let record = ModelCallMetricsRecord {
iteration,
requested_model,
effective_model,
fallback_index,
failure_kind,
duration_ms,
usage: diagnostic_usage(usage),
disclosure,
};
let milestones =
LoopHostMilestoneEmitter::new(self.run_context.clone(), Arc::clone(milestone_sink));
if let Err(error) = milestones.model_call_metrics_recorded(record).await {
tracing::debug!(
kind = ?error.kind,
"loop model call metrics milestone failed; rollout evidence for this call is lost"
);
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Check whether thread_loop_host_contract.rs exercises model_call_metrics_sink / ModelCallMetricsRecorded.
rg -n 'model_call_metrics|ModelCallMetricsRecorded|with_model_call_metrics_sink' crates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rs

Repository: nearai/ironclaw

Length of output: 153


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== candidate files =="
fd -i 'thread_loop_host_contract|lib|model_call|milestone|model_call_metrics' crates/loop/ironclaw_loop_host crates/loop 2>/dev/null | sed -n '1,120p'

echo
echo "== occurrences in ironclaw_loop_host =="
rg -n 'emit_model_call_metrics|model_call_metrics|ModelCallMetricsRecorded|with_model_call_metrics_sink|loop_model_call_metrics|model_call_metrics_recorded|FailOnModel(Metrics|Call)?' crates/loop/ironclaw_loop_host || true

echo
echo "== tests references across repo =="
rg -n 'FailOnModelStartedMilestoneSink|FailOnModelCompletedMilestoneSink|model_call_metrics|ModelCallMetricsRecorded|with_model_call_metrics_sink' crates tests . 2>/dev/null || true

Repository: nearai/ironclaw

Length of output: 32109


Add a caller-level test for emit_model_call_metrics.

emit_model_call_metrics is new production-wired behavior invoked for stream_model calls. Per the loop crate guideline, new production-wired behavior needs a caller-level test at the nearest seam. Add coverage through thread_loop_host_contract.rs for with_model_call_metrics_sink / ModelCallMetricsRecorded, including milestone-sink failure behavior.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/loop/ironclaw_loop_host/src/lib.rs` around lines 1715 - 1764, Add
caller-level coverage in thread_loop_host_contract.rs for the stream_model path
configured via with_model_call_metrics_sink, asserting a
ModelCallMetricsRecorded milestone is emitted with the expected data. Also
configure a failing milestone sink and verify emit_model_call_metrics swallows
the publish error without failing the loop operation.

Source: Coding guidelines

Comment on lines +269 to +299
fn tool_disclosure_metrics(&self) -> Option<ToolDisclosureCallMetrics> {
// Read-only. Every number here was already computed for the shadow
// logs; this returns them instead of recomputing, so the durable
// record and the log line can never disagree.
let guard = self.turn_state().ok()?;
let state = guard.as_ref()?;
let (full_tool_count, full_schema_tokens) = state.catalog.effective_metrics(&self.policy);
let (
tool_search_count,
empty_search_count,
selected_result_rank,
promotions,
recoveries,
outside_surface_attempts,
) = self.counters.snapshot();
Some(ToolDisclosureCallMetrics {
deferred: state.active.deferred,
full_tool_count: u32::try_from(full_tool_count).unwrap_or(u32::MAX),
advertised_tool_count: u32::try_from(state.active.definitions.len())
.unwrap_or(u32::MAX),
full_schema_tokens,
advertised_schema_tokens: state.active.advertised_tokens,
tool_search_count,
empty_search_count,
selected_result_rank,
promotions,
recoveries,
outside_surface_attempts,
})
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Confirm no existing test in this file asserts tool_disclosure_metrics() output.
rg -n 'tool_disclosure_metrics\(\)' crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs

Repository: nearai/ironclaw

Length of output: 153


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== file presence and size =="
wc -l crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs || true

echo "== relevant function/context occurrences in file =="
rg -n "tool_disclosure_metrics|DisclosureRunCounters|record_search|record_promotion|record_recovery|record_outside_surface_attempt|record_selected_rank|mod spy|spy|SpyPort" crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs || true

echo "== test blocks in file =="
rg -n "#\\[cfg\\(test\\)|#\\[test\\]|mod tests|AsyncTrait|mock|expect|assert" crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs || true

echo "== nearby function implementations =="
sed -n '240,310p' crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs
echo "== counter call implementations =="
sed -n '990,1215p' crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs

Repository: nearai/ironclaw

Length of output: 30129


Add a unit test asserting tool_disclosure_metrics() counter values.

tool_disclosure_metrics() reads DisclosureRunCounters populated by record_search, record_promotion, record_recovery, record_outside_surface_attempt, and record_selected_rank, but no caller-level test in this crate-tier SpyPort suite asserts tool_search_count, empty_search_count, promotions, recoveries, selected_result_rank, or outside_surface_attempts. Add a tool_disclosure_metrics() assertion at this seam so the durable counter path stays tied to disclosure telemetry.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/loop/ironclaw_loop_host/src/tool_disclosure_port.rs` around lines 269
- 299, Add a caller-level unit test in the SpyPort suite that invokes the
counter-recording methods record_search, record_promotion, record_recovery,
record_outside_surface_attempt, and record_selected_rank, then calls
tool_disclosure_metrics() and asserts tool_search_count, empty_search_count,
promotions, recoveries, selected_result_rank, and outside_surface_attempts. Keep
the assertion at this seam to verify the durable DisclosureRunCounters telemetry
path.

Source: Coding guidelines

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7385 — b8f9fdf4 Deployed Aug 7, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant