Skip to content

feat(inspector): add model call statistics - #7277

Merged
think-in-universe merged 37 commits into
mainfrom
issue-7223-model-call-stats
Aug 7, 2026
Merged

think-in-universe merged 37 commits into
mainfrom
issue-7223-model-call-stats

Conversation

@italic-jinxin

Copy link
Copy Markdown
Contributor

Summary

  • Adds run-scoped model-call statistics to the Inspector, including call counts, latency, input/output token usage, cache token usage, and per-model breakdowns.
  • Captures the effective provider model for each individual call so concurrent requests, fallbacks, and failed calls are attributed correctly.
  • Clearly reports unavailable or partial usage evidence instead of treating missing metrics as zero.
  • Keeps tool execution metrics out of the UI until tool-level diagnostic capture is implemented.
  • Adds Rust, frontend, and browser E2E coverage. Browser E2E changes are isolated in a separate commit.

Linked Issue

Closes #7223

Depends on #7222

Part of #7218

Security Impact

Diagnostic statistics remain operator-only, caller-scoped, bounded, and process-local. No credentials or unredacted prompt content are added to the statistics response.

Database Impact

None. No schema or migration changes.

Blast Radius

Limited to model-call diagnostic capture, the operator inspection API, and the Inspector statistics tab.

Rollback Plan

Revert this PR to remove model-call aggregation and the Inspector statistics tab without changing the underlying diagnostic session storage.

@italic-jinxin italic-jinxin added size: XL 500+ changed lines contributor: core 20+ merged PRs labels Aug 6, 2026
@coderabbitai

This comment was marked as resolved.

@railway-app

railway-app Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7277 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 7, 2026 at 2:00 pm

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7277 August 6, 2026 10:54 Destroyed
@github-actions github-actions Bot added scope: docs Documentation risk: low Changes to docs, tests, or low-risk modules labels Aug 6, 2026
@italic-jinxin
italic-jinxin marked this pull request as draft August 6, 2026 10:54
@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

@ironloopai

This comment was marked as resolved.

ironloopai[bot]

This comment was marked as resolved.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7277 August 7, 2026 11:40 Destroyed
@think-in-universe

Copy link
Copy Markdown
Collaborator

@ironloopai review

@ironloopai

ironloopai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

🧭 IronLoop Run · Review

This comment updates in place as the Run moves through its stages.

⬛ Final result · Stopped

🟨 Queued → 🟦 Working → ⬛ Stopped

Manual command by think-in-universe · attempt 1 of 3 · stopped after 13m 32s

IronLoop stopped because the pull request target branch or head changed while this Run was active.

Run details

Run: 7f7ea39b-03c1-43c1-8014-f19a346fc8a4
Base: issue-7222-prompt-inspection at d70ff59
Head: issue-7223-model-call-stats at 14352aa
Created: 2026-08-07 12:07 UTC
Updated: 2026-08-07 12:20 UTC

Base automatically changed from issue-7222-prompt-inspection to main August 7, 2026 12:20
@think-in-universe
think-in-universe dismissed their stale review August 7, 2026 12:20

The base branch was changed.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7277 August 7, 2026 12:59 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 6

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/loop/ironclaw_loop_host/src/lib.rs (1)

2101-2125: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Mark dropped model-call captures as partial.

A full queue drops ModelCall captures with only a debug log. The Inspector sink receives no loss signal. Saturation can therefore show lower call, token, and latency totals as complete statistics.

Retain the non-blocking queue. Add a run-scoped loss marker or truncation state that reaches the Inspector snapshot. Add a caller-level saturation test that verifies the Stats tab reports partial diagnostics.

Based on PR objectives, partial and truncated diagnostic data must be explicit. As per path instructions, “Fail loud” prohibits silently continuing with poisoned state.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/loop/ironclaw_loop_host/src/lib.rs` around lines 2101 - 2125, Update
BufferedPromptDiagnosticSink::enqueue to propagate a run-scoped loss or
truncation marker when a ModelCall capture is dropped because the queue is full,
while preserving the non-blocking try_send behavior. Ensure the marker reaches
the Inspector snapshot so model-call statistics are reported as partial rather
than complete, and add a caller-level saturation test verifying the Stats tab
exposes partial diagnostics.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/loop/ironclaw_loop_host/src/lib.rs`:
- Around line 1578-1588: Remove the final model_profile_id fallback from
effective_model construction in the host loop, preserving None when
diagnostic_effective_model and diagnostic_effective_model evidence are
unavailable. Propagate the optional value through capture and the Inspector
persistence/display contract so missing provider evidence is represented as
unavailable rather than as a logical profile. Add a caller-level test covering
absent evidence with a nonzero fallback index.

In `@crates/product/ironclaw_assistant/src/inspector_store.rs`:
- Around line 385-402: In
crates/product/ironclaw_assistant/src/inspector_store.rs:385-402, update the
model-call statistics path to use a bounded identity set of counted call_id
values independent of the retained model_calls deque, preventing evicted Started
records from being recounted on terminal updates. Apply the same approach at
crates/product/ironclaw_assistant/src/inspector_store.rs:413-435 for activity_id
and tool-call statistics. Add a regression test that fills
max_model_calls_per_run, evicts a Started record, records its Succeeded update,
and verifies total_model_calls is unchanged.
- Around line 632-651: Update update_model_count so calls_per_model_truncated is
cleared when removing an existing model causes calls_per_model to fall below
MAX_MODELS_IN_STATS, allowing the breakdown to be reported as complete again. If
the flag is intentionally monotonic because omitted models cannot be recovered,
preserve the latch and document that decision directly beside the flag update.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.tsx`:
- Around line 317-330: Update the calls_per_model row key in the stats rendering
map to use the array index rather than the truncated entry.model.content value.
Correct the partialMetricCount display so it either counts affected calls or
explicitly labels the value as unavailable metric samples across the tracked
metrics; do not present the summed metric count as individual samples when it
represents multiple metrics per call.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts`:
- Around line 201-206: Update useInspector() so diagnostic updates in the
prompt_updated/model_call/tool_execution_updated/stats branch coalesce snapshot
refreshes with a short trailing debounce or pending-refresh flag instead of
incrementing snapshotGeneration for every event. Also serialize the stream SSE
cursor in the snapshot request’s last-event-id header and stop re-querying it
solely through after_cursor.

In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py`:
- Around line 477-490: Move the inline stats selectors used around the inspector
stats assertions into helpers.SEL_V2, adding entries for the stats tab and stats
content and referencing those entries here. Update the “Model calls” and “Input
tokens” assertions to verify the exact rendered values rather than substring
matches, while preserving the existing visibility and other stats checks.

---

Outside diff comments:
In `@crates/loop/ironclaw_loop_host/src/lib.rs`:
- Around line 2101-2125: Update BufferedPromptDiagnosticSink::enqueue to
propagate a run-scoped loss or truncation marker when a ModelCall capture is
dropped because the queue is full, while preserving the non-blocking try_send
behavior. Ensure the marker reaches the Inspector snapshot so model-call
statistics are reported as partial rather than complete, and add a caller-level
saturation test verifying the Stats tab exposes partial diagnostics.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: f9650ab3-9eee-4aea-80f5-b4d4f6294941

📥 Commits

Reviewing files that changed from the base of the PR and between 12e5a30 and a0e8cf9.

📒 Files selected for processing (21)
  • crates/app/ironclaw_composition/src/observability/budget.rs
  • crates/contracts/ironclaw_loop_contracts/src/host/model.rs
  • crates/kernel/ironclaw_turns/tests/agent_loop_host_contract.rs
  • crates/loop/ironclaw_agent_loop/src/executor/failure_explanation.rs
  • crates/loop/ironclaw_agent_loop/src/executor/model.rs
  • crates/loop/ironclaw_agent_loop/src/executor/tests/cancellation.rs
  • crates/loop/ironclaw_hooks/src/middleware/model_port.rs
  • crates/loop/ironclaw_loop_host/src/budget_accountant.rs
  • crates/loop/ironclaw_loop_host/src/lib.rs
  • crates/loop/ironclaw_loop_host/src/model_gateway.rs
  • crates/loop/ironclaw_loop_host/tests/llm_gateway.rs
  • crates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rs
  • crates/loop/ironclaw_turn_runner/src/text_loop_driver.rs
  • crates/product/ironclaw_assistant/src/inspector_store.rs
  • crates/product/ironclaw_assistant/tests/support/planned_agent_loop.rs
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.test.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.test.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts
  • tests/CLAUDE.md
  • tests/e2e/scenarios/test_reborn_webui_v2_smoke.py

Comment thread crates/loop/ironclaw_loop_host/src/lib.rs Outdated
Comment thread crates/product/ironclaw_assistant/src/inspector_store.rs
Comment thread crates/product/ironclaw_assistant/src/inspector_store.rs
Comment thread tests/e2e/scenarios/test_reborn_webui_v2_smoke.py Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7277 August 7, 2026 13:29 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/loop/ironclaw_loop_host/src/lib.rs (1)

2310-2310: 🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift

Validate provider model evidence before diagnostic capture.

The response and error fields accept arbitrary String values. Line 1513 copies those values into HostManagedModelCallDiagnosticCapture. A provider can therefore place credential-like or otherwise unsafe text into Inspector diagnostics.

Use ProviderModelId for both evidence fields, or validate in both builders and omit invalid evidence. Add a caller-level test that invalid provider evidence never reaches record_model_call.

  • crates/loop/ironclaw_loop_host/src/lib.rs#L2310-L2310: replace Option<Arc<String>> with validated ProviderModelId evidence.
  • crates/loop/ironclaw_loop_host/src/lib.rs#L2450-L2450: apply the same validated type to error evidence.
  • crates/loop/ironclaw_loop_host/src/lib.rs#L1513-L1522: preserve only validated evidence when creating the diagnostic capture.

As per coding guidelines, “Prefer strong types such as enums and newtypes over strings.” Based on PR objectives, diagnostic captures must contain no credentials or unredacted content.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/loop/ironclaw_loop_host/src/lib.rs` at line 2310, Replace the response
evidence field at crates/loop/ironclaw_loop_host/src/lib.rs:2310-2310 and the
error evidence field at crates/loop/ironclaw_loop_host/src/lib.rs:2450-2450 with
validated ProviderModelId values instead of Option<Arc<String>>. In the
diagnostic capture construction at
crates/loop/ironclaw_loop_host/src/lib.rs:1513-1522, propagate only validated
evidence so unsafe provider text cannot reach record_model_call. Add a
caller-level test verifying invalid provider evidence is omitted and never
recorded.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/product/ironclaw_assistant/src/inspector_store.rs`:
- Around line 1809-1820: Update
model_breakdown_truncation_remains_latched_after_a_bucket_is_removed to exercise
record_model_call instead of update_model_count: record MAX_MODELS_IN_STATS + 1
distinct effective models, replace one model’s contribution with zero, then
assert snapshot.stats.calls_per_model_truncated remains true. Use the store’s
public record path and its snapshot result so replacement reversal is covered.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts`:
- Around line 151-157: Update scheduleSnapshotRefresh to retain the first
refresh deadline while continuing to debounce subsequent updates, using
SNAPSHOT_REFRESH_MAX_WAIT_MS as the upper bound. Ensure frequent qualifying
events cannot postpone the snapshot indefinitely, while preserving timer cleanup
and the disposed guard.

---

Outside diff comments:
In `@crates/loop/ironclaw_loop_host/src/lib.rs`:
- Line 2310: Replace the response evidence field at
crates/loop/ironclaw_loop_host/src/lib.rs:2310-2310 and the error evidence field
at crates/loop/ironclaw_loop_host/src/lib.rs:2450-2450 with validated
ProviderModelId values instead of Option<Arc<String>>. In the diagnostic capture
construction at crates/loop/ironclaw_loop_host/src/lib.rs:1513-1522, propagate
only validated evidence so unsafe provider text cannot reach record_model_call.
Add a caller-level test verifying invalid provider evidence is omitted and never
recorded.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 33db8bb5-d5b2-435c-b021-7c5cbc3ae74c

📥 Commits

Reviewing files that changed from the base of the PR and between a0e8cf9 and 5912f4b.

📒 Files selected for processing (10)
  • crates/loop/ironclaw_loop_host/src/lib.rs
  • crates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rs
  • crates/product/ironclaw_assistant/src/inspector_store.rs
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.test.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.test.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts
  • tests/e2e/helpers.py
  • tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
  • tests/support/reborn_parity_qa/binary_e2e.rs

Comment thread crates/product/ironclaw_assistant/src/inspector_store.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7277 August 7, 2026 13:50 Destroyed
@think-in-universe
think-in-universe added this pull request to the merge queue Aug 7, 2026
Merged via the queue into main with commit 81724a6 Aug 7, 2026
47 checks passed
@think-in-universe
think-in-universe deleted the issue-7223-model-call-stats branch August 7, 2026 14:16
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* feat(inspector): add operator inspection API

* docs(inspector): assign product service ownership

* test(inspector): ratchet diagnostic contracts

* feat(inspector): add debug panel shell

* test(inspector): cover debug panel shell e2e

* fix(inspector): stop diagnostics when panel closes

* feat(inspector): add prompt inspection

* fix(inspector): follow current webui ownership

* feat(inspector): add model call statistics

* test(inspector): cover model statistics e2e

* fix(inspector): avoid uncollected tool metrics

* test(inspector): cover prompt diagnostics e2e

* test(inspector): align statistics e2e scope

* fix(inspector): redact prompt metadata

* fix(inspector): preserve per-call model identity

* fix(inspector): classify prompt instruction sources

* test(inspector): assert reported token usage

* fix(inspector): address review feedback

* fix(inspector): retry transient snapshot failures

* fix(inspector): address prompt diagnostic review findings

* fix(inspector): follow debug query navigation

* fix(inspector): preserve stream terminal state

* fix(inspector): capture full capability surface

* fix(inspector): address prompt diagnostic review feedback

* fix inspector model call stats review findings

* fix inspector refresh and truncation regressions

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7277 — 88ccedc1 Deployed Aug 7, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Inspector] Add model-call metrics and the Stats tab

2 participants