Skip to content

feat(inspector): complete statistics, navigation, and localization - #7291

Merged
think-in-universe merged 78 commits into
mainfrom
issue-7287-inspector-ui-parity
Aug 10, 2026
Merged

think-in-universe merged 78 commits into
mainfrom
issue-7287-inspector-ui-parity

Conversation

@italic-jinxin

Copy link
Copy Markdown
Contributor

Summary

  • Adds total, successful, and failed tool-call statistics to the Inspector Stats tab.
  • Adds direct latest-turn navigation and browser-observed stream health metrics with bounded, duplicate-safe session storage.
  • Localizes Inspector labels, states, errors, metrics, buttons, and accessibility text, with complete English, Simplified Chinese, and Korean coverage.
  • Updates the Inspector contract, user documentation, frontend tests, and browser E2E coverage.
image image image

Linked Issue

Closes #7287

Depends on #7226

Part of #7218

Validation

  • git diff --check
  • Frontend test suite: 1126 tests passed
  • Inspector focused suite: 21 tests passed
  • TypeScript typecheck and source convention checks
  • Vite production build
  • WebUI locale consistency tests: 5 passed
  • Inspector browser E2E scenarios
  • Manual testing of Prompt, Activity, Stats, navigation, stream metrics, and responsive presentation

The existing bundle-budget check remains above its configured threshold at
218.7 KB on both this branch and the unchanged dependency branch. The Inspector
changes add no gzip bytes to the ordinary /chat initial bundle because they
remain lazy-loaded.

Test Strategy

User behavior:

Operators using debug=true can inspect prompt composition, ordered run
activity, model and tool statistics, previous/latest turns, and browser stream
health. Ordinary chat remains unchanged when diagnostics are disabled.

Risk areas:

  • Browser
  • Cross-component behavior
  • Model behavior
  • Side effect
  • Persistence
  • Security or permissions
  • External provider

Tests added or updated:

  • Unit or contract: Inspector rendering, missing values, navigation, bounded
    counters, cursor deduplication, and translated accessibility labels.
  • Reborn integration: Locale consistency tests.
  • Recorded fixture: Not applicable; no provider protocol changes.
  • Browser E2E: Real diagnostic rendering, latest-turn navigation, reconnect
    metrics, localization, responsive layout, and disabled-debug behavior.
  • Backend or runtime: Not applicable; no backend or runtime changes.
  • Live canary: Not applicable; deterministic local browser coverage reaches
    the changed contract.

What the tests prove:

The Inspector displays host-provided diagnostics without fabricating unavailable
values, keeps browser counters bounded and duplicate-safe, preserves chat during
stream reconnects, and leaves ordinary chat unaffected.

Commands run:

  • Frontend Vitest suite with TZ=America/Los_Angeles
  • Focused Inspector Vitest suite
  • TypeScript typecheck
  • Vite production build
  • SKIP_FRONTEND_BUILD=1 cargo test -p ironclaw_webui --test i18n_consistency
  • Focused Inspector Playwright E2E scenarios

Security Impact

None. Diagnostic authorization, redaction, retention, and backend contracts are
unchanged. Browser-observed counters contain no diagnostic payloads and use
bounded session storage.

Database Impact

None. No migrations, schemas, or persistent storage behavior changed.

Blast Radius

Limited to the opt-in Web Inspector frontend, its localized strings,
documentation, and browser coverage. The ordinary chat path remains lazy and
unchanged when debug=true is absent.

Rollback Plan

Revert the two commits to restore the previous Inspector presentation and E2E
coverage. No data migration or backend rollback is required.

Review Follow-Through

Full review found no blocking issues. Remaining limitations—provider cost
accounting, raw reasoning, and persistent diagnostics—are explicitly outside
the issue scope.


Review track: B

…copy

A `calls_per_model` entry with a negative or non-integer `calls` passed the
statistics decoder and was then coerced to zero during accumulation without
marking the breakdown truncated, presenting a fabricated "0 calls" for a model.
Every entry is now validated before a record is accepted. German turn
navigation used "Zug" (a train, or a game move); it now reads "Runde", with the
determiner agreement that noun requires. Spanish and Portuguese tool-status
values were written feminine against a masculine "Estado"/"Status" label.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7291 August 8, 2026 15:32 Destroyed
coderabbitai[bot]

This comment was marked as resolved.

…ng it

The statistics decoder validated every calls_per_model entry but never the
array length, so an out-of-contract response was scanned in full and then
retained by the accumulator for up to 128 runs. The host truncates this
breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer
array cannot conform; the client now mirrors that ceiling and rejects the
record before the scan.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7291 August 8, 2026 15:44 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (5)
crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts (4)

178-200: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Do not merge truncated model labels as stable identities.

The accumulator keys entries by entry.model.content. When entry.model.truncated is true, two distinct model names can share the same retained prefix and their calls are merged incorrectly.

Use an authoritative model identifier, or mark truncated model breakdowns as non-mergeable and incomplete. Add a collision test.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts`
around lines 178 - 200, Update the model aggregation logic around models and
sourceModels so entries with model.truncated set are not merged using
model.content as a stable key; use an authoritative identifier when available,
otherwise treat the breakdown as non-mergeable and mark
calls_per_model_truncated. Add a test covering distinct truncated model names
sharing the same retained prefix and verify their calls are not incorrectly
combined.

130-135: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Preserve unavailable tool totals instead of fabricating zero values.

Lines 130-135 initialize optional tool totals to zero. Lines 162-170 add omitted source fields as zero. A valid older-host record without tool totals therefore becomes 0 for all three tool metrics in snapshot(). Mixed records also lose the fact that some runs are unknown.

Preserve these fields as unavailable. Track aggregate completeness when records mix omitted and present fields. Add an accumulator test for omitted tool totals.

The PR objective requires explicit missing or unavailable states instead of fabricated zero values.

Also applies to: 162-170, 266-273

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts`
around lines 130 - 135, Update emptyStats, the source-field accumulation around
snapshot(), and the accumulator logic around lines 266-273 so omitted tool
totals remain unavailable rather than becoming zero. Track whether tool totals
are present across records, preserving unknown state for entirely omitted fields
and representing mixed present/omitted records as incomplete. Add an accumulator
test covering records with omitted tool totals.

111-113: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Do not evict deduplication keys while retaining carried totals.

After 1,153 distinct runs, retiredRunIds evicts the first run while carried still contains its statistics. A later historical-navigation replay is accepted and counted again. The aggregate then double-counts that run.

While an ID remains retired, later authoritative snapshots are ignored. Keep exact run identity for the supported session, or define a different replay contract. Add a boundary test.

The PR objective requires bounded session statistics and latest-turn navigation without duplicate totals.

Also applies to: 240-261

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts`
around lines 111 - 113, Update retiredRunIds and carried retention so a run
whose statistics remain in carried can never be evicted from deduplication
tracking; preserve exact run identity for the supported session, or atomically
discard its carried totals under an explicitly defined replay contract. Keep
bounded session statistics and latest-turn navigation without double-counting,
and add a boundary test covering more than MAX_RETIRED_RUN_IDS distinct runs
followed by replay of the earliest run.

67-108: 🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Reject oversized BoundedDiagnosticText before retaining stats.

isBoundedDiagnosticText currently only checks shape/type; content.original_bytes is never used. Apply the owning limit for SessionDiagnosticStats::model before decoder.record(...): reject empty content when original_bytes > 0 or original_bytes exceeds BoundedDiagnosticText’s retained maximum, and enforce an existing field bound for model such as DIAGNOSTIC_LABEL_MAX_BYTES.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts`
around lines 67 - 108, Update decodeSessionDiagnosticStats to validate the model
BoundedDiagnosticText before returning stats for decoder.record. Reject empty
model content when original_bytes is positive, reject original_bytes above the
retained BoundedDiagnosticText maximum, and enforce the existing
DIAGNOSTIC_LABEL_MAX_BYTES bound on retained model content. Preserve valid model
diagnostics and return null for violations.

Source: Path instructions

crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.ts (1)

10-41: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Cover successful and failed tool-call totals.

stats() hardcodes failed_tool_calls: 0, and the accumulator test checks only total_tool_calls. A regression in successful or failed tool-call aggregation can pass.

Parameterize failed calls and assert all three counters after the refreshed run.

The PR objective requires total, successful, and failed tool-call counts.

Also applies to: 44-84

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.ts`
around lines 10 - 41, Update the test helper stats() to accept a
failed-tool-call count and derive successful_tool_calls from total tools minus
failed calls instead of hardcoding it. In the accumulator test covering the
refreshed run, assert total_tool_calls, successful_tool_calls, and
failed_tool_calls so aggregation regressions for each counter are detected.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.ts`:
- Around line 10-41: Update the test helper stats() to accept a failed-tool-call
count and derive successful_tool_calls from total tools minus failed calls
instead of hardcoding it. In the accumulator test covering the refreshed run,
assert total_tool_calls, successful_tool_calls, and failed_tool_calls so
aggregation regressions for each counter are detected.

In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts`:
- Around line 178-200: Update the model aggregation logic around models and
sourceModels so entries with model.truncated set are not merged using
model.content as a stable key; use an authoritative identifier when available,
otherwise treat the breakdown as non-mergeable and mark
calls_per_model_truncated. Add a test covering distinct truncated model names
sharing the same retained prefix and verify their calls are not incorrectly
combined.
- Around line 130-135: Update emptyStats, the source-field accumulation around
snapshot(), and the accumulator logic around lines 266-273 so omitted tool
totals remain unavailable rather than becoming zero. Track whether tool totals
are present across records, preserving unknown state for entirely omitted fields
and representing mixed present/omitted records as incomplete. Add an accumulator
test covering records with omitted tool totals.
- Around line 111-113: Update retiredRunIds and carried retention so a run whose
statistics remain in carried can never be evicted from deduplication tracking;
preserve exact run identity for the supported session, or atomically discard its
carried totals under an explicitly defined replay contract. Keep bounded session
statistics and latest-turn navigation without double-counting, and add a
boundary test covering more than MAX_RETIRED_RUN_IDS distinct runs followed by
replay of the earliest run.
- Around line 67-108: Update decodeSessionDiagnosticStats to validate the model
BoundedDiagnosticText before returning stats for decoder.record. Reject empty
model content when original_bytes is positive, reject original_bytes above the
retained BoundedDiagnosticText maximum, and enforce the existing
DIAGNOSTIC_LABEL_MAX_BYTES bound on retained model content. Preserve valid model
diagnostics and return null for violations.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 93250f3d-9135-47ee-a325-360a440bd616

📥 Commits

Reviewing files that changed from the base of the PR and between 6d6d4f4 and 3d40283.

📒 Files selected for processing (2)
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.ts
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts

The browser offered 32 turns of navigation per thread while the host retained
diagnostics for 2 runs per session, so every turn past the second rendered
blank. Each layer was individually correct and the e2e scenario stopped at two
turns, so nothing saw the dead zone. Retention moves to 4 and the navigation
window mirrors it, pinned by a new architecture gate that reads both constants;
the scenario now walks back two turns and asserts real activity. Retention is a
ceiling as well as a default, and capture is unconditional, so 4 is a resident
memory choice — roughly 80 MB worst case across the eight tracked sessions.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7291 August 8, 2026 16:07 Destroyed
@github-actions github-actions Bot added the scope: dependencies Dependency updates label Aug 8, 2026
@italic-jinxin italic-jinxin added the human-verified Manually tested and verified label Aug 8, 2026
coderabbitai[bot]

This comment was marked as resolved.

…rker

The guard sliced the concatenated chunk bundle from the i18n provider up to
`QueryClient`, a symbol another module owns, so the segment's extent tracked
Rollup's chunk boundaries. A split that merely folded react-query into the
entry chunk removed that marker from everything appended after the provider and
failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES
literal that follows the provider in the same module; string literals survive
minification, and every existing assertion holds against the tighter segment.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7291 August 8, 2026 16:19 Destroyed
…e_path

The gate joined a family-nested literal onto the workspace root, the idiom
crate_path exists to replace: a crate family move would have turned this into a
read failure rather than a resolved path. It now names the SPA file in the
logical flat spelling and resolves it, and the assertion reports the resolved
path so the message still points at a file that exists.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7291 August 8, 2026 16:37 Destroyed
The multi-turn scenario assumed the panel would jump to an arriving turn, but a
selection the operator navigated to is deliberately sticky: the new turn widens
the window without yanking them off the turn they are reading. The scenario now
asserts that guarantee, then clicks Latest to follow, then walks back two turns
as before. Verified by running the inspector scenarios locally rather than by
reading, which is how this slipped through the first time.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7291 August 8, 2026 16:43 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
tests/e2e/scenarios/test_reborn_webui_v2_smoke.py (2)

690-699: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Require a positive update count before testing persistence.

updates_before_reload can be 0. The equality check after reload then passes without proving that reconnect processing recorded any diagnostic update. Assert updates_before_reload > 0 before calling page.reload().

Proposed assertion
         updates_before_reload = int((await updates.inner_text()).replace(",", ""))
+        assert updates_before_reload > 0
         await page.reload()

Based on the PR objective: “Showing ... diagnostic update count” and keeping counters within the current browser session.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py` around lines 690 - 699, In
the smoke test around updates_before_reload, assert that the parsed update count
is greater than zero before calling page.reload(). Keep the existing persistence
assertion unchanged after the reload.

534-542: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Assert exact tool-call counter values.

to_contain_text("0") checks the whole card. It also passes when the value is 10 or 20. A regression that records tool calls can therefore pass this test. Select the value element and use to_have_text("0") for each counter.

Proposed assertion
-        await expect(stats.get_by_text("Tool calls", exact=True).locator("..")).to_contain_text(
-            "0"
-        )
+        await expect(
+            stats.get_by_text("Tool calls", exact=True).locator("..").locator("p").nth(1)
+        ).to_have_text("0")

Apply the same exact-value assertion to Successful tool calls and Failed tool calls.

Based on the PR objective: “Displaying total, successful, and failed tool-call counts.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py` around lines 534 - 542,
Update the tool-call counter assertions in the smoke test to target each card’s
value element rather than the whole card, then use exact text assertions
expecting “0” for Tool calls, Successful tool calls, and Failed tool calls.
Preserve the existing labels and counter coverage while preventing values such
as 10 or 20 from matching.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py`:
- Around line 624-631: After clicking “Latest turn” in the activity navigation
flow, capture the completed third run’s ID and assert that
SEL_V2["inspector_panel"] contains that ID, alongside the existing “Turn 3 of 3”
assertion. Use the third-run identity established by the surrounding test rather
than reusing second_run_id.

---

Outside diff comments:
In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py`:
- Around line 690-699: In the smoke test around updates_before_reload, assert
that the parsed update count is greater than zero before calling page.reload().
Keep the existing persistence assertion unchanged after the reload.
- Around line 534-542: Update the tool-call counter assertions in the smoke test
to target each card’s value element rather than the whole card, then use exact
text assertions expecting “0” for Tool calls, Successful tool calls, and Failed
tool calls. Preserve the existing labels and counter coverage while preventing
values such as 10 or 20 from matching.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c8a949a3-54ec-4bcd-b31f-3b64af8d3065

📥 Commits

Reviewing files that changed from the base of the PR and between d3dd385 and 0f8cf14.

📒 Files selected for processing (1)
  • tests/e2e/scenarios/test_reborn_webui_v2_smoke.py

Comment thread tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
stream: true,
stream_options: None,
stream_options: Some(ChatCompletionStreamOptions {
include_usage: true,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is include_usage needed for all NEAR AI endpoints?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need the token counts back from streaming responses. NEAR's gateway only puts a usage block in the stream if you explicitly ask for it, and when it's absent our parser reads it as zero — so every streaming turn was quietly being recorded as 0 tokens, which then feeds cost, budget, resource accounting, and the usage we hand back to our own OpenAI-compat clients.

@think-in-universe
think-in-universe added this pull request to the merge queue Aug 10, 2026
Merged via the queue into main with commit 9fd1e63 Aug 10, 2026
51 checks passed
@think-in-universe
think-in-universe deleted the issue-7287-inspector-ui-parity branch August 10, 2026 06:26
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…earai#7291)

* feat(inspector): add operator inspection API

* docs(inspector): assign product service ownership

* test(inspector): ratchet diagnostic contracts

* feat(inspector): add debug panel shell

* test(inspector): cover debug panel shell e2e

* fix(inspector): stop diagnostics when panel closes

* feat(inspector): add prompt inspection

* fix(inspector): follow current webui ownership

* feat(inspector): add model call statistics

* test(inspector): cover model statistics e2e

* fix(inspector): avoid uncollected tool metrics

* test(inspector): cover prompt diagnostics e2e

* test(inspector): align statistics e2e scope

* fix(inspector): redact prompt metadata

* fix(inspector): preserve per-call model identity

* fix(inspector): classify prompt instruction sources

* test(inspector): assert reported token usage

* feat(inspector): add activity timeline and turn navigation

* test(inspector): cover activity timeline in browser

* fix(inspector): read current run before publishing activity

* feat(inspector): add bounded tool execution details

* test(inspector): cover bounded tool details in browser

* fix(inspector): validate retained tool result sizes

* test(inspector): add security and operator coverage

* test(inspector): cover browser workflows end to end

* fix(inspector): address review feedback

* fix(inspector): retry transient snapshot failures

* fix(inspector): address prompt diagnostic review findings

* fix(inspector): follow debug query navigation

* feat(inspector): complete frontend diagnostics

* test(inspector): cover frontend parity in browser

* fix(inspector): preserve stream terminal state

* fix(inspector): capture full capability surface

* fix(inspector): scope projection activity to its run

* fix(inspector): harden activity diagnostics

* fix(inspector): bound tool result diagnostic capture

* fix(inspector): harden tool diagnostic pipeline

* fix(llm): request usage for NEAR AI streams

* fix(inspector): address prompt diagnostic review feedback

* fix(webui): harden inspector stream coverage

* fix(inspector): preserve debug session statistics

* fix(inspector): keep diagnostics active while hidden

* test(e2e): cover hidden inspector observation

* fix inspector model call stats review findings

* fix inspector refresh and truncation regressions

* fix(inspector): address activity timeline review feedback

* fix(inspector): harden activity lifecycle handling

* fix(composition): move tool diagnostics to loop host

* fix(inspector): keep a settled stream live and complete locale parity

A live diagnostic update's debounced snapshot refresh was announcing LOADING,
so an open, healthy stream read as "Connecting" indefinitely once a run
settled — the settling stats update is the last one. That refresh is now a
background read. Incomplete snapshot statistics no longer accumulate as real
zeros, browser-session inspector state is namespaced by the authenticated
caller, an evicted pinned run rejoins the latest turn instead of the oldest,
tool status is localized, and the inspector strings now cover all ten locales.

* test(inspector): put the inspector locale sidecar under the parity gate

The inspector's English copy is registered from its lazy chunk instead of
src/i18n/en.ts, so the all-locale parity test — which derives the required key
set from en.ts — never covered those keys; a locale could drop one and fall
back to English silently. The test now treats the English key set as the union
of en.ts and a declared sidecar list. Keeping the copy in en.ts is not an
option: measured, it puts /chat at 217.4 KB gzip against a 217.0 KB budget.

* fix(inspector): reject malformed model breakdowns and correct locale copy

A `calls_per_model` entry with a negative or non-integer `calls` passed the
statistics decoder and was then coerced to zero during accumulation without
marking the breakdown truncated, presenting a fabricated "0 calls" for a model.
Every entry is now validated before a record is accepted. German turn
navigation used "Zug" (a train, or a game move); it now reads "Runde", with the
determiner agreement that noun requires. Spanish and Portuguese tool-status
values were written feminine against a masculine "Estado"/"Status" label.

* fix(inspector): bound the model breakdown before scanning and retaining it

The statistics decoder validated every calls_per_model entry but never the
array length, so an out-of-contract response was scanned in full and then
retained by the accumulator for up to 128 runs. The host truncates this
breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer
array cannot conform; the client now mirrors that ceiling and rejects the
record before the scan.

* fix(inspector): align turn navigation with host diagnostic retention

The browser offered 32 turns of navigation per thread while the host retained
diagnostics for 2 runs per session, so every turn past the second rendered
blank. Each layer was individually correct and the e2e scenario stopped at two
turns, so nothing saw the dead zone. Retention moves to 4 and the navigation
window mirrors it, pinned by a new architecture gate that reads both constants;
the scenario now walks back two turns and asserts real activity. Retention is a
ceiling as well as a default, and capture is unconditional, so 4 is a resident
memory choice — roughly 80 MB worst case across the eight tracked sessions.

* fix(composition): delimit the i18n bundle guard with an i18n-owned marker

The guard sliced the concatenated chunk bundle from the i18n provider up to
`QueryClient`, a symbol another module owns, so the segment's extent tracked
Rollup's chunk boundaries. A split that merely folded react-query into the
entry chunk removed that marker from everything appended after the provider and
failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES
literal that follows the provider in the same module; string literals survive
minification, and every existing assertion holds against the tighter segment.

* fix(architecture): resolve the inspector gate's SPA path through crate_path

The gate joined a family-nested literal onto the workspace root, the idiom
crate_path exists to replace: a crate family move would have turned this into a
read failure rather than a resolved path. It now names the SPA file in the
logical flat spelling and resolves it, and the assertion reports the resolved
path so the message still points at a file that exists.

* test(inspector): follow a pinned turn explicitly when a new turn arrives

The multi-turn scenario assumed the panel would jump to an arriving turn, but a
selection the operator navigated to is deliberately sticky: the new turn widens
the window without yanking them off the turn they are reading. The scenario now
asserts that guarantee, then clicks Latest to follow, then walks back two turns
as before. Verified by running the inspector scenarios locally rather than by
reading, which is how this slipped through the first time.
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…earai#7291)

* feat(inspector): add operator inspection API

* docs(inspector): assign product service ownership

* test(inspector): ratchet diagnostic contracts

* feat(inspector): add debug panel shell

* test(inspector): cover debug panel shell e2e

* fix(inspector): stop diagnostics when panel closes

* feat(inspector): add prompt inspection

* fix(inspector): follow current webui ownership

* feat(inspector): add model call statistics

* test(inspector): cover model statistics e2e

* fix(inspector): avoid uncollected tool metrics

* test(inspector): cover prompt diagnostics e2e

* test(inspector): align statistics e2e scope

* fix(inspector): redact prompt metadata

* fix(inspector): preserve per-call model identity

* fix(inspector): classify prompt instruction sources

* test(inspector): assert reported token usage

* feat(inspector): add activity timeline and turn navigation

* test(inspector): cover activity timeline in browser

* fix(inspector): read current run before publishing activity

* feat(inspector): add bounded tool execution details

* test(inspector): cover bounded tool details in browser

* fix(inspector): validate retained tool result sizes

* test(inspector): add security and operator coverage

* test(inspector): cover browser workflows end to end

* fix(inspector): address review feedback

* fix(inspector): retry transient snapshot failures

* fix(inspector): address prompt diagnostic review findings

* fix(inspector): follow debug query navigation

* feat(inspector): complete frontend diagnostics

* test(inspector): cover frontend parity in browser

* fix(inspector): preserve stream terminal state

* fix(inspector): capture full capability surface

* fix(inspector): scope projection activity to its run

* fix(inspector): harden activity diagnostics

* fix(inspector): bound tool result diagnostic capture

* fix(inspector): harden tool diagnostic pipeline

* fix(llm): request usage for NEAR AI streams

* fix(inspector): address prompt diagnostic review feedback

* fix(webui): harden inspector stream coverage

* fix(inspector): preserve debug session statistics

* fix(inspector): keep diagnostics active while hidden

* test(e2e): cover hidden inspector observation

* fix inspector model call stats review findings

* fix inspector refresh and truncation regressions

* fix(inspector): address activity timeline review feedback

* fix(inspector): harden activity lifecycle handling

* fix(composition): move tool diagnostics to loop host

* fix(inspector): keep a settled stream live and complete locale parity

A live diagnostic update's debounced snapshot refresh was announcing LOADING,
so an open, healthy stream read as "Connecting" indefinitely once a run
settled — the settling stats update is the last one. That refresh is now a
background read. Incomplete snapshot statistics no longer accumulate as real
zeros, browser-session inspector state is namespaced by the authenticated
caller, an evicted pinned run rejoins the latest turn instead of the oldest,
tool status is localized, and the inspector strings now cover all ten locales.

* test(inspector): put the inspector locale sidecar under the parity gate

The inspector's English copy is registered from its lazy chunk instead of
src/i18n/en.ts, so the all-locale parity test — which derives the required key
set from en.ts — never covered those keys; a locale could drop one and fall
back to English silently. The test now treats the English key set as the union
of en.ts and a declared sidecar list. Keeping the copy in en.ts is not an
option: measured, it puts /chat at 217.4 KB gzip against a 217.0 KB budget.

* fix(inspector): reject malformed model breakdowns and correct locale copy

A `calls_per_model` entry with a negative or non-integer `calls` passed the
statistics decoder and was then coerced to zero during accumulation without
marking the breakdown truncated, presenting a fabricated "0 calls" for a model.
Every entry is now validated before a record is accepted. German turn
navigation used "Zug" (a train, or a game move); it now reads "Runde", with the
determiner agreement that noun requires. Spanish and Portuguese tool-status
values were written feminine against a masculine "Estado"/"Status" label.

* fix(inspector): bound the model breakdown before scanning and retaining it

The statistics decoder validated every calls_per_model entry but never the
array length, so an out-of-contract response was scanned in full and then
retained by the accumulator for up to 128 runs. The host truncates this
breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer
array cannot conform; the client now mirrors that ceiling and rejects the
record before the scan.

* fix(inspector): align turn navigation with host diagnostic retention

The browser offered 32 turns of navigation per thread while the host retained
diagnostics for 2 runs per session, so every turn past the second rendered
blank. Each layer was individually correct and the e2e scenario stopped at two
turns, so nothing saw the dead zone. Retention moves to 4 and the navigation
window mirrors it, pinned by a new architecture gate that reads both constants;
the scenario now walks back two turns and asserts real activity. Retention is a
ceiling as well as a default, and capture is unconditional, so 4 is a resident
memory choice — roughly 80 MB worst case across the eight tracked sessions.

* fix(composition): delimit the i18n bundle guard with an i18n-owned marker

The guard sliced the concatenated chunk bundle from the i18n provider up to
`QueryClient`, a symbol another module owns, so the segment's extent tracked
Rollup's chunk boundaries. A split that merely folded react-query into the
entry chunk removed that marker from everything appended after the provider and
failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES
literal that follows the provider in the same module; string literals survive
minification, and every existing assertion holds against the tighter segment.

* fix(architecture): resolve the inspector gate's SPA path through crate_path

The gate joined a family-nested literal onto the workspace root, the idiom
crate_path exists to replace: a crate family move would have turned this into a
read failure rather than a resolved path. It now names the SPA file in the
logical flat spelling and resolves it, and the assertion reports the resolved
path so the message still points at a file that exists.

* test(inspector): follow a pinned turn explicitly when a new turn arrives

The multi-turn scenario assumed the panel would jump to an arriving turn, but a
selection the operator navigated to is deliberately sticky: the new turn widens
the window without yanking them off the turn they are reading. The scenario now
asserts that guarantee, then clicks Latest to follow, then walks back two turns
as before. Verified by running the inspector scenarios locally rather than by
reading, which is how this slipped through the first time.
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…earai#7291)

* feat(inspector): add operator inspection API

* docs(inspector): assign product service ownership

* test(inspector): ratchet diagnostic contracts

* feat(inspector): add debug panel shell

* test(inspector): cover debug panel shell e2e

* fix(inspector): stop diagnostics when panel closes

* feat(inspector): add prompt inspection

* fix(inspector): follow current webui ownership

* feat(inspector): add model call statistics

* test(inspector): cover model statistics e2e

* fix(inspector): avoid uncollected tool metrics

* test(inspector): cover prompt diagnostics e2e

* test(inspector): align statistics e2e scope

* fix(inspector): redact prompt metadata

* fix(inspector): preserve per-call model identity

* fix(inspector): classify prompt instruction sources

* test(inspector): assert reported token usage

* feat(inspector): add activity timeline and turn navigation

* test(inspector): cover activity timeline in browser

* fix(inspector): read current run before publishing activity

* feat(inspector): add bounded tool execution details

* test(inspector): cover bounded tool details in browser

* fix(inspector): validate retained tool result sizes

* test(inspector): add security and operator coverage

* test(inspector): cover browser workflows end to end

* fix(inspector): address review feedback

* fix(inspector): retry transient snapshot failures

* fix(inspector): address prompt diagnostic review findings

* fix(inspector): follow debug query navigation

* feat(inspector): complete frontend diagnostics

* test(inspector): cover frontend parity in browser

* fix(inspector): preserve stream terminal state

* fix(inspector): capture full capability surface

* fix(inspector): scope projection activity to its run

* fix(inspector): harden activity diagnostics

* fix(inspector): bound tool result diagnostic capture

* fix(inspector): harden tool diagnostic pipeline

* fix(llm): request usage for NEAR AI streams

* fix(inspector): address prompt diagnostic review feedback

* fix(webui): harden inspector stream coverage

* fix(inspector): preserve debug session statistics

* fix(inspector): keep diagnostics active while hidden

* test(e2e): cover hidden inspector observation

* fix inspector model call stats review findings

* fix inspector refresh and truncation regressions

* fix(inspector): address activity timeline review feedback

* fix(inspector): harden activity lifecycle handling

* fix(composition): move tool diagnostics to loop host

* fix(inspector): keep a settled stream live and complete locale parity

A live diagnostic update's debounced snapshot refresh was announcing LOADING,
so an open, healthy stream read as "Connecting" indefinitely once a run
settled — the settling stats update is the last one. That refresh is now a
background read. Incomplete snapshot statistics no longer accumulate as real
zeros, browser-session inspector state is namespaced by the authenticated
caller, an evicted pinned run rejoins the latest turn instead of the oldest,
tool status is localized, and the inspector strings now cover all ten locales.

* test(inspector): put the inspector locale sidecar under the parity gate

The inspector's English copy is registered from its lazy chunk instead of
src/i18n/en.ts, so the all-locale parity test — which derives the required key
set from en.ts — never covered those keys; a locale could drop one and fall
back to English silently. The test now treats the English key set as the union
of en.ts and a declared sidecar list. Keeping the copy in en.ts is not an
option: measured, it puts /chat at 217.4 KB gzip against a 217.0 KB budget.

* fix(inspector): reject malformed model breakdowns and correct locale copy

A `calls_per_model` entry with a negative or non-integer `calls` passed the
statistics decoder and was then coerced to zero during accumulation without
marking the breakdown truncated, presenting a fabricated "0 calls" for a model.
Every entry is now validated before a record is accepted. German turn
navigation used "Zug" (a train, or a game move); it now reads "Runde", with the
determiner agreement that noun requires. Spanish and Portuguese tool-status
values were written feminine against a masculine "Estado"/"Status" label.

* fix(inspector): bound the model breakdown before scanning and retaining it

The statistics decoder validated every calls_per_model entry but never the
array length, so an out-of-contract response was scanned in full and then
retained by the accumulator for up to 128 runs. The host truncates this
breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer
array cannot conform; the client now mirrors that ceiling and rejects the
record before the scan.

* fix(inspector): align turn navigation with host diagnostic retention

The browser offered 32 turns of navigation per thread while the host retained
diagnostics for 2 runs per session, so every turn past the second rendered
blank. Each layer was individually correct and the e2e scenario stopped at two
turns, so nothing saw the dead zone. Retention moves to 4 and the navigation
window mirrors it, pinned by a new architecture gate that reads both constants;
the scenario now walks back two turns and asserts real activity. Retention is a
ceiling as well as a default, and capture is unconditional, so 4 is a resident
memory choice — roughly 80 MB worst case across the eight tracked sessions.

* fix(composition): delimit the i18n bundle guard with an i18n-owned marker

The guard sliced the concatenated chunk bundle from the i18n provider up to
`QueryClient`, a symbol another module owns, so the segment's extent tracked
Rollup's chunk boundaries. A split that merely folded react-query into the
entry chunk removed that marker from everything appended after the provider and
failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES
literal that follows the provider in the same module; string literals survive
minification, and every existing assertion holds against the tighter segment.

* fix(architecture): resolve the inspector gate's SPA path through crate_path

The gate joined a family-nested literal onto the workspace root, the idiom
crate_path exists to replace: a crate family move would have turned this into a
read failure rather than a resolved path. It now names the SPA file in the
logical flat spelling and resolves it, and the assertion reports the resolved
path so the message still points at a file that exists.

* test(inspector): follow a pinned turn explicitly when a new turn arrives

The multi-turn scenario assumed the panel would jump to an arriving turn, but a
selection the operator navigated to is deliberately sticky: the new turn widens
the window without yanking them off the turn they are reading. The scenario now
asserts that guarantee, then clicks Latest to follow, then walks back two turns
as before. Verified by running the inspector scenarios locally rather than by
reading, which is how this slipped through the first time.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7291 — 0f8cf14f Deployed Aug 8, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs human-verified Manually tested and verified risk: low Changes to docs, tests, or low-risk modules scope: dependencies Dependency updates scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Complete Web Inspector statistics, navigation, and localization

2 participants