feat(inspector): complete statistics, navigation, and localization - #7291
Conversation
…pector-panel-shell
…copy A `calls_per_model` entry with a negative or non-integer `calls` passed the statistics decoder and was then coerced to zero during accumulation without marking the breakdown truncated, presenting a fabricated "0 calls" for a model. Every entry is now validated before a record is accepted. German turn navigation used "Zug" (a train, or a game move); it now reads "Runde", with the determiner agreement that noun requires. Spanish and Portuguese tool-status values were written feminine against a masculine "Estado"/"Status" label.
…ng it The statistics decoder validated every calls_per_model entry but never the array length, so an out-of-contract response was scanned in full and then retained by the accumulator for up to 128 runs. The host truncates this breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer array cannot conform; the client now mirrors that ceiling and rejects the record before the scan.
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (5)
crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts (4)
178-200: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy liftDo not merge truncated model labels as stable identities.
The accumulator keys entries by
entry.model.content. Whenentry.model.truncatedis true, two distinct model names can share the same retained prefix and their calls are merged incorrectly.Use an authoritative model identifier, or mark truncated model breakdowns as non-mergeable and incomplete. Add a collision test.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts` around lines 178 - 200, Update the model aggregation logic around models and sourceModels so entries with model.truncated set are not merged using model.content as a stable key; use an authoritative identifier when available, otherwise treat the breakdown as non-mergeable and mark calls_per_model_truncated. Add a test covering distinct truncated model names sharing the same retained prefix and verify their calls are not incorrectly combined.
130-135: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy liftPreserve unavailable tool totals instead of fabricating zero values.
Lines 130-135 initialize optional tool totals to zero. Lines 162-170 add omitted source fields as zero. A valid older-host record without tool totals therefore becomes
0for all three tool metrics insnapshot(). Mixed records also lose the fact that some runs are unknown.Preserve these fields as unavailable. Track aggregate completeness when records mix omitted and present fields. Add an accumulator test for omitted tool totals.
The PR objective requires explicit missing or unavailable states instead of fabricated zero values.
Also applies to: 162-170, 266-273
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts` around lines 130 - 135, Update emptyStats, the source-field accumulation around snapshot(), and the accumulator logic around lines 266-273 so omitted tool totals remain unavailable rather than becoming zero. Track whether tool totals are present across records, preserving unknown state for entirely omitted fields and representing mixed present/omitted records as incomplete. Add an accumulator test covering records with omitted tool totals.
111-113: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy liftDo not evict deduplication keys while retaining carried totals.
After 1,153 distinct runs,
retiredRunIdsevicts the first run whilecarriedstill contains its statistics. A later historical-navigation replay is accepted and counted again. The aggregate then double-counts that run.While an ID remains retired, later authoritative snapshots are ignored. Keep exact run identity for the supported session, or define a different replay contract. Add a boundary test.
The PR objective requires bounded session statistics and latest-turn navigation without duplicate totals.
Also applies to: 240-261
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts` around lines 111 - 113, Update retiredRunIds and carried retention so a run whose statistics remain in carried can never be evicted from deduplication tracking; preserve exact run identity for the supported session, or atomically discard its carried totals under an explicitly defined replay contract. Keep bounded session statistics and latest-turn navigation without double-counting, and add a boundary test covering more than MAX_RETIRED_RUN_IDS distinct runs followed by replay of the earliest run.
67-108: 🔒 Security & Privacy | 🟡 Minor | ⚡ Quick winReject oversized
BoundedDiagnosticTextbefore retaining stats.
isBoundedDiagnosticTextcurrently only checks shape/type;content.original_bytesis never used. Apply the owning limit forSessionDiagnosticStats::modelbeforedecoder.record(...): reject emptycontentwhenoriginal_bytes > 0ororiginal_bytesexceedsBoundedDiagnosticText’s retained maximum, and enforce an existing field bound formodelsuch asDIAGNOSTIC_LABEL_MAX_BYTES.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts` around lines 67 - 108, Update decodeSessionDiagnosticStats to validate the model BoundedDiagnosticText before returning stats for decoder.record. Reject empty model content when original_bytes is positive, reject original_bytes above the retained BoundedDiagnosticText maximum, and enforce the existing DIAGNOSTIC_LABEL_MAX_BYTES bound on retained model content. Preserve valid model diagnostics and return null for violations.Source: Path instructions
crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.ts (1)
10-41: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winCover successful and failed tool-call totals.
stats()hardcodesfailed_tool_calls: 0, and the accumulator test checks onlytotal_tool_calls. A regression in successful or failed tool-call aggregation can pass.Parameterize failed calls and assert all three counters after the refreshed run.
The PR objective requires total, successful, and failed tool-call counts.
Also applies to: 44-84
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.ts` around lines 10 - 41, Update the test helper stats() to accept a failed-tool-call count and derive successful_tool_calls from total tools minus failed calls instead of hardcoding it. In the accumulator test covering the refreshed run, assert total_tool_calls, successful_tool_calls, and failed_tool_calls so aggregation regressions for each counter are detected.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.ts`:
- Around line 10-41: Update the test helper stats() to accept a failed-tool-call
count and derive successful_tool_calls from total tools minus failed calls
instead of hardcoding it. In the accumulator test covering the refreshed run,
assert total_tool_calls, successful_tool_calls, and failed_tool_calls so
aggregation regressions for each counter are detected.
In
`@crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts`:
- Around line 178-200: Update the model aggregation logic around models and
sourceModels so entries with model.truncated set are not merged using
model.content as a stable key; use an authoritative identifier when available,
otherwise treat the breakdown as non-mergeable and mark
calls_per_model_truncated. Add a test covering distinct truncated model names
sharing the same retained prefix and verify their calls are not incorrectly
combined.
- Around line 130-135: Update emptyStats, the source-field accumulation around
snapshot(), and the accumulator logic around lines 266-273 so omitted tool
totals remain unavailable rather than becoming zero. Track whether tool totals
are present across records, preserving unknown state for entirely omitted fields
and representing mixed present/omitted records as incomplete. Add an accumulator
test covering records with omitted tool totals.
- Around line 111-113: Update retiredRunIds and carried retention so a run whose
statistics remain in carried can never be evicted from deduplication tracking;
preserve exact run identity for the supported session, or atomically discard its
carried totals under an explicitly defined replay contract. Keep bounded session
statistics and latest-turn navigation without double-counting, and add a
boundary test covering more than MAX_RETIRED_RUN_IDS distinct runs followed by
replay of the earliest run.
- Around line 67-108: Update decodeSessionDiagnosticStats to validate the model
BoundedDiagnosticText before returning stats for decoder.record. Reject empty
model content when original_bytes is positive, reject original_bytes above the
retained BoundedDiagnosticText maximum, and enforce the existing
DIAGNOSTIC_LABEL_MAX_BYTES bound on retained model content. Preserve valid model
diagnostics and return null for violations.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 93250f3d-9135-47ee-a325-360a440bd616
📒 Files selected for processing (2)
crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.test.tscrates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-session-stats.ts
The browser offered 32 turns of navigation per thread while the host retained diagnostics for 2 runs per session, so every turn past the second rendered blank. Each layer was individually correct and the e2e scenario stopped at two turns, so nothing saw the dead zone. Retention moves to 4 and the navigation window mirrors it, pinned by a new architecture gate that reads both constants; the scenario now walks back two turns and asserts real activity. Retention is a ceiling as well as a default, and capture is unconditional, so 4 is a resident memory choice — roughly 80 MB worst case across the eight tracked sessions.
…rker The guard sliced the concatenated chunk bundle from the i18n provider up to `QueryClient`, a symbol another module owns, so the segment's extent tracked Rollup's chunk boundaries. A split that merely folded react-query into the entry chunk removed that marker from everything appended after the provider and failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES literal that follows the provider in the same module; string literals survive minification, and every existing assertion holds against the tighter segment.
…e_path The gate joined a family-nested literal onto the workspace root, the idiom crate_path exists to replace: a crate family move would have turned this into a read failure rather than a resolved path. It now names the SPA file in the logical flat spelling and resolves it, and the assertion reports the resolved path so the message still points at a file that exists.
The multi-turn scenario assumed the panel would jump to an arriving turn, but a selection the operator navigated to is deliberately sticky: the new turn widens the window without yanking them off the turn they are reading. The scenario now asserts that guarantee, then clicks Latest to follow, then walks back two turns as before. Verified by running the inspector scenarios locally rather than by reading, which is how this slipped through the first time.
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
tests/e2e/scenarios/test_reborn_webui_v2_smoke.py (2)
690-699: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winRequire a positive update count before testing persistence.
updates_before_reloadcan be0. The equality check after reload then passes without proving that reconnect processing recorded any diagnostic update. Assertupdates_before_reload > 0before callingpage.reload().Proposed assertion
updates_before_reload = int((await updates.inner_text()).replace(",", "")) + assert updates_before_reload > 0 await page.reload()Based on the PR objective: “Showing ... diagnostic update count” and keeping counters within the current browser session.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py` around lines 690 - 699, In the smoke test around updates_before_reload, assert that the parsed update count is greater than zero before calling page.reload(). Keep the existing persistence assertion unchanged after the reload.
534-542: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winAssert exact tool-call counter values.
to_contain_text("0")checks the whole card. It also passes when the value is10or20. A regression that records tool calls can therefore pass this test. Select the value element and useto_have_text("0")for each counter.Proposed assertion
- await expect(stats.get_by_text("Tool calls", exact=True).locator("..")).to_contain_text( - "0" - ) + await expect( + stats.get_by_text("Tool calls", exact=True).locator("..").locator("p").nth(1) + ).to_have_text("0")Apply the same exact-value assertion to
Successful tool callsandFailed tool calls.Based on the PR objective: “Displaying total, successful, and failed tool-call counts.”
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py` around lines 534 - 542, Update the tool-call counter assertions in the smoke test to target each card’s value element rather than the whole card, then use exact text assertions expecting “0” for Tool calls, Successful tool calls, and Failed tool calls. Preserve the existing labels and counter coverage while preventing values such as 10 or 20 from matching.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py`:
- Around line 624-631: After clicking “Latest turn” in the activity navigation
flow, capture the completed third run’s ID and assert that
SEL_V2["inspector_panel"] contains that ID, alongside the existing “Turn 3 of 3”
assertion. Use the third-run identity established by the surrounding test rather
than reusing second_run_id.
---
Outside diff comments:
In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py`:
- Around line 690-699: In the smoke test around updates_before_reload, assert
that the parsed update count is greater than zero before calling page.reload().
Keep the existing persistence assertion unchanged after the reload.
- Around line 534-542: Update the tool-call counter assertions in the smoke test
to target each card’s value element rather than the whole card, then use exact
text assertions expecting “0” for Tool calls, Successful tool calls, and Failed
tool calls. Preserve the existing labels and counter coverage while preventing
values such as 10 or 20 from matching.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: c8a949a3-54ec-4bcd-b31f-3b64af8d3065
📒 Files selected for processing (1)
tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
| stream: true, | ||
| stream_options: None, | ||
| stream_options: Some(ChatCompletionStreamOptions { | ||
| include_usage: true, |
There was a problem hiding this comment.
Why is include_usage needed for all NEAR AI endpoints?
There was a problem hiding this comment.
We need the token counts back from streaming responses. NEAR's gateway only puts a usage block in the stream if you explicitly ask for it, and when it's absent our parser reads it as zero — so every streaming turn was quietly being recorded as 0 tokens, which then feeds cost, budget, resource accounting, and the usage we hand back to our own OpenAI-compat clients.
…earai#7291) * feat(inspector): add operator inspection API * docs(inspector): assign product service ownership * test(inspector): ratchet diagnostic contracts * feat(inspector): add debug panel shell * test(inspector): cover debug panel shell e2e * fix(inspector): stop diagnostics when panel closes * feat(inspector): add prompt inspection * fix(inspector): follow current webui ownership * feat(inspector): add model call statistics * test(inspector): cover model statistics e2e * fix(inspector): avoid uncollected tool metrics * test(inspector): cover prompt diagnostics e2e * test(inspector): align statistics e2e scope * fix(inspector): redact prompt metadata * fix(inspector): preserve per-call model identity * fix(inspector): classify prompt instruction sources * test(inspector): assert reported token usage * feat(inspector): add activity timeline and turn navigation * test(inspector): cover activity timeline in browser * fix(inspector): read current run before publishing activity * feat(inspector): add bounded tool execution details * test(inspector): cover bounded tool details in browser * fix(inspector): validate retained tool result sizes * test(inspector): add security and operator coverage * test(inspector): cover browser workflows end to end * fix(inspector): address review feedback * fix(inspector): retry transient snapshot failures * fix(inspector): address prompt diagnostic review findings * fix(inspector): follow debug query navigation * feat(inspector): complete frontend diagnostics * test(inspector): cover frontend parity in browser * fix(inspector): preserve stream terminal state * fix(inspector): capture full capability surface * fix(inspector): scope projection activity to its run * fix(inspector): harden activity diagnostics * fix(inspector): bound tool result diagnostic capture * fix(inspector): harden tool diagnostic pipeline * fix(llm): request usage for NEAR AI streams * fix(inspector): address prompt diagnostic review feedback * fix(webui): harden inspector stream coverage * fix(inspector): preserve debug session statistics * fix(inspector): keep diagnostics active while hidden * test(e2e): cover hidden inspector observation * fix inspector model call stats review findings * fix inspector refresh and truncation regressions * fix(inspector): address activity timeline review feedback * fix(inspector): harden activity lifecycle handling * fix(composition): move tool diagnostics to loop host * fix(inspector): keep a settled stream live and complete locale parity A live diagnostic update's debounced snapshot refresh was announcing LOADING, so an open, healthy stream read as "Connecting" indefinitely once a run settled — the settling stats update is the last one. That refresh is now a background read. Incomplete snapshot statistics no longer accumulate as real zeros, browser-session inspector state is namespaced by the authenticated caller, an evicted pinned run rejoins the latest turn instead of the oldest, tool status is localized, and the inspector strings now cover all ten locales. * test(inspector): put the inspector locale sidecar under the parity gate The inspector's English copy is registered from its lazy chunk instead of src/i18n/en.ts, so the all-locale parity test — which derives the required key set from en.ts — never covered those keys; a locale could drop one and fall back to English silently. The test now treats the English key set as the union of en.ts and a declared sidecar list. Keeping the copy in en.ts is not an option: measured, it puts /chat at 217.4 KB gzip against a 217.0 KB budget. * fix(inspector): reject malformed model breakdowns and correct locale copy A `calls_per_model` entry with a negative or non-integer `calls` passed the statistics decoder and was then coerced to zero during accumulation without marking the breakdown truncated, presenting a fabricated "0 calls" for a model. Every entry is now validated before a record is accepted. German turn navigation used "Zug" (a train, or a game move); it now reads "Runde", with the determiner agreement that noun requires. Spanish and Portuguese tool-status values were written feminine against a masculine "Estado"/"Status" label. * fix(inspector): bound the model breakdown before scanning and retaining it The statistics decoder validated every calls_per_model entry but never the array length, so an out-of-contract response was scanned in full and then retained by the accumulator for up to 128 runs. The host truncates this breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer array cannot conform; the client now mirrors that ceiling and rejects the record before the scan. * fix(inspector): align turn navigation with host diagnostic retention The browser offered 32 turns of navigation per thread while the host retained diagnostics for 2 runs per session, so every turn past the second rendered blank. Each layer was individually correct and the e2e scenario stopped at two turns, so nothing saw the dead zone. Retention moves to 4 and the navigation window mirrors it, pinned by a new architecture gate that reads both constants; the scenario now walks back two turns and asserts real activity. Retention is a ceiling as well as a default, and capture is unconditional, so 4 is a resident memory choice — roughly 80 MB worst case across the eight tracked sessions. * fix(composition): delimit the i18n bundle guard with an i18n-owned marker The guard sliced the concatenated chunk bundle from the i18n provider up to `QueryClient`, a symbol another module owns, so the segment's extent tracked Rollup's chunk boundaries. A split that merely folded react-query into the entry chunk removed that marker from everything appended after the provider and failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES literal that follows the provider in the same module; string literals survive minification, and every existing assertion holds against the tighter segment. * fix(architecture): resolve the inspector gate's SPA path through crate_path The gate joined a family-nested literal onto the workspace root, the idiom crate_path exists to replace: a crate family move would have turned this into a read failure rather than a resolved path. It now names the SPA file in the logical flat spelling and resolves it, and the assertion reports the resolved path so the message still points at a file that exists. * test(inspector): follow a pinned turn explicitly when a new turn arrives The multi-turn scenario assumed the panel would jump to an arriving turn, but a selection the operator navigated to is deliberately sticky: the new turn widens the window without yanking them off the turn they are reading. The scenario now asserts that guarantee, then clicks Latest to follow, then walks back two turns as before. Verified by running the inspector scenarios locally rather than by reading, which is how this slipped through the first time.
…earai#7291) * feat(inspector): add operator inspection API * docs(inspector): assign product service ownership * test(inspector): ratchet diagnostic contracts * feat(inspector): add debug panel shell * test(inspector): cover debug panel shell e2e * fix(inspector): stop diagnostics when panel closes * feat(inspector): add prompt inspection * fix(inspector): follow current webui ownership * feat(inspector): add model call statistics * test(inspector): cover model statistics e2e * fix(inspector): avoid uncollected tool metrics * test(inspector): cover prompt diagnostics e2e * test(inspector): align statistics e2e scope * fix(inspector): redact prompt metadata * fix(inspector): preserve per-call model identity * fix(inspector): classify prompt instruction sources * test(inspector): assert reported token usage * feat(inspector): add activity timeline and turn navigation * test(inspector): cover activity timeline in browser * fix(inspector): read current run before publishing activity * feat(inspector): add bounded tool execution details * test(inspector): cover bounded tool details in browser * fix(inspector): validate retained tool result sizes * test(inspector): add security and operator coverage * test(inspector): cover browser workflows end to end * fix(inspector): address review feedback * fix(inspector): retry transient snapshot failures * fix(inspector): address prompt diagnostic review findings * fix(inspector): follow debug query navigation * feat(inspector): complete frontend diagnostics * test(inspector): cover frontend parity in browser * fix(inspector): preserve stream terminal state * fix(inspector): capture full capability surface * fix(inspector): scope projection activity to its run * fix(inspector): harden activity diagnostics * fix(inspector): bound tool result diagnostic capture * fix(inspector): harden tool diagnostic pipeline * fix(llm): request usage for NEAR AI streams * fix(inspector): address prompt diagnostic review feedback * fix(webui): harden inspector stream coverage * fix(inspector): preserve debug session statistics * fix(inspector): keep diagnostics active while hidden * test(e2e): cover hidden inspector observation * fix inspector model call stats review findings * fix inspector refresh and truncation regressions * fix(inspector): address activity timeline review feedback * fix(inspector): harden activity lifecycle handling * fix(composition): move tool diagnostics to loop host * fix(inspector): keep a settled stream live and complete locale parity A live diagnostic update's debounced snapshot refresh was announcing LOADING, so an open, healthy stream read as "Connecting" indefinitely once a run settled — the settling stats update is the last one. That refresh is now a background read. Incomplete snapshot statistics no longer accumulate as real zeros, browser-session inspector state is namespaced by the authenticated caller, an evicted pinned run rejoins the latest turn instead of the oldest, tool status is localized, and the inspector strings now cover all ten locales. * test(inspector): put the inspector locale sidecar under the parity gate The inspector's English copy is registered from its lazy chunk instead of src/i18n/en.ts, so the all-locale parity test — which derives the required key set from en.ts — never covered those keys; a locale could drop one and fall back to English silently. The test now treats the English key set as the union of en.ts and a declared sidecar list. Keeping the copy in en.ts is not an option: measured, it puts /chat at 217.4 KB gzip against a 217.0 KB budget. * fix(inspector): reject malformed model breakdowns and correct locale copy A `calls_per_model` entry with a negative or non-integer `calls` passed the statistics decoder and was then coerced to zero during accumulation without marking the breakdown truncated, presenting a fabricated "0 calls" for a model. Every entry is now validated before a record is accepted. German turn navigation used "Zug" (a train, or a game move); it now reads "Runde", with the determiner agreement that noun requires. Spanish and Portuguese tool-status values were written feminine against a masculine "Estado"/"Status" label. * fix(inspector): bound the model breakdown before scanning and retaining it The statistics decoder validated every calls_per_model entry but never the array length, so an out-of-contract response was scanned in full and then retained by the accumulator for up to 128 runs. The host truncates this breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer array cannot conform; the client now mirrors that ceiling and rejects the record before the scan. * fix(inspector): align turn navigation with host diagnostic retention The browser offered 32 turns of navigation per thread while the host retained diagnostics for 2 runs per session, so every turn past the second rendered blank. Each layer was individually correct and the e2e scenario stopped at two turns, so nothing saw the dead zone. Retention moves to 4 and the navigation window mirrors it, pinned by a new architecture gate that reads both constants; the scenario now walks back two turns and asserts real activity. Retention is a ceiling as well as a default, and capture is unconditional, so 4 is a resident memory choice — roughly 80 MB worst case across the eight tracked sessions. * fix(composition): delimit the i18n bundle guard with an i18n-owned marker The guard sliced the concatenated chunk bundle from the i18n provider up to `QueryClient`, a symbol another module owns, so the segment's extent tracked Rollup's chunk boundaries. A split that merely folded react-query into the entry chunk removed that marker from everything appended after the provider and failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES literal that follows the provider in the same module; string literals survive minification, and every existing assertion holds against the tighter segment. * fix(architecture): resolve the inspector gate's SPA path through crate_path The gate joined a family-nested literal onto the workspace root, the idiom crate_path exists to replace: a crate family move would have turned this into a read failure rather than a resolved path. It now names the SPA file in the logical flat spelling and resolves it, and the assertion reports the resolved path so the message still points at a file that exists. * test(inspector): follow a pinned turn explicitly when a new turn arrives The multi-turn scenario assumed the panel would jump to an arriving turn, but a selection the operator navigated to is deliberately sticky: the new turn widens the window without yanking them off the turn they are reading. The scenario now asserts that guarantee, then clicks Latest to follow, then walks back two turns as before. Verified by running the inspector scenarios locally rather than by reading, which is how this slipped through the first time.
…earai#7291) * feat(inspector): add operator inspection API * docs(inspector): assign product service ownership * test(inspector): ratchet diagnostic contracts * feat(inspector): add debug panel shell * test(inspector): cover debug panel shell e2e * fix(inspector): stop diagnostics when panel closes * feat(inspector): add prompt inspection * fix(inspector): follow current webui ownership * feat(inspector): add model call statistics * test(inspector): cover model statistics e2e * fix(inspector): avoid uncollected tool metrics * test(inspector): cover prompt diagnostics e2e * test(inspector): align statistics e2e scope * fix(inspector): redact prompt metadata * fix(inspector): preserve per-call model identity * fix(inspector): classify prompt instruction sources * test(inspector): assert reported token usage * feat(inspector): add activity timeline and turn navigation * test(inspector): cover activity timeline in browser * fix(inspector): read current run before publishing activity * feat(inspector): add bounded tool execution details * test(inspector): cover bounded tool details in browser * fix(inspector): validate retained tool result sizes * test(inspector): add security and operator coverage * test(inspector): cover browser workflows end to end * fix(inspector): address review feedback * fix(inspector): retry transient snapshot failures * fix(inspector): address prompt diagnostic review findings * fix(inspector): follow debug query navigation * feat(inspector): complete frontend diagnostics * test(inspector): cover frontend parity in browser * fix(inspector): preserve stream terminal state * fix(inspector): capture full capability surface * fix(inspector): scope projection activity to its run * fix(inspector): harden activity diagnostics * fix(inspector): bound tool result diagnostic capture * fix(inspector): harden tool diagnostic pipeline * fix(llm): request usage for NEAR AI streams * fix(inspector): address prompt diagnostic review feedback * fix(webui): harden inspector stream coverage * fix(inspector): preserve debug session statistics * fix(inspector): keep diagnostics active while hidden * test(e2e): cover hidden inspector observation * fix inspector model call stats review findings * fix inspector refresh and truncation regressions * fix(inspector): address activity timeline review feedback * fix(inspector): harden activity lifecycle handling * fix(composition): move tool diagnostics to loop host * fix(inspector): keep a settled stream live and complete locale parity A live diagnostic update's debounced snapshot refresh was announcing LOADING, so an open, healthy stream read as "Connecting" indefinitely once a run settled — the settling stats update is the last one. That refresh is now a background read. Incomplete snapshot statistics no longer accumulate as real zeros, browser-session inspector state is namespaced by the authenticated caller, an evicted pinned run rejoins the latest turn instead of the oldest, tool status is localized, and the inspector strings now cover all ten locales. * test(inspector): put the inspector locale sidecar under the parity gate The inspector's English copy is registered from its lazy chunk instead of src/i18n/en.ts, so the all-locale parity test — which derives the required key set from en.ts — never covered those keys; a locale could drop one and fall back to English silently. The test now treats the English key set as the union of en.ts and a declared sidecar list. Keeping the copy in en.ts is not an option: measured, it puts /chat at 217.4 KB gzip against a 217.0 KB budget. * fix(inspector): reject malformed model breakdowns and correct locale copy A `calls_per_model` entry with a negative or non-integer `calls` passed the statistics decoder and was then coerced to zero during accumulation without marking the breakdown truncated, presenting a fabricated "0 calls" for a model. Every entry is now validated before a record is accepted. German turn navigation used "Zug" (a train, or a game move); it now reads "Runde", with the determiner agreement that noun requires. Spanish and Portuguese tool-status values were written feminine against a masculine "Estado"/"Status" label. * fix(inspector): bound the model breakdown before scanning and retaining it The statistics decoder validated every calls_per_model entry but never the array length, so an out-of-contract response was scanned in full and then retained by the accumulator for up to 128 runs. The host truncates this breakdown at MAX_MODELS_IN_STATS and reports it as truncated, so a longer array cannot conform; the client now mirrors that ceiling and rejects the record before the scan. * fix(inspector): align turn navigation with host diagnostic retention The browser offered 32 turns of navigation per thread while the host retained diagnostics for 2 runs per session, so every turn past the second rendered blank. Each layer was individually correct and the e2e scenario stopped at two turns, so nothing saw the dead zone. Retention moves to 4 and the navigation window mirrors it, pinned by a new architecture gate that reads both constants; the scenario now walks back two turns and asserts real activity. Retention is a ceiling as well as a default, and capture is unconditional, so 4 is a resident memory choice — roughly 80 MB worst case across the eight tracked sessions. * fix(composition): delimit the i18n bundle guard with an i18n-owned marker The guard sliced the concatenated chunk bundle from the i18n provider up to `QueryClient`, a symbol another module owns, so the segment's extent tracked Rollup's chunk boundaries. A split that merely folded react-query into the entry chunk removed that marker from everything appended after the provider and failed an i18n guard with no i18n change. It now ends on the AVAILABLE_LANGUAGES literal that follows the provider in the same module; string literals survive minification, and every existing assertion holds against the tighter segment. * fix(architecture): resolve the inspector gate's SPA path through crate_path The gate joined a family-nested literal onto the workspace root, the idiom crate_path exists to replace: a crate family move would have turned this into a read failure rather than a resolved path. It now names the SPA file in the logical flat spelling and resolves it, and the assertion reports the resolved path so the message still points at a file that exists. * test(inspector): follow a pinned turn explicitly when a new turn arrives The multi-turn scenario assumed the panel would jump to an arriving turn, but a selection the operator navigated to is deliberately sticky: the new turn widens the window without yanking them off the turn they are reading. The scenario now asserts that guarantee, then clicks Latest to follow, then walks back two turns as before. Verified by running the inspector scenarios locally rather than by reading, which is how this slipped through the first time.
Summary
Linked Issue
Closes #7287
Depends on #7226
Part of #7218
Validation
git diff --checkThe existing bundle-budget check remains above its configured threshold at
218.7 KB on both this branch and the unchanged dependency branch. The Inspector
changes add no gzip bytes to the ordinary
/chatinitial bundle because theyremain lazy-loaded.
Test Strategy
User behavior:
Operators using
debug=truecan inspect prompt composition, ordered runactivity, model and tool statistics, previous/latest turns, and browser stream
health. Ordinary chat remains unchanged when diagnostics are disabled.
Risk areas:
Tests added or updated:
counters, cursor deduplication, and translated accessibility labels.
metrics, localization, responsive layout, and disabled-debug behavior.
the changed contract.
What the tests prove:
The Inspector displays host-provided diagnostics without fabricating unavailable
values, keeps browser counters bounded and duplicate-safe, preserves chat during
stream reconnects, and leaves ordinary chat unaffected.
Commands run:
TZ=America/Los_AngelesSKIP_FRONTEND_BUILD=1 cargo test -p ironclaw_webui --test i18n_consistencySecurity Impact
None. Diagnostic authorization, redaction, retention, and backend contracts are
unchanged. Browser-observed counters contain no diagnostic payloads and use
bounded session storage.
Database Impact
None. No migrations, schemas, or persistent storage behavior changed.
Blast Radius
Limited to the opt-in Web Inspector frontend, its localized strings,
documentation, and browser coverage. The ordinary chat path remains lazy and
unchanged when
debug=trueis absent.Rollback Plan
Revert the two commits to restore the previous Inspector presentation and E2E
coverage. No data migration or backend rollback is required.
Review Follow-Through
Full review found no blocking issues. Remaining limitations—provider cost
accounting, raw reasoning, and persistent diagnostics—are explicitly outside
the issue scope.
Review track: B