Skip to content

fix(reborn): surface real failure detail instead of generic "invalid_input" - #5338

Merged
think-in-universe merged 42 commits into
mainfrom
issue-5289-surface-capability-failure-detail
Jul 1, 2026
Merged

think-in-universe merged 42 commits into
mainfrom
issue-5289-surface-capability-failure-detail

Conversation

@italic-jinxin

@italic-jinxin italic-jinxin commented Jun 26, 2026 •

Copy link
Copy Markdown
Contributor
  • Fixes [Reborn] Run ends with generic "driver protocol error" after builtin.json invalid_input failure #5289 by surfacing user-friendly capability/tool failure details in Reborn WebUI instead of showing only stable error-kind tokens like invalid_input or operation_failed.
  • This is intentionally a cross-layer fix because the failing UI state can be produced through multiple paths:
    • refresh/history uses capability display previews;
    • live chat uses SSE/runtime capability activity events;
    • WebUI state merging can receive those frames out of order.
  • Carries sanitized failure summaries through those paths: failure preview staging, runtime events, event projections, event-stream/product outbound activity payloads, and WebUI tool cards.
  • Adds optional error_summary / error_detail fields where needed so live SSE and refresh/history render the same host-authored failure detail.
  • Preserves safety boundaries by using bounded host-authored summaries, dropping unsafe runtime event summaries, avoiding raw received tool input, and re-validating display text at product boundaries.
  • Prevents late bare error-kind activity frames from clobbering richer failure details already shown by display previews.
  • The larger file count is mostly schema/fixture fallout from adding optional activity error-detail fields plus caller-level regression tests across the live and history paths.
image image

Linked Issue

Closes #5289

Validation

  • cargo fmt --check
  • git diff --check
  • cargo test -p ironclaw_events
  • cargo test -p ironclaw_event_projections --test replay_projection_contract
  • cargo test -p ironclaw_event_streams --test event_stream_manager_contract
  • cargo test -p ironclaw_reborn capability_failed_milestone_projects_to_dispatch_failed --lib
  • cargo test -p ironclaw_reborn_composition runtime_stream --lib
  • cargo clippy -p ironclaw_reborn_composition --tests --all-features -- -D warnings
  • cargo clippy --all --benches --tests --examples --all-features - blocked locally before checking this branch by toolchain/dependency MSRV mismatch: local rustc 1.92.0; monty@0.0.18 requires rustc 1.95, and ruff_* crates require rustc 1.93

Security Impact

Yes. This changes user-visible failure rendering for tool/capability errors. The implementation only surfaces bounded, host-authored summaries; sanitizes runtime event summaries on construction, serialization, deserialization, and projection replay; omits raw tool input values from invalid-input summaries; and re-validates display text at product outbound boundaries before sending it to the browser.

Database Impact

No schema or migration changes.

Blast Radius

Limited to Reborn capability/tool failure reporting and WebUI activity rendering: dispatch failure summaries, loop capability failure preview staging, runtime event/projection activity payloads, event-stream DTOs, product outbound activity views, and Reborn WebUI tool-card state merging. Many touched files are test fixtures or schema literals updated to include optional error-detail fields.

Rollback Plan

Revert this PR to return failed tool activity cards to the prior behavior of showing only stable error-kind tokens or generic failure text. Existing events remain compatible because the new fields are optional and omitted when absent.


Review track: C

italic-jinxin and others added 2 commits June 26, 2026 21:01
…ies (#5289)

Terminal failures that reach the WebUI projection via the normal
loop-exit path carry a category from `LoopFailureKind::as_str()`
(e.g. `capability_protocol_error`). `reborn_failure_summary_for_category`
only mapped the driver-error and scheduler categories plus three loop
kinds, so the rest — `capability_protocol_error`, `model_error`,
`invalid_model_output`, `checkpoint_*`, `transcript_write_failed`,
`driver_bug`, `policy_denied`, `compaction_unavailable`, and
`driver_protocol_violation` — degraded to the generic
"The run failed before producing a reply." The LLM failure explainer,
fed only that generic fallback, then paraphrased it into the vague
"driver protocol error" the user saw, masking the real tool failure.

Map each loop-exit category to a specific, honest, user-facing summary.
This also improves the explainer's input, since the fallback it receives
now describes the actual failure stage.

Regression coverage:
- unit: `reborn_failure_summary_describes_capability_protocol_error` and
  `reborn_failure_summary_maps_loop_failure_categories_specifically`
  (no loop-exit category degrades to the generic fallback).
- caller-level: `webui_event_stream_projects_capability_protocol_error_summary`
  drives the category through the projection with no explainer wired and
  asserts the specific summary reaches the run-status item.
…#5289)

A failed capability (e.g. the `json` builtin returning `invalid_input`)
showed only the bare error-kind string in the WebUI Activity panel's
per-tool Error tab. The rich `CapabilityFailureDetail::InvalidInput`
field issues reached the model transcript but never the display-preview
path the UI renders: failures don't call `write_capability_result`, so
no preview record was staged and the projection fell back to
`failed_capability_display_preview`, which only renders the kind.

Stage a failure display-preview record (approach B), so the existing
rich projection path surfaces the detail:

- ironclaw_loop_support: `LoopCapabilityResultWriter` gains a default
  no-op `stage_capability_failure_preview`. `runtime_outcome_to_loop`
  renders a bounded, host-authored summary from the InvalidInput issues
  (`capability_failure_display_summary`) and stages it on the Failed
  arm. Only schema-derived fields (path/code/expected) are rendered;
  `received` (raw tool input) is deliberately omitted.
- ironclaw_reborn_composition: implement the writer hook on
  `LocalDevCapabilityIo` and `ProductLiveCapabilityIo`; add
  `CapabilityDisplayPreviewStore::record_failure_preview`, which mirrors
  the success path (title/input pulled from the staged input) and stores
  the rendered summary as the output, no result ref.
- frontend (history-messages.js): `toolCardFromPreview` now prefers the
  backend `output_summary`/`output_preview` over the bare error kind for
  the Error tab. Bundle rebuilt (static/dist/app.js).

Regression coverage:
- unit: `capability_failure_display_summary_renders_invalid_input_issues`
  (+ asserts `received` never leaks) and `_is_none_for_non_invalid_input`.
- caller-level: `capability_display_preview_uses_staged_failure_summary_over_bare_kind`
  drives the projection chokepoint and asserts the detailed summary
  reaches the preview view instead of "tool failed: <kind>".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@italic-jinxin italic-jinxin added size: L 200-499 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Jun 26, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 June 26, 2026 14:55 Destroyed
@github-actions github-actions Bot added risk: low Changes to docs, tests, or low-risk modules and removed risk: medium Business logic, config, or moderate-risk modules labels Jun 26, 2026
@coderabbitai

coderabbitai Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@italic-jinxin, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 12 seconds

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 349c038a-9ba7-4124-9266-b114b68aee6b

📥 Commits

Reviewing files that changed from the base of the PR and between 933e168 and 4915a16.

📒 Files selected for processing (1)
  • crates/ironclaw_event_projections/CLAUDE.md
📝 Walkthrough

Walkthrough

Capability failure details now propagate as sanitized summaries from runtime classification into milestones, display previews, projections, and WebUI rendering. Failed tool paths prefer staged or detailed text over bare kinds, with tests updated across the pipeline.

Changes

Capability failure detail propagation

Layer / File(s) Summary
Runtime summaries and safe text
crates/ironclaw_host_api/src/dispatch.rs, crates/ironclaw_host_runtime/src/production.rs, crates/ironclaw_turns/src/run_profile/host.rs, crates/ironclaw_events/src/runtime_event.rs, crates/ironclaw_events/src/lib.rs, crates/ironclaw_reborn/src/milestone_events.rs, crates/ironclaw_reborn_composition/src/failure_summary.rs, crates/ironclaw_reborn_composition/src/projection/tests/failure_explanation.rs, crates/ironclaw_agent_loop/src/executor/capabilities.rs, crates/ironclaw_agent_loop/src/executor/tests.rs, crates/ironclaw_loop_support/src/capability_port.rs, crates/ironclaw_threads/src/tool_result_reference.rs
Dispatch and runtime failures gain plain-language summaries; RuntimeEvent gains sanitized error_summary; LoopSafeSummary adds failure-summary constructors and broader secret-token rejection; loop-exit failure categories map to explicit summaries and tests pin the new strings.
Preview staging and milestone emission
crates/ironclaw_turns/src/run_profile/milestones.rs, crates/ironclaw_loop_support/src/capability_port.rs, crates/ironclaw_reborn/src/loop_driver_host/port_adapters.rs, crates/ironclaw_reborn_composition/src/product_live_adapters.rs, crates/ironclaw_reborn_composition/src/runtime/local_dev.rs, crates/ironclaw_reborn_composition/src/runtime/local_dev/tests/tests/display_preview.rs
CapabilityFailed milestones and progress events carry safe_summary; failed runtime outcomes stage display previews; live and local-dev adapters persist failure previews; durable preview appends accept explicit status; tests cover the new failure-preview path.
Projection and UI error detail
crates/ironclaw_event_streams/src/types.rs, crates/ironclaw_event_projections/src/lib.rs, crates/ironclaw_event_projections/src/runtime_projection.rs, crates/ironclaw_reborn_composition/src/projection/display_preview.rs, crates/ironclaw_reborn_composition/src/projection.rs, crates/ironclaw_reborn_composition/src/projection/live_progress.rs, crates/ironclaw_product_adapters/src/outbound.rs, crates/ironclaw_webui_v2/tests/webui_v2_handlers_contract.rs, crates/ironclaw_webui_v2/tests/webui_v2_schema_contract.rs, crates/ironclaw_event_streams/tests/event_stream_manager_contract/support/builders.rs, crates/ironclaw_reborn_composition/src/projection/tests/display_preview.rs, crates/ironclaw_reborn_composition/src/projection/tests/display_preview_runtime.rs, crates/ironclaw_reborn_composition/src/projection/tests/runtime_stream.rs, crates/ironclaw_reborn_composition/src/projection/tests/live_progress_stream.rs, crates/ironclaw_reborn/tests/loop_milestone_event_projection.rs, crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/history-messages.js, crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/tool-activity-state.js, crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/tool-activity-state.test.mjs, crates/ironclaw_event_projections/CLAUDE.md
Capability activity payloads gain error_detail; projections derive it from runtime events or staged previews; live updates and product views forward it; the chat UI prefers backend failure text and preserves detailed errors across merges.

Estimated code review effort: 5 (Critical) | ~120 minutes

Possibly related PRs

  • nearai/ironclaw#5140: Shares the InvalidInput capability-failure path that feeds the new preview summary rendering.
  • nearai/ironclaw#5230: Touches the same durable preview staging path used by stage_capability_failure_preview.
  • nearai/ironclaw#5207: Also changes failure-category to user-facing summary mapping in crates/ironclaw_reborn_composition/src/failure_summary.rs.

Suggested reviewers: think-in-universe

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Conventional Commits style is used and the title accurately summarizes the failure-detail surfacing change.
Description check ✅ Passed Mostly matches the template; core sections are present, but Change Type and the Reborn Trust-Boundary Checklist are incomplete.
Linked Issues check ✅ Passed The PR directly addresses #5289 by propagating sanitized failure details through previews, events, projections, and the WebUI.
Out of Scope Changes check ✅ Passed No unrelated functionality stands out; the schema, fixtures, tests, and docs all support the same failure-detail propagation path.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

gemini-code-assist[bot]

This comment was marked as resolved.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_loop_support/src/capability_port.rs`:
- Around line 264-277: The failure-preview staging hook in capability_port.rs is
only updating transient state and never persists a durable preview record,
unlike the success path in LocalDevCapabilityIo::write_capability_result. Update
stage_capability_failure_preview (or add a paired async persistence method with
the same capability display preview semantics) so failed capability invocations
also append a persisted capability_display_preview entry and survive
refresh/timeline replay.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: c7a2ec6f-3176-45c4-906b-6481b83b45a8

📥 Commits

Reviewing files that changed from the base of the PR and between 535c29c and 97f5024.

⛔ Files ignored due to path filters (1)
  • crates/ironclaw_webui_v2_static/static/dist/app.js is excluded by !**/dist/**
📒 Files selected for processing (8)
  • crates/ironclaw_loop_support/src/capability_port.rs
  • crates/ironclaw_reborn_composition/src/failure_summary.rs
  • crates/ironclaw_reborn_composition/src/product_live_adapters.rs
  • crates/ironclaw_reborn_composition/src/projection/display_preview.rs
  • crates/ironclaw_reborn_composition/src/projection/tests/display_preview.rs
  • crates/ironclaw_reborn_composition/src/projection/tests/failure_explanation.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev.rs
  • crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/history-messages.js

Comment thread crates/ironclaw_loop_support/src/capability_port.rs
@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

…ues (#5289)

The previous change only rendered a failure display preview for
`InvalidInput` failures carrying structured field issues. Builtin tools
like `json` report invalid_input with a descriptive message (e.g.
"invalid JSON: expected value at line 1 column 1") but no structured
issues, so `capability_failure_display_summary` returned `None`, nothing
was staged, and the per-tool preview fell back to the bare error kind.

Extend the helper: when there are no structured issues, surface the
failure's host-authored `safe_summary` (already sanitized) unless it is a
generic placeholder ("capability invocation failed" /
"capability authorization denied") that adds nothing over the kind.

Tests: replaced the non-invalid-input case with
`capability_failure_display_summary_uses_safe_summary_without_issues`
(asserts the json-style message is surfaced) and
`_is_none_for_generic_placeholder`.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@railway-app

railway-app Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5338 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 1, 2026 at 9:48 am

@italic-jinxin
italic-jinxin marked this pull request as draft June 26, 2026 15:23
italic-jinxin and others added 2 commits June 26, 2026 23:57
Three review findings on the per-tool failure-preview change:

- Reuse `ironclaw_host_api::truncate_capability_display_text` for UTF-8
  boundary truncation in `capability_failure_display_summary` instead of
  a hand-rolled helper (gemini-code-assist).
- Fix a TOCTOU between `record_failure_preview` and `prune_run`:
  acquire the pending and completed locks together and hold them across
  the remove-from-pending + insert-into-completed pair so a concurrent
  prune cannot interleave and leak an unprunable completed record. Lock
  order (pending before completed) matches every other site, so holding
  both cannot deadlock (gemini-code-assist).
- Persist the failure preview to the durable timeline, not just the
  in-memory store: `stage_capability_failure_preview` is now async, and
  `LocalDevCapabilityIo` appends a durable display-preview message with
  status `Failed`, mirroring the success path so the detail survives
  refresh/replay. `try_append_durable_display_preview` takes a status
  parameter. `ProductLiveCapabilityIo` stays in-memory, matching its own
  success path which does not persist previews durably (coderabbitai).

Test: `capability_io_writes_failure_display_preview_to_durable_history`
asserts the durable timeline carries a Failed-status preview with the
rendered summary.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…5289)

The live, in-progress per-tool activity card showed only the bare error
kind (e.g. "invalid_input") on failure. The display-preview store fix
covered the history/timeline path, but the live card is driven by a
different projection path: the `CapabilityFailed` loop milestone ->
`ThreadLiveProjectionItem::CapabilityActivity` -> `CapabilityActivityView`,
which carried `error_kind` only — the sanitized failure message never
reached it.

Plumb the host-authored sanitized failure summary additively through that
live path so the card shows the real reason (e.g. "invalid JSON: ..."):

- ironclaw_turns: add `safe_summary: Option<String>` to
  `LoopHostMilestoneKind::CapabilityFailed` and
  `LoopProgressEvent::CapabilityActivityFailed` (additive, serde-default).
- ironclaw_loop_support: `runtime_terminal_milestone` populates it from
  the `RuntimeCapabilityFailure` message on the model-visible failure arm;
  host/infra and gate-denied paths emit `None`.
- ironclaw_event_streams: add `error_detail` to the live
  `CapabilityActivity` projection item.
- ironclaw_product_adapters: add `error_detail` to `CapabilityActivityView`
  (+ input, wire ser/de) with the same bounded/sanitized boundary
  validation as the other display fields.
- ironclaw_reborn_composition: live progress sanitizes the milestone
  summary and maps it onto the view; the history/runtime-payload path
  keeps carrying detail via the separate CapabilityDisplayPreview.
- frontend (history-messages.js): `toolCardFromActivity` prefers
  `error_detail` over the bare kind. Bundle rebuilt.

The durable runtime event log still records only the failure kind (no
backend-detail persistence), per ironclaw_turns guardrails.

Regression: `webui_event_stream_projects_live_tool_failure` now drives a
`CapabilityFailed` milestone with a `safe_summary` and asserts the live
`CapabilityActivityView.error_detail` carries it end-to-end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 June 27, 2026 05:00 Destroyed
…egory tokens (#5289)

When a capability dispatch fails without a host-authored safe_summary,
the failure message fell back to the error's Display, which exposed the
stable redacted category token (e.g. "dispatch failed: InputEncode") in
the per-tool UI Error tab.

Add `human_summary()` to `RuntimeDispatchErrorKind` and
`DispatchFailureKind` (fixed host-authored sentences, no raw content) and
use it as the fallback in `sanitized_failure_message`'s dispatch arm. The
stable `as_str()` token stays the contract for routing/metrics/audit; only
the user-facing message changes. So "InputEncode" now reads "the tool
input could not be encoded".

Tests: `dispatch_failure_kind_human_summary_is_plain_language_not_category_token`;
updated the two production.rs tests that pinned the old token wording
(they now assert the human summary and still verify no raw backend string
or secret leaks).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 June 27, 2026 05:11 Destroyed
@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Jun 27, 2026
Code review flagged a race between the display-preview store's
remove-from-pending + insert-into-completed pair and `prune_run`: the two
locks were acquired independently, so a concurrent prune could interleave
and leak a completed preview record that is never pruned.

The failure path (`record_failure_preview`) was already fixed to hold both
locks; apply the same fix to the pre-existing success path
(`record_result_with_preview`) so the whole pattern is consistent. Both
sites and `prune_run` acquire pending-before-completed, so holding both
cannot deadlock.

(The other two review findings — non-durable failure staging, and a
hand-rolled char-boundary truncation helper — were already resolved in
earlier commits on this branch.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 June 27, 2026 05:20 Destroyed
@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 June 27, 2026 05:34 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 June 27, 2026 05:51 Destroyed
@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

@think-in-universe

Copy link
Copy Markdown
Collaborator

@claude review

@claude

This comment was marked as resolved.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 July 1, 2026 09:01 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_loop_support/src/capability_port.rs`:
- Around line 372-386: Reject sensitive markers in rendered issue fields:
capability_input_issue_display_text currently only filters
control/non-ASCII/delimiter characters, so sensitive names like secret_api_key,
api_key, password, and tool_input can still reach the WebUI preview. Update this
function to apply the same sensitive-marker class used by safe-summary
validation before returning display text, and ensure any structured issue
rendering path falls back to sanitized output instead of exposing these
identifiers.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: b3ce2ef6-28ed-471b-b76c-920d70574194

📥 Commits

Reviewing files that changed from the base of the PR and between e0baf96 and 9cd8e1c.

📒 Files selected for processing (10)
  • crates/ironclaw_agent_loop/src/executor/capabilities.rs
  • crates/ironclaw_events/src/runtime_event.rs
  • crates/ironclaw_host_api/src/dispatch.rs
  • crates/ironclaw_loop_support/src/capability_port.rs
  • crates/ironclaw_reborn_composition/src/projection/tests/runtime_stream.rs
  • crates/ironclaw_threads/src/tool_result_reference.rs
  • crates/ironclaw_turns/src/run_profile/host.rs
  • crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/history-messages.js
  • crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/tool-activity-state.js
  • crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/tool-activity-state.test.mjs
💤 Files with no reviewable changes (3)
  • crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/tool-activity-state.js
  • crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/history-messages.js
  • crates/ironclaw_webui_v2_static/static/js/pages/chat/lib/tool-activity-state.test.mjs

Comment thread crates/ironclaw_loop_support/src/capability_port.rs
@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

@claude

This comment was marked as resolved.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 July 1, 2026 09:18 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 July 1, 2026 09:18 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_host_runtime/src/production.rs (1)

425-451: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Validate scope before emitting scoped telemetry.

context.validate() runs only at Line 480, but the new latency events at Lines 446-451 and 463-468 already emit tenant_id, user_id, agent_id, project_id, thread_id, and invocation_id from context.resource_scope. The nearby comment explicitly notes malformed requests can forge this scope, so policy/trust-rejected requests can now poison cross-tenant observability labels.

Move the scope consistency validation before any scoped trace, or emit pre-validation rejection traces without request-derived scope labels. As per coding guidelines, crates/**/*.rs must “Preserve tenant/user/agent/project/mission/thread scope on authority, state, memory, process, network, outbound, resource, and event records”.

Also applies to: 463-468, 480-482

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_host_runtime/src/production.rs` around lines 425 - 451,
Validate the request scope before any telemetry that reads from
`context.resource_scope`. In `invoke_capability`, move `context.validate()`
ahead of the new latency/tracing calls, or change the rejection paths to emit
only unscoped pre-validation traces; do not attach `tenant_id`, `user_id`,
`agent_id`, `project_id`, `thread_id`, or `invocation_id` until the scope has
been verified. Keep the `enforce_runtime_policy` rejection path and the
`trace_capability_latency_ok` calls consistent with this ordering so malformed
requests cannot poison scoped observability labels.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_loop_support/src/capability_port.rs`:
- Around line 397-438: `contains_capability_input_issue_sensitive_marker` misses
hyphenated and camelCase secret markers, so values like x-api-key, accessToken,
auth_token, and toolInput can still slip through. Expand the forbidden-marker
checks and token matching in this helper to recognize these additional variants,
not just spaced or snake_case forms. Add table-driven test cases covering
hyphenated and camelCase inputs alongside the existing secret marker patterns to
verify they are detected and redacted.

---

Outside diff comments:
In `@crates/ironclaw_host_runtime/src/production.rs`:
- Around line 425-451: Validate the request scope before any telemetry that
reads from `context.resource_scope`. In `invoke_capability`, move
`context.validate()` ahead of the new latency/tracing calls, or change the
rejection paths to emit only unscoped pre-validation traces; do not attach
`tenant_id`, `user_id`, `agent_id`, `project_id`, `thread_id`, or
`invocation_id` until the scope has been verified. Keep the
`enforce_runtime_policy` rejection path and the `trace_capability_latency_ok`
calls consistent with this ordering so malformed requests cannot poison scoped
observability labels.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: d7e5cd33-5adf-4e78-ab8f-683c1b004f0e

📥 Commits

Reviewing files that changed from the base of the PR and between 9cd8e1c and 644e35c.

📒 Files selected for processing (2)
  • crates/ironclaw_host_runtime/src/production.rs
  • crates/ironclaw_loop_support/src/capability_port.rs

Comment thread crates/ironclaw_loop_support/src/capability_port.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 July 1, 2026 09:30 Destroyed
@github-actions github-actions Bot added the scope: docs Documentation label Jul 1, 2026
@italic-jinxin

Copy link
Copy Markdown
Contributor Author

@claude review

@claude

This comment was marked as resolved.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5338 July 1, 2026 09:42 Destroyed

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5338 — 4915a169 Deployed Jul 1, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs human-verified Manually tested and verified risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Reborn] Run ends with generic "driver protocol error" after builtin.json invalid_input failure

2 participants