Skip to content

test(inspector): add browser, security, and operator coverage - #7280

Merged
think-in-universe merged 63 commits into
mainfrom
issue-7226-inspector-coverage-docs
Aug 8, 2026
Merged

think-in-universe merged 63 commits into
mainfrom
issue-7226-inspector-coverage-docs

Conversation

@italic-jinxin

Copy link
Copy Markdown
Contributor

Summary

  • Adds security coverage for operator authorization, missing authenticated context, cross-scope isolation, invalid cursors, connection limits, and verbose-data stream exclusion.
  • Gives each browser tab a stable Inspector connection identity and monotonically increasing generation so reconnects replace stale streams without consuming additional capacity.
  • Adds browser coverage for activation, responsive layout, prompt inspection, activity ordering, turn navigation, statistics, tool details, 50 KiB truncation, reconnect behavior, and chat safety.
  • Documents Inspector activation, security boundaries, retention limits, operational behavior, troubleshooting, and rollback.
  • Updates feature and testing documentation to reflect the completed Inspector surface.

Linked Issue

Closes #7226

Depends on #7225

Part of #7218

Validation

  • cargo fmt --all -- --check
  • Relevant clippy checks pass with warnings denied
  • WebUI route and security contracts pass
  • Assistant scope-isolation tests pass
  • Frontend typecheck passes
  • Full frontend suite passes: 1123 tests
  • Focused browser E2E passes
  • Full review completed

Test Strategy

User behavior:

  • Inspector mode remains opt-in and does not affect ordinary chat behavior.
  • Operators can reconnect without duplicate activity rows or leaking connection capacity.
  • Security failures remain fail-closed and do not expose diagnostic content.

Risk areas:

  • Browser
  • Security or permissions
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: authorization gates, missing context, scope isolation, cursor validation, stream replacement, and verbose-data exclusion.
  • Reborn integration: Not applicable; no new durable workflow or provider integration.
  • Recorded fixture: Not applicable; no provider payload changes.
  • Browser E2E: complete Inspector workflow, reconnect, navigation, tool details, and truncation.
  • Backend or runtime: authenticated route dispatch and bounded stream capacity.
  • Live canary: Not applicable; no external provider behavior changed.

What the tests prove:

  • Every Inspector route requires both operator gates and authenticated scope.
  • Prompt, activity, statistics, and tool details cannot cross tenant, user, thread, run, or invocation boundaries.
  • Normal activity streams never contain raw tool arguments or results.
  • Same-tab reconnects replace stale streams without consuming another connection slot.
  • Reconnect and replay do not duplicate lifecycle entries.

Security Impact

Strengthens verification of existing operator-only boundaries. The production stream change adds opaque browser connection identity and generation values; it does not weaken authentication or expose diagnostic content.

Database Impact

None.

Blast Radius

Limited to Inspector SSE connection management, operator route contracts, browser tests, and operator documentation.

Rollback Plan

Revert this PR to restore the previous Inspector stream lifecycle and remove the added coverage/documentation. The activity and tool-detail data contracts remain independently functional.

Review Follow-Through

Risk level: Medium. Review should focus on reconnect generation ordering, same-tab stream replacement, fail-closed authorization, and ensuring normal chat remains unaffected.


Review track: B

@italic-jinxin italic-jinxin added size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Aug 6, 2026
…o issue-7224-activity-timeline

# Conflicts:
#	crates/loop/ironclaw_loop_host/src/lib.rs
#	crates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rs
#	crates/product/ironclaw_assistant/src/inspector_store.rs
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts
#	tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
…imeline

# Conflicts:
#	crates/loop/ironclaw_loop_host/src/lib.rs
#	crates/loop/ironclaw_loop_host/tests/thread_loop_host_contract.rs
#	crates/product/ironclaw_assistant/src/inspector_store.rs
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.test.tsx
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts
#	tests/CLAUDE.md
#	tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
…to issue-7225-bounded-tool-details

# Conflicts:
#	crates/product/ironclaw_assistant/src/inspector_store.rs
…to issue-7225-bounded-tool-details

# Conflicts:
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.test.tsx
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.tsx
… into issue-7226-inspector-coverage-docs

# Conflicts:
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7280 August 7, 2026 15:30 Destroyed
@italic-jinxin
italic-jinxin changed the base branch from issue-7225-bounded-tool-details to main August 7, 2026 15:43
…coverage-docs

# Conflicts:
#	crates/app/ironclaw_composition/src/runtime.rs
#	crates/app/ironclaw_composition/src/runtime/capability_host/tests.rs
#	crates/contracts/ironclaw_product_contracts/src/inspector.rs
#	crates/loop/ironclaw_loop_host/src/lib.rs
#	crates/loop/ironclaw_loop_host/src/tool_diagnostics.rs
#	crates/product/ironclaw_assistant/src/inspector_store.rs
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.test.tsx
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/inspector-panel.tsx
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.test.tsx
#	crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts
#	tests/CLAUDE.md
#	tests/e2e/scenarios/test_reborn_webui_v2_tool_gates.py
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7280 August 8, 2026 10:49 Destroyed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/product/ironclaw_assistant/src/reborn_services/inspector.rs`:
- Around line 190-199: Strengthen the authorized snapshot assertion in the test
around snapshot() so it verifies the snapshot payload contains the recorded
“owner-only activity” or expected activity record, not merely that
payload["snapshot"] is an object. Keep this assertion before the
unauthorized-scope isolation checks.

In `@crates/product/ironclaw_webui/src/webui_v2/inspector.rs`:
- Around line 56-63: Update the WebUI request ingress around
stream_connection_id to distinguish absent metadata from invalid connection_id
values, returning HTTP 400 for malformed IDs before dispatch. Require
connection_generation only when a valid connection ID is present, likewise
returning 400 otherwise, and add route tests covering both invalid connection_id
and generation-without-valid-ID cases.

In `@tests/e2e/scenarios/test_reborn_webui_v2_smoke.py`:
- Line 529: Define the authentication headers in the test scope before the
second httpx.AsyncClient context is created, then pass that existing headers
value to the client so the multi-turn navigation and reconnect flow runs without
a NameError.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: a740d2a2-fee0-49ec-94f5-4d2b8cf8d91b

📥 Commits

Reviewing files that changed from the base of the PR and between 7be9b84 and 4a1282d.

📒 Files selected for processing (16)
  • FEATURE_PARITY.md
  • crates/contracts/ironclaw_product_contracts/src/inspector.rs
  • crates/product/ironclaw_assistant/src/reborn_services/inspector.rs
  • crates/product/ironclaw_webui/README.md
  • crates/product/ironclaw_webui/frontend/src/pages/chat/chat.inspector-navigation.test.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.test.tsx
  • crates/product/ironclaw_webui/frontend/src/pages/chat/inspector/useInspector.ts
  • crates/product/ironclaw_webui/src/webui_v2/inspector.rs
  • crates/product/ironclaw_webui/tests/webui_v2_inspector_contract.rs
  • docs/reborn/README.md
  • docs/reborn/contracts/web-debug-inspector.md
  • docs/using/webui.mdx
  • tests/CLAUDE.md
  • tests/e2e/CLAUDE.md
  • tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
  • tests/e2e/scenarios/test_reborn_webui_v2_tool_gates.py

Comment thread crates/product/ironclaw_assistant/src/reborn_services/inspector.rs Outdated
Comment thread crates/product/ironclaw_webui/src/webui_v2/inspector.rs Outdated
Comment thread tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7280 August 8, 2026 11:03 Destroyed
coderabbitai[bot]

This comment was marked as resolved.

@think-in-universe
think-in-universe added this pull request to the merge queue Aug 8, 2026
Merged via the queue into main with commit 9edf3fe Aug 8, 2026
45 checks passed
@think-in-universe
think-in-universe deleted the issue-7226-inspector-coverage-docs branch August 8, 2026 13:25
serrrfirat added a commit that referenced this pull request Aug 9, 2026
- Merge origin/main (#7377 run-acts-as-invoker, #7323, #7382, #6938,
  #7280, #7393, #7389, #7364, #7228, #7371, #7399).
- main's #7377 landed a narrower terminal arm (generic failure notice for
  TurnStatus::Failed only); keep the #6896 arm, which covers Failed and
  RecoveryRequired with sanitized per-category summaries plus Cancelled
  and the timeout grace path, and adapt to the Option<String>
  notice_discriminator main introduced.
- Re-seed the composition budget to the merged-tree measurement
  (40811 -> 40861, the run-failure settlement observer lands +50 governed
  LOC); the arch-test record moves with the manifest.
personal-upstream-sync Bot pushed a commit to theredspoon/ironclaw that referenced this pull request Aug 10, 2026
…rai#6896) (nearai#7131)

* fix(run_delivery): deliver triggered run failures to the creator (nearai#6896)

Scheduled/triggered runs that ended in Failed, Cancelled, or
RecoveryRequired produced no user-visible notification: the triggered
delivery driver minted notifications only for Completed /
BlockedApproval / BlockedAuth and recorded every other terminal status
as Skipped. A run that timed out before reaching an actionable state
only logged a warn and recorded Failed, leaving the creator in silence.

Delivery:
- triggered_notification_for_state now mints a FinalReplyReady
  notification for Failed and RecoveryRequired using the existing
  per-category failure summaries (reborn_failure_summary_for_category)
  over state.failure.category(), with a generic fallback when no
  category is present.
- Cancelled mints the same notification, preferring a failure-category
  summary when one is present and falling back to a fixed cancellation
  notice otherwise.
- The RunWaitTimedOut branch with no prior blocked marker now delivers
  the timeout notice as a terminal reply instead of recording Failed.
- The wildcard arm is replaced with explicit non-actionable statuses
  (Queued, Running, CancelRequested, BlockedResource,
  BlockedDependentRun, BlockedExternalTool) so a future status fails to
  compile rather than silently skipping.

Observer:
- TriggerFireSettlementObserver gains on_failed_fire_settled as a
  default no-op method, plus a TriggerFailedFireSettlement event
  carrying tenant/trigger/fire-slot/run-id/history-status. Noop and
  existing implementors keep compiling.
- The active-cleanup sweep fires on_failed_fire_settled when
  clear_active_fire succeeds with TriggerRunHistoryStatus::Error, so
  post-accept failures are observable for automation health. Ok,
  Running, and already-cleared fires do not fire the hook.

Tests:
- run_delivery_contract: Failed+model_error, Failed without category,
  Cancelled, and timeout-before-actionable all assert a Delivered
  outcome with the expected notice text and footer.
- worker tests: a terminal-Error active fire fires exactly one
  on_failed_fire_settled; a terminal-Ok active fire fires none.

The larger retry/redrive budget for failed post-accept fires
(retry_disposition has zero production callers) is intentionally left
for a follow-up; it is out of scope for this surgical delivery fix.

* style: cargo fmt the nearai#6896 delivery fix

* fix(triggers): address terminal delivery review feedback

* fix(assistant): drop unused UserId import after merge

* fix(run_delivery): address multi-agent review findings

- Extract shared terminal-notice helpers (final_reply_notice,
  outcome_for_delivery_failure, deliver_terminal_notice) so the
  timeout, OAuth-backstop, and generic failure arms share one notice
  shape and outcome taxonomy instead of a third hand-rolled copy.
- Add a bounded race-grace window after the wait backstop: a run that
  crosses into a terminal state during the final wait (cancellation in
  flight, failure landing after the last poll) now delivers the correct
  terminal notice instead of the timeout copy.
- Cancelled runs always deliver the fixed cancellation notice; the
  failure-category branch was unreachable in production and would have
  mislabeled a host/operator cancel as a failure.
- Update the stale invariant doc, the five-output surface contract
  count, and the exhaustiveness-only comment on the non-actionable arm.
- Document the cheap/non-blocking contract on
  TriggerFireSettlementObserver (the worker awaits it inline in the
  poller sweep) and note it at the active-cleanup call site.
- Add contract coverage for the timeout arm's delivery-failure outcome
  (Failed) and a regression test proving the race-grace path delivers
  the cancellation notice; the cancelled-with-category test now asserts
  the cancellation notice wins.

* fix(run_delivery): address review comments and restore CI gates

Review fixes (CodeRabbit on 01e887f/f8af109):
- Grace loop fails loud: log the bound TurnError on state-poll failure and
  the RunDeliveryError on terminal-notice build failure before falling back
  to the timeout copy, with silent-ok markers on both intentional fallbacks.
- Hoist TriggeredReplyTargetAuthority, CodecChannelTargetResolver, and
  TriggeredNotificationContext to one construction before the watcher loop;
  the race-grace arm, timeout arm, and loop body now share it.
- Collapse the duplicated failure-summary expression into one closure and
  name TurnStatus::Failed explicitly so future statuses are compiler-visible.
- Drop the stale "Only three states" count from the surface-contract doc.
- Test fixture: encode the late-terminal flip as one Option<(usize,
  ScriptedRunState)> field instead of two correlated Options with an expect.
- Terminal-crossing test: document why flip_after=30 deterministically
  outruns the wait poll budget and assert the grace loop issues no
  cancellation (cancel_calls == 0).

CI:
- composition-budget: re-seed loc_ceiling 40432 -> 40593 (measured on the
  merged tree; the nearai#7131 settlement observer adds +161 governed LOC of
  wiring) and move the arch-test record with it.
- trigger_poller: use the colon-form tracing target required by nearai#7146.

* ci: re-trigger pull_request workflows for c2460ed

* fix(composition): capture the settlement health warn in the observer test

The traced_test default filter is {crate}=trace, which drops events whose
metadata target is `ironclaw::reborn::…`. The observer warning is emitted
with the colon-form target (required by nearai#7146 — the equals form recorded a
field and never matched RUST_LOG target filters), so the test saw an empty
buffer. Enable tracing-test's no-env-filter feature, the same pattern the
capabilities/host-runtime/mcp/loop crates use for cross-target assertions.

Re-seed the composition budget to the merged-tree measurement (40747 ->
40867): nearai#7131's observer wiring lands on top of post-measurement mainline
inflow; measured with the gate, set to current. The arch-test record moves
with the manifest.

* fix(run_delivery): merge main and adapt to notice_discriminator String

- Merge origin/main (nearai#7377 run-acts-as-invoker, nearai#7323, nearai#7382, nearai#6938,
  nearai#7280, nearai#7393, nearai#7389, nearai#7364, nearai#7228, nearai#7371, nearai#7399).
- main's nearai#7377 landed a narrower terminal arm (generic failure notice for
  TurnStatus::Failed only); keep the nearai#6896 arm, which covers Failed and
  RecoveryRequired with sanitized per-category summaries plus Cancelled
  and the timeout grace path, and adapt to the Option<String>
  notice_discriminator main introduced.
- Re-seed the composition budget to the merged-tree measurement
  (40811 -> 40861, the run-failure settlement observer lands +50 governed
  LOC); the arch-test record moves with the manifest.
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…#7280)

* feat(inspector): add operator inspection API

* docs(inspector): assign product service ownership

* test(inspector): ratchet diagnostic contracts

* feat(inspector): add debug panel shell

* test(inspector): cover debug panel shell e2e

* fix(inspector): stop diagnostics when panel closes

* feat(inspector): add prompt inspection

* fix(inspector): follow current webui ownership

* feat(inspector): add model call statistics

* test(inspector): cover model statistics e2e

* fix(inspector): avoid uncollected tool metrics

* test(inspector): cover prompt diagnostics e2e

* test(inspector): align statistics e2e scope

* fix(inspector): redact prompt metadata

* fix(inspector): preserve per-call model identity

* fix(inspector): classify prompt instruction sources

* test(inspector): assert reported token usage

* feat(inspector): add activity timeline and turn navigation

* test(inspector): cover activity timeline in browser

* fix(inspector): read current run before publishing activity

* feat(inspector): add bounded tool execution details

* test(inspector): cover bounded tool details in browser

* fix(inspector): validate retained tool result sizes

* test(inspector): add security and operator coverage

* test(inspector): cover browser workflows end to end

* fix(inspector): address review feedback

* fix(inspector): retry transient snapshot failures

* fix(inspector): address prompt diagnostic review findings

* fix(inspector): follow debug query navigation

* fix(inspector): preserve stream terminal state

* fix(inspector): capture full capability surface

* fix(inspector): scope projection activity to its run

* fix(inspector): harden activity diagnostics

* fix(inspector): bound tool result diagnostic capture

* fix(inspector): harden tool diagnostic pipeline

* fix(inspector): address prompt diagnostic review feedback

* fix(webui): harden inspector stream coverage

* fix inspector model call stats review findings

* fix inspector refresh and truncation regressions

* fix(inspector): address activity timeline review feedback

* fix(inspector): harden activity lifecycle handling

* fix(composition): move tool diagnostics to loop host

* fix(inspector): address review findings
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…#7280)

* feat(inspector): add operator inspection API

* docs(inspector): assign product service ownership

* test(inspector): ratchet diagnostic contracts

* feat(inspector): add debug panel shell

* test(inspector): cover debug panel shell e2e

* fix(inspector): stop diagnostics when panel closes

* feat(inspector): add prompt inspection

* fix(inspector): follow current webui ownership

* feat(inspector): add model call statistics

* test(inspector): cover model statistics e2e

* fix(inspector): avoid uncollected tool metrics

* test(inspector): cover prompt diagnostics e2e

* test(inspector): align statistics e2e scope

* fix(inspector): redact prompt metadata

* fix(inspector): preserve per-call model identity

* fix(inspector): classify prompt instruction sources

* test(inspector): assert reported token usage

* feat(inspector): add activity timeline and turn navigation

* test(inspector): cover activity timeline in browser

* fix(inspector): read current run before publishing activity

* feat(inspector): add bounded tool execution details

* test(inspector): cover bounded tool details in browser

* fix(inspector): validate retained tool result sizes

* test(inspector): add security and operator coverage

* test(inspector): cover browser workflows end to end

* fix(inspector): address review feedback

* fix(inspector): retry transient snapshot failures

* fix(inspector): address prompt diagnostic review findings

* fix(inspector): follow debug query navigation

* fix(inspector): preserve stream terminal state

* fix(inspector): capture full capability surface

* fix(inspector): scope projection activity to its run

* fix(inspector): harden activity diagnostics

* fix(inspector): bound tool result diagnostic capture

* fix(inspector): harden tool diagnostic pipeline

* fix(inspector): address prompt diagnostic review feedback

* fix(webui): harden inspector stream coverage

* fix inspector model call stats review findings

* fix inspector refresh and truncation regressions

* fix(inspector): address activity timeline review feedback

* fix(inspector): harden activity lifecycle handling

* fix(composition): move tool diagnostics to loop host

* fix(inspector): address review findings
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…#7280)

* feat(inspector): add operator inspection API

* docs(inspector): assign product service ownership

* test(inspector): ratchet diagnostic contracts

* feat(inspector): add debug panel shell

* test(inspector): cover debug panel shell e2e

* fix(inspector): stop diagnostics when panel closes

* feat(inspector): add prompt inspection

* fix(inspector): follow current webui ownership

* feat(inspector): add model call statistics

* test(inspector): cover model statistics e2e

* fix(inspector): avoid uncollected tool metrics

* test(inspector): cover prompt diagnostics e2e

* test(inspector): align statistics e2e scope

* fix(inspector): redact prompt metadata

* fix(inspector): preserve per-call model identity

* fix(inspector): classify prompt instruction sources

* test(inspector): assert reported token usage

* feat(inspector): add activity timeline and turn navigation

* test(inspector): cover activity timeline in browser

* fix(inspector): read current run before publishing activity

* feat(inspector): add bounded tool execution details

* test(inspector): cover bounded tool details in browser

* fix(inspector): validate retained tool result sizes

* test(inspector): add security and operator coverage

* test(inspector): cover browser workflows end to end

* fix(inspector): address review feedback

* fix(inspector): retry transient snapshot failures

* fix(inspector): address prompt diagnostic review findings

* fix(inspector): follow debug query navigation

* fix(inspector): preserve stream terminal state

* fix(inspector): capture full capability surface

* fix(inspector): scope projection activity to its run

* fix(inspector): harden activity diagnostics

* fix(inspector): bound tool result diagnostic capture

* fix(inspector): harden tool diagnostic pipeline

* fix(inspector): address prompt diagnostic review feedback

* fix(webui): harden inspector stream coverage

* fix inspector model call stats review findings

* fix inspector refresh and truncation regressions

* fix(inspector): address activity timeline review feedback

* fix(inspector): harden activity lifecycle handling

* fix(composition): move tool diagnostics to loop host

* fix(inspector): address review findings
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…rai#6896) (nearai#7131)

* fix(run_delivery): deliver triggered run failures to the creator (nearai#6896)

Scheduled/triggered runs that ended in Failed, Cancelled, or
RecoveryRequired produced no user-visible notification: the triggered
delivery driver minted notifications only for Completed /
BlockedApproval / BlockedAuth and recorded every other terminal status
as Skipped. A run that timed out before reaching an actionable state
only logged a warn and recorded Failed, leaving the creator in silence.

Delivery:
- triggered_notification_for_state now mints a FinalReplyReady
  notification for Failed and RecoveryRequired using the existing
  per-category failure summaries (reborn_failure_summary_for_category)
  over state.failure.category(), with a generic fallback when no
  category is present.
- Cancelled mints the same notification, preferring a failure-category
  summary when one is present and falling back to a fixed cancellation
  notice otherwise.
- The RunWaitTimedOut branch with no prior blocked marker now delivers
  the timeout notice as a terminal reply instead of recording Failed.
- The wildcard arm is replaced with explicit non-actionable statuses
  (Queued, Running, CancelRequested, BlockedResource,
  BlockedDependentRun, BlockedExternalTool) so a future status fails to
  compile rather than silently skipping.

Observer:
- TriggerFireSettlementObserver gains on_failed_fire_settled as a
  default no-op method, plus a TriggerFailedFireSettlement event
  carrying tenant/trigger/fire-slot/run-id/history-status. Noop and
  existing implementors keep compiling.
- The active-cleanup sweep fires on_failed_fire_settled when
  clear_active_fire succeeds with TriggerRunHistoryStatus::Error, so
  post-accept failures are observable for automation health. Ok,
  Running, and already-cleared fires do not fire the hook.

Tests:
- run_delivery_contract: Failed+model_error, Failed without category,
  Cancelled, and timeout-before-actionable all assert a Delivered
  outcome with the expected notice text and footer.
- worker tests: a terminal-Error active fire fires exactly one
  on_failed_fire_settled; a terminal-Ok active fire fires none.

The larger retry/redrive budget for failed post-accept fires
(retry_disposition has zero production callers) is intentionally left
for a follow-up; it is out of scope for this surgical delivery fix.

* style: cargo fmt the nearai#6896 delivery fix

* fix(triggers): address terminal delivery review feedback

* fix(assistant): drop unused UserId import after merge

* fix(run_delivery): address multi-agent review findings

- Extract shared terminal-notice helpers (final_reply_notice,
  outcome_for_delivery_failure, deliver_terminal_notice) so the
  timeout, OAuth-backstop, and generic failure arms share one notice
  shape and outcome taxonomy instead of a third hand-rolled copy.
- Add a bounded race-grace window after the wait backstop: a run that
  crosses into a terminal state during the final wait (cancellation in
  flight, failure landing after the last poll) now delivers the correct
  terminal notice instead of the timeout copy.
- Cancelled runs always deliver the fixed cancellation notice; the
  failure-category branch was unreachable in production and would have
  mislabeled a host/operator cancel as a failure.
- Update the stale invariant doc, the five-output surface contract
  count, and the exhaustiveness-only comment on the non-actionable arm.
- Document the cheap/non-blocking contract on
  TriggerFireSettlementObserver (the worker awaits it inline in the
  poller sweep) and note it at the active-cleanup call site.
- Add contract coverage for the timeout arm's delivery-failure outcome
  (Failed) and a regression test proving the race-grace path delivers
  the cancellation notice; the cancelled-with-category test now asserts
  the cancellation notice wins.

* fix(run_delivery): address review comments and restore CI gates

Review fixes (CodeRabbit on 01e887f/f8af109):
- Grace loop fails loud: log the bound TurnError on state-poll failure and
  the RunDeliveryError on terminal-notice build failure before falling back
  to the timeout copy, with silent-ok markers on both intentional fallbacks.
- Hoist TriggeredReplyTargetAuthority, CodecChannelTargetResolver, and
  TriggeredNotificationContext to one construction before the watcher loop;
  the race-grace arm, timeout arm, and loop body now share it.
- Collapse the duplicated failure-summary expression into one closure and
  name TurnStatus::Failed explicitly so future statuses are compiler-visible.
- Drop the stale "Only three states" count from the surface-contract doc.
- Test fixture: encode the late-terminal flip as one Option<(usize,
  ScriptedRunState)> field instead of two correlated Options with an expect.
- Terminal-crossing test: document why flip_after=30 deterministically
  outruns the wait poll budget and assert the grace loop issues no
  cancellation (cancel_calls == 0).

CI:
- composition-budget: re-seed loc_ceiling 40432 -> 40593 (measured on the
  merged tree; the nearai#7131 settlement observer adds +161 governed LOC of
  wiring) and move the arch-test record with it.
- trigger_poller: use the colon-form tracing target required by nearai#7146.

* ci: re-trigger pull_request workflows for c2460ed

* fix(composition): capture the settlement health warn in the observer test

The traced_test default filter is {crate}=trace, which drops events whose
metadata target is `ironclaw::reborn::…`. The observer warning is emitted
with the colon-form target (required by nearai#7146 — the equals form recorded a
field and never matched RUST_LOG target filters), so the test saw an empty
buffer. Enable tracing-test's no-env-filter feature, the same pattern the
capabilities/host-runtime/mcp/loop crates use for cross-target assertions.

Re-seed the composition budget to the merged-tree measurement (40747 ->
40867): nearai#7131's observer wiring lands on top of post-measurement mainline
inflow; measured with the gate, set to current. The arch-test record moves
with the manifest.

* fix(run_delivery): merge main and adapt to notice_discriminator String

- Merge origin/main (nearai#7377 run-acts-as-invoker, nearai#7323, nearai#7382, nearai#6938,
  nearai#7280, nearai#7393, nearai#7389, nearai#7364, nearai#7228, nearai#7371, nearai#7399).
- main's nearai#7377 landed a narrower terminal arm (generic failure notice for
  TurnStatus::Failed only); keep the nearai#6896 arm, which covers Failed and
  RecoveryRequired with sanitized per-category summaries plus Cancelled
  and the timeout grace path, and adapt to the Option<String>
  notice_discriminator main introduced.
- Re-seed the composition budget to the merged-tree measurement
  (40811 -> 40861, the run-failure settlement observer lands +50 governed
  LOC); the arch-test record moves with the manifest.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7280 — 6caa8804 Deployed Aug 8, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Inspector] Add browser, security, and documentation coverage

2 participants