Skip to content

Expose outbound delivery targets to Reborn model - #4779

Merged
serrrfirat merged 17 commits into
mainfrom
codex/channel-manifest-surfaces
Jun 15, 2026
Merged

serrrfirat merged 17 commits into
mainfrom
codex/channel-manifest-surfaces

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jun 11, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • build on [codex] Represent Slack as a product-adapter extension #4778's Slack-as-product-adapter extension surface by adding the runtime/model layer for outbound delivery target selection
  • add local-dev model-visible capabilities to list outbound delivery targets and set the caller's default final-reply target, with the setter going through the normal approval path outside local-dev-yolo
  • add a mutable Reborn runtime outbound target registry so post-build product adapters can register providers after runtime assembly
  • register Slack host-beta DM/shared-channel targets into that registry and use the same caller-scoped provider path from WebUI outbound preferences and model capabilities
  • persist triggered-run delivery outcomes through the local-dev outbound store so scheduled runs can deliver final replies to the configured target
  • address review feedback around provider error propagation, caller-sensitive target providers, approval consumption before preference persistence, and hiding provider tools when the outbound facade is absent

Stack / Reviewer Notes

  • [codex] Represent Slack as a product-adapter extension #4778 is merged and owns the Slack extension manifest/channel UI surface. This PR is the next layer: model/runtime selection of where final replies should be delivered.
  • The Slack extension card can still show No capabilities; Slack is a channel/product adapter, while the model-visible outbound target tools are synthetic runtime built-ins.
  • local-dev-yolo intentionally bypasses approval gates. Use non-yolo local-dev to exercise the approval prompt for changing the final-reply target.
  • Manual local testing exposed a separate local-dev SSO caveat: if a signed SSO session is reused after restart, local trigger access may not be reseeded for that user until re-login. Fresh OAuth login seeds it correctly; follow-up should seed local trigger access on successful session authentication as well.

Manual Slack E2E

  • Created a real Slack app from manifest, pointed Events API at local Reborn through ngrok, and paired a personal Slack DM target.
  • Created a scheduled automation: every 5 minutes, GET https://api.github.com with a User-Agent header and send the result as the final reply to Slack.
  • After seeding the local-dev trigger access row for the active SSO user, the trigger fired and delivered GitHub API check: HTTP 200 OK to the Slack DM.
  • Persisted evidence from the local run:
    • trigger_run_history.status = ok
    • triggered-run delivery outcome delivered
    • outbound delivery status delivered
    • delivered candidate kind final_reply

Verification

  • cargo check -p ironclaw_reborn_composition --tests
  • cargo clippy -p ironclaw_reborn_composition --tests -- -D warnings
  • targeted outbound/local-dev regression tests for provider errors, caller-scoped targets, approval/persistence ordering, and facade-absent tool hiding
  • bash scripts/pre-commit-safety.sh
  • GitHub checks on Expose outbound delivery targets to Reborn model #4779 are green: Code Style, Reborn E2E, Tests (Reborn), Tests (Legacy), and CodeRabbit

@github-actions github-actions Bot added scope: docs Documentation size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jun 11, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces outbound delivery target capabilities for local development, adding a MutableOutboundDeliveryTargetRegistry to dynamically register target providers and exposing synthetic capabilities to list and set delivery targets. It also refactors Slack integration to register its provider dynamically. The review feedback suggests simplifying mapping closures and match arms in runtime.rs for cleaner code, and correcting the return value of register_provider in outbound_preferences.rs to avoid misleading callers when overwriting an existing provider.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread crates/ironclaw_reborn_composition/src/runtime.rs Outdated
Comment thread crates/ironclaw_reborn_composition/src/runtime.rs Outdated
Comment thread crates/ironclaw_reborn_composition/src/outbound_preferences.rs Outdated

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed draft PR #4779 at 038903ba7ea303c77e8014ea0bedf9b7b34b9f58.

Existing live threads are Gemini style/nit comments, so I did not duplicate them. I found four current behavioral issues below.

Verification:

  • PASS: cargo test -p ironclaw_reborn_composition local_dev_outbound_delivery_capabilities_use_late_registered_targets -- --nocapture
  • PASS: cargo test -p ironclaw_reborn_composition --features slack-v2-host-beta slack_host_beta_targets_wire_through_outbound_preferences_facade -- --nocapture

Comment thread crates/ironclaw_reborn_composition/src/runtime/local_dev.rs Outdated
Comment thread crates/ironclaw_reborn_composition/src/webui.rs Outdated
Comment thread crates/ironclaw_reborn_composition/src/slack_host_beta.rs Outdated

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (re-review)

Re-reviewed draft PR #4779 at 038903ba7ea303c77e8014ea0bedf9b7b34b9f58 against the live unresolved threads.

The four previous Henry threads are still current and not outdated, so I did not duplicate them:

  1. model-controlled final_reply_target_id durable mutation needs an approval/ephemeral boundary
  2. local-dev exposes outbound delivery tools even when the provider registry is empty
  3. runtime-global Slack provider registration leaks targets into WebUI bundles without explicit Slack opt-in
  4. Slack mount construction ignores registry registration failure/replacement

This pass found one additional distinct validation issue inline. I also noted the registry/lifecycle/order concerns, but they are mostly consequences of the same runtime-global provider side effect already covered by the existing WebUI/Slack threads.

Verification:

  • PASS: cargo test -p ironclaw_reborn_composition local_dev_outbound_delivery_capabilities_use_late_registered_targets -- --nocapture
  • PASS: cargo test -p ironclaw_reborn_composition --features slack-v2-host-beta slack_host_beta_targets_wire_through_outbound_preferences_facade -- --nocapture

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review (multi-agent)

Intent: Expose outbound delivery targets to the Reborn model via local-dev capabilities, runtime registry, Slack registration, and WebUI preference reuse.

Stats: 5 findings (from 11 raw, 5 after filter/dedup) across 3 files. Reviewers run: security, bugs, performance, tests, conventions, local-patterns, maintainability, approach. Reviewers failed: none. Body-only: 0.

Bugs

  1. Medium List tool accepts malformed JSON as valid input (crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs:243-257, confidence 93) — anchor: crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs:243
    parse_optional_channel treats non-object inputs as if no channel was supplied, so malformed list-tool invocations can run despite the declared object schema.

Tests

  1. Medium Cover malformed outbound target IDs (crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs:260-272, confidence 83) — anchor: crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs:260
    The invalid-format branch of parse_target_id is not covered, so malformed model input can regress silently.
  2. Medium Test explicit-owner precedence in caller selection (crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs:191-213, confidence 79) — anchor: crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs:191
    The capability test covers the actor-only path but not the branch where explicit_owner_user_id must override the actor.

Approach

  1. Medium Runtime mutation API cuts across the WebUI bundle boundary (crates/ironclaw_reborn_composition/src/runtime.rs:1030-1054, confidence 80) — anchor: crates/ironclaw_reborn_composition/CLAUDE.md:3-6
    The post-build runtime registration hook makes RebornRuntime a mutable handoff point for a Slack/WebUI dependency that already has an explicit provider-list seam.

Local Patterns

  1. Low Registration warning omits the provider key (crates/ironclaw_reborn_composition/src/outbound_preferences.rs:121-124, confidence 74) — anchor: crates/ironclaw_reborn_composition/src/outbound_preferences.rs:121-124
    The poisoned-lock warning does not include the provider key, which weakens diagnostics if registration fails.

Comment thread crates/ironclaw_reborn_composition/src/runtime.rs Outdated
Comment thread crates/ironclaw_reborn_composition/src/outbound_preferences.rs Outdated
@henrypark133
henrypark133 force-pushed the codex/slack-extension-manifest branch from 62101f4 to e34744c Compare June 12, 2026 21:23
henrypark133 and others added 11 commits June 14, 2026 10:19
- Reserve "slack" as host-bundled extension id (prevent filesystem shadowing)
- Add builtin_first_party_trust_policy regression test for Slack admin entry
- Activate manifest-backed channel packages from WebUI (suppress only wasm_channel)
- Preserve legacy Slack connect controls for pre-install deployments
- Project only ProductSurfaceKind::ExternalChannel to channel kind
- Parse each manifest once via ExtensionManifestRecord
- Rename product_adapter.host_beta section to stable product_adapter.inbound
- Add caller-level tests: list_extension_registry, extension_info, ChannelsTab render

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rface_kinds

#4778 lands the ProductAdapter surface projection, so the lifecycle summary
now carries surface_kinds. Flip CHANNEL_CLASSIFICATION_AVAILABLE to true and
make extension_is_channel_surface a real predicate (ExternalChannel), so
connected channel names render in the model runtime context instead of unknown.

Convert the two stubbed tests to positive cases: empty list -> Known([]),
mixed list -> only active channel-surface extensions reported.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ranch

Remove the dead 'if !CHANNEL_CLASSIFICATION_AVAILABLE' arm inside the
Some(Ok(response)) match: lifecycle_fut only issues the ExtensionList call
when classification is enabled, so a present response always means it is on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…nnel refactor

- assets.rs smoke test: assert showLegacySlackConnectActions (the refactor's
  built-in Slack status path) instead of the removed slackBuiltinStatus helper
- add extensions.kind.channel to all 10 non-en locales (en gained the key with
  the new channel surface kind; locale-parity test requires all locales match)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…try heading

Review findings (PR #4778 review 4493764666):
- F2: add extension-card.test.mjs proving kind=channel/wasm_channel surface
  Setup (setup_required/failed) and Reconfigure (active/ready) overflow actions
  on the real component; channels-tab.test stubbed ExtensionCard so this was
  uncovered.
- F4: render the 'Available channels' registry heading via t(channels.availableChannels)
  instead of a hardcoded literal; add the key to all 11 locales. Update the
  channels-tab test to locate the registry section by the RegistryCard component
  (heading is now an interpolated value, not a template literal); drop the now
  unused renderedValueAfter helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…nt surface_kinds cache

Review findings (PR #4778 review 4493764666):
- F5: remove CHANNEL_CLASSIFICATION_AVAILABLE (permanently true after the stub
  flip) and run the lifecycle ExtensionList fetch unconditionally when a
  lifecycle facade is wired.
- F3: document AvailableExtensionPackage.surface_kinds as an intentional
  single-parse cache (re-deriving in summary() would re-run the manifest
  projection, undoing the parse-once optimization).

F1 (per-turn ExtensionList cost) accepted as-is: the fetch is spawned off the
critical path under a 500 ms budget with abort-on-drop; caching deferred to a
follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Review finding (PR #4778 review 4494623523, finding 1): primaryExtensionAction
only suppressed the primary Activate button for legacy wasm_channel, so a
manifest-backed kind=channel Slack card fell through to 'activate' in
pairing_required/pairing states where the dedicated pairing section already
owns the flow. Suppress the primary action for channel-surface kinds in those
states (via isChannelExtensionKind); installed channels still return 'activate'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…kind projection

Review findings (PR #4778 review 06:28):
- Add missing channels.slack i18n key to all 11 locales (legacy Slack row
  rendered the raw key because the i18n helper returns the key on miss; the
  || "Slack" fallback never fired).
- Add filesystem-path test: a /system manifest with
  product_adapter.inbound.surface_kind = external_channel projects to
  ExternalChannel surface (previously only the bundled catalog path was covered).
- Add extension_kind regression test: non-channel summaries keep their runtime
  wire kind (wasm_tool, mcp_server) while channel surfaces map to "channel".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… wasm_channel registry

Review findings (PR #4778 review 07:01):
- Gate the Slack first-party catalog entry, its only-Slack symbols (slack_package,
  slack_assets, SLACK_MANIFEST, slack_manifest_digest), the factory trust-policy
  Slack AdminEntry, and Slack-asserting tests behind slack-v2-host-beta. Without
  the feature the Slack route/runtime/WebUI mounts don't exist, so the catalog
  must not advertise an unrunnable Slack extension. Clean clippy + tests in both
  feature-on and feature-off configs.
- Add useExtensions hook test proving an uninstalled kind=wasm_channel registry
  entry lands in channelRegistry, not toolRegistry (isChannelExtensionKind covers
  both channel and wasm_channel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@serrrfirat
serrrfirat marked this pull request as ready for review June 15, 2026 08:06
@coderabbitai

coderabbitai Bot commented Jun 15, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Threads optional EffectiveRuntimePolicy through local-dev composition structs. Adds local_dev_effects_require_approval and a TOML policy grant. Implements two synthetic local-dev capabilities (builtin.outbound_delivery_targets_list, builtin.outbound_delivery_target_set) with approval-resume orchestration. Conditionally appends a Slack manifest AdminEntry to the first-party trust policy under slack-v2-host-beta. Updates the trigger delivery contract.

Changes

Outbound Delivery Capability and Runtime Policy Wiring

Layer / File(s) Summary
Runtime policy threading and Slack trust entry
src/factory.rs
EffectiveRuntimePolicy added to RebornLocalRuntimeServices and RebornLocalDevStoreGraphInput, threaded through both store-graph builders. builtin_first_party_trust_policy refactored to mutable entries vector; under slack-v2-host-beta, appends Slack manifest AdminEntry at /system/extensions/slack/manifest.toml with digest from slack_manifest_digest(). Test verifies first-party trust only on exact path+digest match.
Approval policy helper and capability policy grant
src/local_dev_authorization.rs, src/local_dev_capability_policy.toml
New local_dev_effects_require_approval helper resolves (ApprovalPolicy, RuntimeProfile) from optional runtime policy, builds RuntimeProfileApprovalGatePolicy, and returns .effects_require_approval(...). local_dev_authorizer refactored to share local_dev_approval_policy. TOML grant added for builtin.outbound_delivery_target_set with dispatch_capability+external_write under ambient mounts.
Outbound delivery wiring through local-dev stack
src/runtime.rs, src/runtime/local_dev.rs, src/runtime/local_dev/refreshing_capability_port.rs, src/runtime/local_dev/synthetic_capability.rs, src/slack_host_beta.rs
capability_wiring gains optional OutboundPreferencesProductFacade param; clones approval/lease stores, computes approval-required flag, stores all in LocalDevLoopCapabilityPortFactory. New fields propagate through RefreshingLocalDevCapabilityPortConfig. invoke_capability gains resume-mode exclusivity guard and uses approval_resume.input directly when present. runtime.rs call site updated. Slack host-beta hoists provider into local before struct construction.
Outbound delivery list/set capability handlers
src/runtime/local_dev/outbound_delivery.rs
New 704-line module implements builtin.outbound_delivery_targets_list (optional channel filter, caller derivation, facade call) and builtin.outbound_delivery_target_set (fingerprint-based approval-request creation on first call; fingerprint+correlation+lease claim on approval resume; facade set_outbound_preferences call). Includes strict JSON validators, scope derivation, resume token helpers, and error mappers.
Integration tests, harness updates, and delivery contract
src/runtime/local_dev/tests.rs, src/runtime/local_dev/shell_tests.rs, docs/reborn/contracts/triggers.md
Adds StaticOutboundDeliveryTargetProvider test double; full approval-gate flow test asserting owner-preference persistence, actor non-persistence, lease consumption; LocalYolo bypass; hidden-when-no-facade assertions; harness updates setting outbound_preferences_facade: None. Trigger delivery contract updated to route target selection through outbound delivery track.

Sequence Diagram(s)

sequenceDiagram
  rect rgba(100, 149, 237, 0.5)
    Note over Agent, OutboundPreferencesProductFacade: builtin.outbound_delivery_target_set (approval-gated)
  end
  participant Agent
  participant set_handler as outbound_delivery.set
  participant ApprovalRequestStore
  participant CapabilityLeaseStore
  participant OutboundPreferencesProductFacade

  Agent->>set_handler: invoke(target_id)
  set_handler->>set_handler: compute fingerprint(scope/capability/input)
  set_handler->>ApprovalRequestStore: save pending dispatch approval + resume token
  set_handler-->>Agent: ApprovalRequired(resume_token)

  Agent->>set_handler: invoke(approval_resume=token)
  set_handler->>set_handler: verify replay input equality + recompute fingerprint
  set_handler->>ApprovalRequestStore: load + validate status/action/correlation
  set_handler->>CapabilityLeaseStore: find active lease for requester+fingerprint
  set_handler->>CapabilityLeaseStore: claim lease
  set_handler->>OutboundPreferencesProductFacade: set_outbound_preferences(final_reply_target_id)
  set_handler->>CapabilityLeaseStore: consume claimed lease
  set_handler-->>Agent: Completed
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related issues

Possibly related PRs

  • nearai/ironclaw#4778: Directly overlaps — both PRs modify builtin_first_party_trust_policy() in factory.rs and add slack-v2-host-beta-gated trust evaluation tests for Slack manifest identity.
  • nearai/ironclaw#4769: Slack channel-delivery E2E tests in that PR exercise the final_reply_target_id outbound preference that this PR's builtin.outbound_delivery_target_set capability writes.

Poem

🦀 Two builtins rise from the local-dev floor,
targets_list to scout, target_set to implore.
A fingerprint guards the approval-gated gate,
Resume with a token — the lease seals your fate.
Slack earns first-party trust by path and by hash,
Runtime policy threaded through factory in a flash. ⚙️

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 inconclusive)

Check name Status Explanation Resolution
Title check ❓ Inconclusive Title is vague and doesn't follow Conventional Commits style; lacks type/scope prefix and details what's exposed rather than what changed. Use Conventional Commits format: feat(local-dev): expose outbound delivery targets or similar, clarifying scope and change type.
✅ Passed checks (3 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description check ✅ Passed PR description comprehensively covers summary, change type, validation, security impact, trust-boundary checklist, blast radius, and stack notes.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@serrrfirat
serrrfirat force-pushed the codex/channel-manifest-surfaces branch 2 times, most recently from c334a1f to 6b4efab Compare June 15, 2026 11:20
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

CodeRabbit follow-up on the current head (6b4efabe8):

  • Addressed the valid duplication nit: caller_for_run and resource_scope_for_run now share effective_user_id(...).
  • Verified the registry/provider-registration comments are stale on this branch: MutableOutboundDeliveryTargetRegistry, register_provider(...), and the post-build RebornRuntime provider-registration hook no longer exist. Slack providers are passed as composition-time inputs into build_webui_services_with_connectable_channels(...).
  • Verified the product-workflow list-path coverage already exists via list_projects_external_channel_surface_kind_through_extension_info; reran it locally.
  • Did not remove builtin__outbound_delivery_target_set: this PR intentionally exposes the setter through the model path with an approval gate for default local-dev, and local-dev-yolo bypasses that gate by runtime policy.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_reborn_composition/src/local_dev_capability_policy.toml`:
- Around line 224-228: The grant block authorizing the
builtin.outbound_delivery_target_set capability with external_write and ambient
mounts is stale and should be removed entirely. Since the durable setter was
removed and delivery-target selection is now product-owned via WebUI, this grant
reopens an unauthorized LLM-controlled write path and violates fail-closed
security principles. Delete the entire grants block that contains capability =
"builtin.outbound_delivery_target_set" with its associated effects and mounts
entries.

In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs`:
- Around line 186-200: The approval lease is being consumed after the durable
outbound preference is set via the facade set_outbound_preferences method,
creating a state consistency issue where the durable preference persists even if
lease consumption fails. Per the coding guidelines to fail closed for approvals,
move the capability_leases.consume operation for the approved_lease BEFORE the
self.facade.set_outbound_preferences call, so that approval validation and lease
consumption happen atomically before committing any durable state changes. This
ensures that if lease consumption fails, the durable preference write never
occurs and state remains consistent.
- Around line 116-121: In
crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs at
lines 116-121 (anchor), 202-207 (sibling), 543-563 (sibling), and 590-599
(sibling), replace all `.map_err(|_| ...)` patterns that discard error
information with patterns that bind the error and preserve it. For each site,
change the closure from ignoring the error (|_|) to capturing it (|e|), then
include the error details in the AgentLoopHostError construction by using
something like error.to_string() or similar to carry the original cause
information into the error message, so that debugging information is not lost
when these serialization and parsing failures occur.

In `@crates/ironclaw_reborn_composition/src/runtime/local_dev/tests.rs`:
- Around line 177-184: The StaticOutboundDeliveryTargetProvider implementation
currently ignores the _caller parameter in the list_outbound_delivery_targets
method, which means the test cannot verify that the handler is using the correct
caller scope. Modify the StaticOutboundDeliveryTargetProvider to record or
validate the caller parameter instead of discarding it. Store the expected
caller (the owner created in the test) in the provider, and only return the
outbound delivery target entry when the caller matches the expected owner. Then
add an assertion in the test to verify that the discovery operation used the
correct caller scope by checking that targets were retrieved only when queried
by the authorized owner.
- Around line 1641-1648: The test currently only validates that certain
capabilities are absent from the visible_capabilities surface through
descriptor_ids checks, but it should also verify that the corresponding provider
tools are absent from the tool_definitions() to ensure proper fail-closed
security behavior. Add assertions after the existing descriptor_ids checks that
verify builtin__outbound_delivery_targets_list and
builtin__outbound_delivery_target_set tool names are not present in the
surface's tool definitions, following the same pattern as the existing
capability ID assertions to confirm the missing-facade path completely hides
these provider tools from model exposure.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7fb62aac-cb98-4542-bc36-0f5acc9cb559

📥 Commits

Reviewing files that changed from the base of the PR and between 6525f1f and c334a1f.

📒 Files selected for processing (12)
  • crates/ironclaw_reborn_composition/src/factory.rs
  • crates/ironclaw_reborn_composition/src/local_dev_authorization.rs
  • crates/ironclaw_reborn_composition/src/local_dev_capability_policy.toml
  • crates/ironclaw_reborn_composition/src/runtime.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/refreshing_capability_port.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/shell_tests.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/synthetic_capability.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/tests.rs
  • crates/ironclaw_reborn_composition/src/slack_host_beta.rs
  • docs/reborn/contracts/triggers.md

Comment thread crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs Outdated
Comment thread crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs Outdated
Comment thread crates/ironclaw_reborn_composition/src/runtime/local_dev/tests.rs
Comment thread crates/ironclaw_reborn_composition/src/runtime/local_dev/tests.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_reborn_composition/src/factory.rs (1)

4901-4964: 🧹 Nitpick | 🔵 Trivial | ⚡ Quick win

Add the missing-digest fail-closed case.

This test covers wrong digest/path, but not a matching Slack path with digest: None. For a digest-bound first-party entry, that identity must stay sandbox.

Proposed test addition
         assert_eq!(
             wrong_digest.provenance,
             ironclaw_trust::TrustProvenance::Default
         );
+
+        let missing_digest = ironclaw_trust::TrustPolicy::evaluate(
+            &policy,
+            &ironclaw_trust::TrustPolicyInput {
+                identity: slack_identity("/system/extensions/slack/manifest.toml", None),
+                requested_trust: ironclaw_host_api::RequestedTrustClass::FirstPartyRequested,
+                requested_authority: Default::default(),
+            },
+        )
+        .expect("missing digest slack identity should evaluate");
+
+        assert_eq!(missing_digest.effective_trust.class(), TrustClass::Sandbox);
+        assert_eq!(
+            missing_digest.provenance,
+            ironclaw_trust::TrustProvenance::Default
+        );
 
         let wrong_path = ironclaw_trust::TrustPolicy::evaluate(
             &policy,

As per coding guidelines, “Fail closed for auth, approvals, trust, filesystem containment, network policy, secret leases, runtime selection, and adapter identity.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_composition/src/factory.rs` around lines 4901 - 4964,
The test builtin_first_party_trust_policy_includes_slack_local_manifest_entry()
is missing a fail-closed test case for when the Slack manifest path is correct
but the digest is None (missing). Add a fourth test case following the same
pattern as wrong_digest and wrong_path that evaluates the trust policy with the
correct path "/system/extensions/slack/manifest.toml" but with None as the
digest parameter, then assert that it evaluates to TrustClass::Sandbox with
TrustProvenance::Default, ensuring the policy fails closed when digest is
missing.

Source: Coding guidelines

♻️ Duplicate comments (2)
crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs (2)

186-200: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Consume the approval lease before the durable preference write.

set_outbound_preferences persists final_reply_target_id, then consume() can fail and return an error after product state changed. Move lease consumption before Line 186, or add rollback. As per coding guidelines, “Fail closed for auth, approvals, trust, filesystem containment, network policy, secret leases, runtime selection, and adapter identity.”

Fail closed before committing preference
         let target_summary = target_id.as_str().to_string();
         let caller = caller_for_run(&invocation, &self.fallback_user_id);
+        if let Some(approved_lease) = approved_lease {
+            self.capability_leases
+                .consume(&approved_lease.scope, approved_lease.lease_id)
+                .await
+                .map_err(|error| approval_lease_error("consume_approval_lease", error))?;
+        }
         let response = self
             .facade
             .set_outbound_preferences(
                 caller,
@@
             .await
             .map_err(|error| outbound_delivery_host_error("set_target", error))?;
-        if let Some(approved_lease) = approved_lease {
-            self.capability_leases
-                .consume(&approved_lease.scope, approved_lease.lease_id)
-                .await
-                .map_err(|error| approval_lease_error("consume_approval_lease", error))?;
-        }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs`
around lines 186 - 200, The approval lease consumption must happen before the
durable state write to satisfy fail-closed security principles. Move the lease
consumption block (the if let Some(approved_lease) conditional that calls
self.capability_leases.consume()) to execute before the set_outbound_preferences
call, so that if lease consumption fails, no durable preference state has been
written. This ensures the system fails closed for approval verification before
committing any preference changes.

Source: Coding guidelines


116-121: 🛠️ Refactor suggestion | 🟠 Major | ⚡ Quick win

Preserve the mapped error cause.

These .map_err(|_| ...) sites still discard the source failure, which violates Fail loud diagnostics for boundary/runtime errors. Bind the error and carry reason: error.to_string() or equivalent typed context into AgentLoopHostError. As per coding guidelines, “Do not use .map_err(|_| OtherError) — a closure that ignores its error binding and substitutes a generic error drops the underlying cause.”

#!/bin/bash
# Verify remaining source-dropping map_err closures in changed Rust files.
rg -n -C2 --type=rust '\.map_err\(\|_\|' crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs

Also applies to: 202-207, 535-540, 549-554, 562-567, 584-589

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs`
around lines 116 - 121, The `.map_err(|_| ...)` closure discards the original
error source, violating error diagnostics requirements. Fix this by binding the
error parameter instead of ignoring it with `|_|`, then pass the error details
to the AgentLoopHostError constructor via a reason field using
`error.to_string()` or equivalent context. This issue appears at multiple sites
throughout the file: the serialization failure at lines 116-121, and also at
lines 202-207, 535-540, 549-554, 562-567, and 584-589. Update all of these
`.map_err` closures to capture and preserve the actual error cause rather than
substituting a generic message that hides the underlying failure.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/ironclaw_reborn_composition/src/factory.rs`:
- Around line 4901-4964: The test
builtin_first_party_trust_policy_includes_slack_local_manifest_entry() is
missing a fail-closed test case for when the Slack manifest path is correct but
the digest is None (missing). Add a fourth test case following the same pattern
as wrong_digest and wrong_path that evaluates the trust policy with the correct
path "/system/extensions/slack/manifest.toml" but with None as the digest
parameter, then assert that it evaluates to TrustClass::Sandbox with
TrustProvenance::Default, ensuring the policy fails closed when digest is
missing.

---

Duplicate comments:
In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs`:
- Around line 186-200: The approval lease consumption must happen before the
durable state write to satisfy fail-closed security principles. Move the lease
consumption block (the if let Some(approved_lease) conditional that calls
self.capability_leases.consume()) to execute before the set_outbound_preferences
call, so that if lease consumption fails, no durable preference state has been
written. This ensures the system fails closed for approval verification before
committing any preference changes.
- Around line 116-121: The `.map_err(|_| ...)` closure discards the original
error source, violating error diagnostics requirements. Fix this by binding the
error parameter instead of ignoring it with `|_|`, then pass the error details
to the AgentLoopHostError constructor via a reason field using
`error.to_string()` or equivalent context. This issue appears at multiple sites
throughout the file: the serialization failure at lines 116-121, and also at
lines 202-207, 535-540, 549-554, 562-567, and 584-589. Update all of these
`.map_err` closures to capture and preserve the actual error cause rather than
substituting a generic message that hides the underlying failure.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 583db78a-c36e-44a4-9d98-79e00e34a6b1

📥 Commits

Reviewing files that changed from the base of the PR and between c334a1f and 6b4efab.

📒 Files selected for processing (12)
  • crates/ironclaw_reborn_composition/src/factory.rs
  • crates/ironclaw_reborn_composition/src/local_dev_authorization.rs
  • crates/ironclaw_reborn_composition/src/local_dev_capability_policy.toml
  • crates/ironclaw_reborn_composition/src/runtime.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/refreshing_capability_port.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/shell_tests.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/synthetic_capability.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/tests.rs
  • crates/ironclaw_reborn_composition/src/slack_host_beta.rs
  • docs/reborn/contracts/triggers.md

@serrrfirat
serrrfirat force-pushed the codex/channel-manifest-surfaces branch from 6b4efab to 7b2f622 Compare June 15, 2026 11:42

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs`:
- Around line 333-349: The authorization check in the outbound_delivery.rs file
at lines 333–349 uses CapabilityLeaseStore::leases_for_scope() which silently
converts store errors to empty results, masking transient failures. To fix the
"fail loud" violation, you must either: (1) refactor the CapabilityLeaseStore
trait to return Result<Vec<CapabilityLease>, Error> instead of
Vec<CapabilityLease>, then propagate the error from leases_for_scope() with the
? operator in the approval lease lookup; or (2) keep the trait signature but add
explicit error logging in the trait implementation in
crates/ironclaw_authorization/src/lib.rs and modify the caller to return a
distinct InternalError or Unavailable response (not Unauthorized) when
leases_for_scope() indicates a store failure. Either approach ensures store
failures are surfaced loudly instead of being masked as missing leases.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 538a61ca-2e84-4c9f-9723-d113d37df7fe

📥 Commits

Reviewing files that changed from the base of the PR and between 6b4efab and 7b2f622.

📒 Files selected for processing (12)
  • crates/ironclaw_reborn_composition/src/factory.rs
  • crates/ironclaw_reborn_composition/src/local_dev_authorization.rs
  • crates/ironclaw_reborn_composition/src/local_dev_capability_policy.toml
  • crates/ironclaw_reborn_composition/src/runtime.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/refreshing_capability_port.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/shell_tests.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/synthetic_capability.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/tests.rs
  • crates/ironclaw_reborn_composition/src/slack_host_beta.rs
  • docs/reborn/contracts/triggers.md

Comment on lines +333 to +349
let lease = self
.capability_leases
.leases_for_scope(&scope)
.await
.into_iter()
.find(|lease| {
lease.status == CapabilityLeaseStatus::Active
&& lease.grant.capability == capability_id
&& lease.grant.grantee == approval_record.request.requested_by
&& lease.invocation_fingerprint.as_ref() == Some(&fingerprint)
})
.ok_or_else(|| {
AgentLoopHostError::new(
AgentLoopHostErrorKind::Unauthorized,
"outbound delivery target approval lease is unavailable",
)
})?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Check CapabilityLeaseStore trait signature for leases_for_scope
ast-grep --pattern $'trait CapabilityLeaseStore {
  $$$
  async fn leases_for_scope($$$) -> $RET;
  $$$
}'
# Fallback: grep for return type
rg -n 'fn leases_for_scope' --type rust -A 3

Repository: nearai/ironclaw

Length of output: 8028


Authorization gate silently masks store failures — fail-loud violation.

CapabilityLeaseStore::leases_for_scope() returns Vec<CapabilityLease> (not Result). The trait implementation in crates/ironclaw_authorization/src/lib.rs:931–934 explicitly catches store errors and returns Vec::new() instead of propagating them. This means a transient DB failure silently converts to "no leases found," causing the code at lines 333–349 to return Unauthorized instead of surfacing the actual Unavailable/InternalError. Violates "Fail loud" invariant.

The approval gate should either:

  1. Change the trait to return Result<Vec<…>, …> and propagate errors with ?, or
  2. Log and return a distinct error for store unavailability (not "lease is unavailable").
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/outbound_delivery.rs`
around lines 333 - 349, The authorization check in the outbound_delivery.rs file
at lines 333–349 uses CapabilityLeaseStore::leases_for_scope() which silently
converts store errors to empty results, masking transient failures. To fix the
"fail loud" violation, you must either: (1) refactor the CapabilityLeaseStore
trait to return Result<Vec<CapabilityLease>, Error> instead of
Vec<CapabilityLease>, then propagate the error from leases_for_scope() with the
? operator in the approval lease lookup; or (2) keep the trait signature but add
explicit error logging in the trait implementation in
crates/ironclaw_authorization/src/lib.rs and modify the caller to return a
distinct InternalError or Unavailable response (not Unauthorized) when
leases_for_scope() indicates a store failure. Either approach ensures store
failures are surfaced loudly instead of being masked as missing leases.

@serrrfirat
serrrfirat merged commit a9e94b6 into main Jun 15, 2026
69 checks passed
@serrrfirat
serrrfirat deleted the codex/channel-manifest-surfaces branch June 15, 2026 13:39
serrrfirat added a commit that referenced this pull request Jun 15, 2026
…merge; lock retry_run

- remove stray thread_scope arg from the local-yolo-outbound test call site so it
  matches the merged 9-param capability_wiring signature (my thread_scope removal +
  main's outbound_preferences_facade + trajectory_observer params)
- serialize retry_run with lock_thread_operation before retry_turn, matching
  submit_turn/delete_thread, to close the retry-vs-delete race (CodeRabbit)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
elliotBraem added a commit to NEARBuilders/ironclaw that referenced this pull request Jun 17, 2026
* fix(runtime-context): key comms preferences by run owner + cover JoinError gaps (#4895)

Addresses post-merge review findings on #4836.

- Bug (Medium): the communication-context provider keyed outbound
  delivery preferences by the *actor* instead of the run *owner*. Product
  inbound and trusted-trigger runs can carry an explicit thread owner
  (subject/creator) distinct from the actor; the stored preference belongs
  to the owner. Resolve the caller's user_id via
  `scope.explicit_owner_user_id()` with actor fallback — matching
  `TurnScope::to_resource_scope` — so shared/channel inbound and trigger
  runs render the owner's delivery target, not the actor's. Adds two
  regression tests asserting the lookup is keyed by owner vs actor through
  a caller-capturing facade.

- Tests (Medium): cover `CommunicationContextFetch::resolve`'s JoinError
  branches that were previously unexercised — actorless failure degrades
  to `None`, actor-present failure degrades to `Some(Unknown)`.

- Docs (Low): the product-context-factory plan still specified the
  rejected 256-byte `RunOriginAdapter` bound; update to the as-built
  512-byte cap (mirroring `AdapterKind`) so follow-up work does not
  reintroduce the narrowing.

Not addressed (verified out of scope): the `model_safe_label` "injection"
finding is a false positive — label sources are admin/system-set and the
sanitizer strips structural characters (existing hostile-input tests pass);
the shared-route surface assertion finding is already covered by
`shared_user_message_records_channel_surface_type`; the string-based
trigger-helper finding is owned by issue #4851's plan; the
RunOriginAdapter byte-mirror is already mitigated by a named const + docs
+ tests.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: surface missing-credential auth gate before the approval gate (#4840)

* fix(host_runtime): surface AuthRequired before approval gate on missing credentials (Fix B)

Extracts capability_credential_requirements() as the single source of truth
for credential requirements derived from the capability manifest descriptor.
Both the new credential pre-flight check and the existing dispatch-time
obligation check call this function — no second computation added.

Adds credential_preflight_check() on DefaultHostRuntime and calls it in
invoke_capability() and spawn_capability() BEFORE apply_persistent_approval_policy(),
so AuthRequired is returned without persisting an approval request when a
required credential is absent.

Wires the optional secret store from HostRuntimeServices.build_host_runtime()
into DefaultHostRuntime via with_credential_preflight_store(). The pre-flight
is skipped in minimal/test graphs that don't supply a secret store; the
dispatch-time obligation check remains the enforcement backstop regardless.

Tests (host_runtime_services_contract):
- invoke_capability_missing_credential_returns_auth_before_approval
- invoke_capability_present_credential_proceeds_to_approval
- invoke_capability_no_credential_requirement_proceeds_normally
- credential_requirements_preflight_and_dispatch_agree_on_same_handles

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(host_runtime): correct capability_credential_requirements docstring and strengthen test

The docstring claimed "both the pre-flight check and the dispatch-time
obligation check call this function" — that was false. The obligation
handler in BuiltinObligationHandler derives required handles by iterating
descriptor.runtime_credentials directly; it does not call this function.
Correct the docstring to accurately describe the agreement at source-data
level (both iterate the same required-true entries) and explain why gate-ID
divergence between pre-flight and backstop is moot in practice (pre-flight
fires first when a secret store is wired).

Rename `credential_requirements_preflight_and_dispatch_agree_on_same_handles`
to `credential_requirements_extraction_matches_descriptor_required_credentials`
with a scope note clarifying what the test actually verifies (canonical fn
vs descriptor, not vs obligation handler), add an assertion that
credential_requirements is empty for secret_handle source type, and reference
the caller-level test that covers the gate-ordering guarantee.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(host-runtime): skip credential pre-flight on store errors per review

On a transient SecretStore Err, credential_preflight_check previously
returned AuthRequired (treating backend failure as credential absent),
burning a user auth interaction. Now it returns None on Err so the
dispatch-time obligation check remains the sole enforcement backstop.

Also applied:
- FIX 2: add trust-class-agnostic comment at both pre-flight call sites
- FIX 3: rename field secret_store → credential_preflight_store to match
  the with_credential_preflight_store builder
- FIX 4: take registry.snapshot() once in invoke_capability and
  spawn_capability and pass it into credential_preflight_check, removing
  the redundant internal snapshot

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(host-runtime): add credential pre-flight edge-case tests per review

- FIX 5: change section separator to triple-dash style matching
  memory_prompt_context.rs
- FIX 6: four new tests:
  a. spawn_capability_missing_credential_returns_auth_before_approval —
     spawn path mirrors invoke_capability pre-flight behavior
  b. invoke_capability_no_credential_requirement_with_wired_store_proceeds_normally —
     wired store + zero required credentials hits is_empty() branch, not
     no-store early exit
  c. invoke_capability_secret_store_error_skips_preflight — erroring
     store stub confirms pre-flight skips on Err (FIX 1) and flow
     reaches the approval gate
  d. credential_requirements_extraction_returns_empty_for_all_optional_credentials —
     descriptor with only required=false credentials yields empty vecs

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(host-runtime): address #4840 review — product-auth preflight skip, scope validation, test wiring

- capability_credential_requirements no longer treats ProductAuthAccount injection
  slots as presence-checkable secrets, fixing a false-positive AuthRequired preflight
  for capabilities with already-connected product-auth accounts.
- Validate context/resource-scope consistency before the credential preflight queries
  the secret store, closing a forged-scope presence-probe window (invoke + spawn).
- Wire the preflight store in contract fixtures and seed credentials on the request's
  own ResourceScope; add product-auth and scope-validation regression tests.
- silent-ok annotation on the store-error skip; docs name both invoke and spawn;
  moved preflight tests to host_runtime_credential_preflight_contract.rs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(host-runtime): #4840 review round 2 — test fidelity + tighten public surface

- store-error regression now drives the manifest-backed credential backstop via
  ApprovalThenGrantAuthorizer (was masked by a test-authorizer-injected obligation).
- add spawn_capability present-credential happy-path test (proceeds to ApprovalRequired).
- make the credential-preflight-store setter test-only/crate-private; tests wire it
  through HostRuntimeServices::build_host_runtime.
- stop publishing capability_credential_requirements from the crate root; keep it
  crate-private with equivalent coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(host-runtime): exercise the real dispatch-time obligation backstop on store error (#4840)

The store-error regression now grants the required secret so dispatch authorization
passes and the resumed call reaches BuiltinObligationHandler::preflight_secret_injection,
where the AlwaysErrorSecretStore metadata() probe errors and the handler fails closed
(secret_obligation_failed). Previously the grant omitted the secret, so the block came
from grant-matching authorization — not the dispatch-time credential backstop the PR
contract relies on (obligations.rs preflight_secret_injection fails closed on store error).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(host-runtime): single owner for the secret-presence rule (#4840)

Extract obligations::secret_present as the one definition of 'is this required
secret present in scope', shared by the credential pre-flight (ordering) and the
dispatch-time obligation backstop (enforcement). Removes the duplicated metadata()
presence rule across production.rs and obligations.rs so the two paths cannot drift.
Each caller still owns its store-error policy (pre-flight fails open and skips; the
backstop fails closed). The happy-path double-read remains (documented) — addresses
the split-ownership half of the review; collapsing the read is a separate follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(host-runtime): make store-error backstop test airtight + cover required product-auth (#4840)

Addresses PR #4840 review round 3:
- The store-error regression now uses a metadata()-call-counting error store and
  asserts the resume drove at least one further probe after a counter reset — proving
  the resume reaches BuiltinObligationHandler::preflight_secret_injection (authorization
  passed) and fails closed there, not at a premature authorization denial. Both paths map
  to RuntimeFailureKind::Authorization, so the probe count is the distinguishing signal.
  Removes the now-unused AlwaysErrorSecretStore and the stale test doc.
- Corrects the comment wording: the secret-injection obligation comes from grant
  evaluation against the manifest, not from ApprovalThenGrantAuthorizer.
- Adds credential_requirements_extraction_excludes_required_product_auth_account unit
  test: a required product_auth_account credential is excluded from required_secrets
  (no false-positive pre-flight AuthRequired) but still surfaces in credential_requirements.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(host-runtime): preserve secret-store failure cause when failing closed (#4840)

The dispatch-time obligation backstop dropped SecretStoreError with map_err(|_|),
collapsing outages into an opaque secret-obligation failure with no server-side
trail. Bind the error and log it at debug! (SecretStoreError Display carries no raw
secret material) before returning the sanitized secret_obligation_failed(); the
caller still receives the opaque error. Mirror the same cause-logging in the
fail-open pre-flight skip arm.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* fix(agent-loop): surface resume-origin capability failures instead of dying as scope_mismatch (#4899)

A 2nd+ Slack approval (or auth) resume of a capability could terminally
fail with scope_mismatch ("capability input ref is not scoped to this
loop run"). On a resume dispatch returning a transient Backend error,
handle_capability_error cleared the pending resume slot, then the
RecoveryOutcome::Retry path re-dispatched via
capability_invocation_from_candidate(call, None) — dropping the resume
context. The non-resume path resolved the original-run input_ref against
the resuming run, failed ensure_ref_scoped_to_run, and killed the run as
HostUnavailable.

Intercept Retry for approval-resume AND auth-resume origin failures:
surface the real backend error to the model as a tool result and
continue the loop (so the user can re-approve / re-auth) instead of
re-dispatching. Kills the scope_mismatch (S1) and avoids
double-executing a side effect whose lease is one-shot after a Backend
error (S2).

Adds regression tests for both resume origins.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui): keep code block overflow local (#4791)

* fix(webui): keep code block overflow local

* fix(webui): keep code block typography readable after merge

* fix(auth-resume): preserve input replay across gate boundaries (#4910)

* fix auth resume input replay

* fix auth resume approval token carryover

* fix(reborn): normalize bare workspace tool paths (#4846)

* fix(reborn): normalize bare workspace tool paths

* Preserve scoped path URL validation for workspace aliases

* Normalize empty workspace alias segments

---------

Co-authored-by: Robert Yan <mstr.raphael@gmail.com>

* [codex] Represent Slack as a product-adapter extension (#4778)

* feat(reborn): declare Slack channel as extension manifest

* fix(reborn): address extension review feedback

* fix(reborn): address Slack product-adapter extension review feedback

- Reserve "slack" as host-bundled extension id (prevent filesystem shadowing)
- Add builtin_first_party_trust_policy regression test for Slack admin entry
- Activate manifest-backed channel packages from WebUI (suppress only wasm_channel)
- Preserve legacy Slack connect controls for pre-install deployments
- Project only ProductSurfaceKind::ExternalChannel to channel kind
- Parse each manifest once via ExtensionManifestRecord
- Rename product_adapter.host_beta section to stable product_adapter.inbound
- Add caller-level tests: list_extension_registry, extension_info, ChannelsTab render

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(runtime-context): enable connected-channel classification via surface_kinds

#4778 lands the ProductAdapter surface projection, so the lifecycle summary
now carries surface_kinds. Flip CHANNEL_CLASSIFICATION_AVAILABLE to true and
make extension_is_channel_surface a real predicate (ExternalChannel), so
connected channel names render in the model runtime context instead of unknown.

Convert the two stubbed tests to positive cases: empty list -> Known([]),
mixed list -> only active channel-surface extensions reported.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(runtime-context): cargo fmt + drop unreachable classification branch

Remove the dead 'if !CHANNEL_CLASSIFICATION_AVAILABLE' arm inside the
Some(Ok(response)) match: lifecycle_fut only issues the ExtensionList call
when classification is enabled, so a present response always means it is on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): repair Slack extension asset + locale checks after channel refactor

- assets.rs smoke test: assert showLegacySlackConnectActions (the refactor's
  built-in Slack status path) instead of the removed slackBuiltinStatus helper
- add extensions.kind.channel to all 10 non-en locales (en gained the key with
  the new channel surface kind; locale-parity test requires all locales match)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(webui-v2): cover ExtensionCard channel overflow + localize registry heading

Review findings (PR #4778 review 4493764666):
- F2: add extension-card.test.mjs proving kind=channel/wasm_channel surface
  Setup (setup_required/failed) and Reconfigure (active/ready) overflow actions
  on the real component; channels-tab.test stubbed ExtensionCard so this was
  uncovered.
- F4: render the 'Available channels' registry heading via t(channels.availableChannels)
  instead of a hardcoded literal; add the key to all 11 locales. Update the
  channels-tab test to locate the registry section by the RegistryCard component
  (heading is now an interpolated value, not a template literal); drop the now
  unused renderedValueAfter helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(runtime-context): drop permanent classification flag, document surface_kinds cache

Review findings (PR #4778 review 4493764666):
- F5: remove CHANNEL_CLASSIFICATION_AVAILABLE (permanently true after the stub
  flip) and run the lifecycle ExtensionList fetch unconditionally when a
  lifecycle facade is wired.
- F3: document AvailableExtensionPackage.surface_kinds as an intentional
  single-parse cache (re-deriving in summary() would re-run the manifest
  projection, undoing the parse-once optimization).

F1 (per-turn ExtensionList cost) accepted as-is: the fetch is spawned off the
critical path under a 500 ms budget with abort-on-drop; caching deferred to a
follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): suppress Activate for channel kinds during pairing

Review finding (PR #4778 review 4494623523, finding 1): primaryExtensionAction
only suppressed the primary Activate button for legacy wasm_channel, so a
manifest-backed kind=channel Slack card fell through to 'activate' in
pairing_required/pairing states where the dedicated pairing section already
owns the flow. Suppress the primary action for channel-surface kinds in those
states (via isChannelExtensionKind); installed channels still return 'activate'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): add channels.slack key + regression tests for surface-kind projection

Review findings (PR #4778 review 06:28):
- Add missing channels.slack i18n key to all 11 locales (legacy Slack row
  rendered the raw key because the i18n helper returns the key on miss; the
  || "Slack" fallback never fired).
- Add filesystem-path test: a /system manifest with
  product_adapter.inbound.surface_kind = external_channel projects to
  ExternalChannel surface (previously only the bundled catalog path was covered).
- Add extension_kind regression test: non-channel summaries keep their runtime
  wire kind (wasm_tool, mcp_server) while channel surfaces map to "channel".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): gate Slack catalog entry behind slack-v2-host-beta; test wasm_channel registry

Review findings (PR #4778 review 07:01):
- Gate the Slack first-party catalog entry, its only-Slack symbols (slack_package,
  slack_assets, SLACK_MANIFEST, slack_manifest_digest), the factory trust-policy
  Slack AdminEntry, and Slack-asserting tests behind slack-v2-host-beta. Without
  the feature the Slack route/runtime/WebUI mounts don't exist, so the catalog
  must not advertise an unrunnable Slack extension. Clean clippy + tests in both
  feature-on and feature-off configs.
- Add useExtensions hook test proving an uninstalled kind=wasm_channel registry
  entry lands in channelRegistry, not toolRegistry (isChannelExtensionKind covers
  both channel and wasm_channel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): consolidate Slack trust policy tests (#4778)

---------

Co-authored-by: Henry Park <henrypark133@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(agent-loop): gated final-answer nudge (reborn empty/canned turn endings) (#4837)

* feat(agent-loop): gated final-answer nudge to avoid empty/canned turn endings

When the reborn loop would otherwise end a turn with no real assistant
answer — an empty/trailed-off reply, the model-call budget exhausted, or
NoProgressDetected (which emits a canned "I stopped repeating the same
step" reply) — issue ONE extra tool-free model call asking the model to
synthesize a closing answer from the work it already did. This is the
reborn equivalent of the legacy loop's on_tool_intent_nudge /
force-text-recovery.

Gated by the (previously unimplemented) SteeringPolicy
`allow_driver_specific_nudges` flag, which defaults to false — so
production behavior is unchanged. Capped at one nudge per run
(`LoopExecutionState.final_answer_nudges_used`) so it can't issue
unbounded extra model calls. Wired into all three exit modes
(assistant_reply empty/trailed, budget IterationLimit, NoProgressDetected);
falls back to existing behavior when disabled, capped, or the model still
declines to answer.

Mechanism note: the tool-free call uses an EMPTY capability_view
(visible_capability_ids: []), not surface_version=None — the reborn model
gateway attaches tools from the capability port regardless of
surface_version, so only an empty view yields a true text-only request.

Evidence (PinchBench, Qwen3.5-122B, claude-haiku judge, controlled A/B on
the same merged tree): nudge OFF 0.682 vs nudge ON 0.768 (+0.086), closing
reborn to v2 parity (0.793). It specifically rescues tasks that do the
work but trip the no-progress detector (stock, events, spreadsheet,
polymarket), turning the canned give-up into a real answer.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* review: route nudge through admission + accounting, drop dead branch, add tests

Addresses review feedback on the final-answer nudge:

- Remove the empty/trailed-reply nudge branch in AssistantReplyStage. As the
  reviewer noted, DefaultReplyAdmissionStrategy rejects empty/artifact replies
  before they reach AssistantReplyStage, so that branch was dead for the default
  family. Empty replies are rejected → the loop continues → NoProgressDetected,
  which is where the nudge still fires. assistant_reply.rs is back to main.
- The nudged reply now goes through the SAME admission policy as a normal reply
  (ctx.planner.reply_admission().admit_reply) instead of a bare !is_empty()
  check — so blank text and provider-transcript artifacts can't be finalized.
- Preserve canonical assistant-reply token accounting: record the nudge turn's
  output tokens (provider usage, else the same estimate AssistantReplyStage
  uses) into recent_output_token_counts so the diminishing-returns window isn't
  fed stale data.
- Add caller-level tests with the gate enabled at the boundaries the nudge
  affects: no-progress (synthesizes via one tool-free model call), budget
  iteration-limit (completes instead of failing closed), gate-disabled (no model
  call, canned fallback), and the one-shot cap (no second call).

All 302 ironclaw_agent_loop lib tests pass.

Note on the remaining structural point (model call still issued from the exit
boundary via the host primitive rather than ModelStage): see PR discussion —
the exit-boundary stages don't own Prompt/Model, so a full "typed stage outcome"
move means restructuring the terminal-exit flow to re-enter the loop for one
tool-free turn. Happy to do that if preferred over this factoring.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* review: address CodeRabbit findings on final-answer nudge

- Move FINAL_ANSWER_NUDGE prompt out of Rust source into
  prompts/final_answer_nudge.md, loaded via include_str! (repo prompt-template
  invariant).
- Fix stale tool-free comment: clarify that the empty capability view on the
  model request (not the surface_version/capability_view None assignments) is
  what suppresses provider tools.
- Make MockHost::with_driver_nudges_enabled flip the steering flag in-place so
  it composes with other context-level builders regardless of order; drop the
  now-unused test_run_context_with_driver_nudges fixture.
- Add legacy-checkpoint decode regression test asserting a payload missing
  final_answer_nudges_used decodes to 0 (#[serde(default)] contract).
- Apply rustfmt to the new stage tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Pranav Raja <pranav.raja@near.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: serrrfirat <f@nuff.tech>

* Expose outbound delivery targets to Reborn model (#4779)

* feat(reborn): declare Slack channel as extension manifest

* fix(reborn): address extension review feedback

* fix(reborn): address Slack product-adapter extension review feedback

- Reserve "slack" as host-bundled extension id (prevent filesystem shadowing)
- Add builtin_first_party_trust_policy regression test for Slack admin entry
- Activate manifest-backed channel packages from WebUI (suppress only wasm_channel)
- Preserve legacy Slack connect controls for pre-install deployments
- Project only ProductSurfaceKind::ExternalChannel to channel kind
- Parse each manifest once via ExtensionManifestRecord
- Rename product_adapter.host_beta section to stable product_adapter.inbound
- Add caller-level tests: list_extension_registry, extension_info, ChannelsTab render

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(runtime-context): enable connected-channel classification via surface_kinds

#4778 lands the ProductAdapter surface projection, so the lifecycle summary
now carries surface_kinds. Flip CHANNEL_CLASSIFICATION_AVAILABLE to true and
make extension_is_channel_surface a real predicate (ExternalChannel), so
connected channel names render in the model runtime context instead of unknown.

Convert the two stubbed tests to positive cases: empty list -> Known([]),
mixed list -> only active channel-surface extensions reported.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(runtime-context): cargo fmt + drop unreachable classification branch

Remove the dead 'if !CHANNEL_CLASSIFICATION_AVAILABLE' arm inside the
Some(Ok(response)) match: lifecycle_fut only issues the ExtensionList call
when classification is enabled, so a present response always means it is on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): repair Slack extension asset + locale checks after channel refactor

- assets.rs smoke test: assert showLegacySlackConnectActions (the refactor's
  built-in Slack status path) instead of the removed slackBuiltinStatus helper
- add extensions.kind.channel to all 10 non-en locales (en gained the key with
  the new channel surface kind; locale-parity test requires all locales match)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(webui-v2): cover ExtensionCard channel overflow + localize registry heading

Review findings (PR #4778 review 4493764666):
- F2: add extension-card.test.mjs proving kind=channel/wasm_channel surface
  Setup (setup_required/failed) and Reconfigure (active/ready) overflow actions
  on the real component; channels-tab.test stubbed ExtensionCard so this was
  uncovered.
- F4: render the 'Available channels' registry heading via t(channels.availableChannels)
  instead of a hardcoded literal; add the key to all 11 locales. Update the
  channels-tab test to locate the registry section by the RegistryCard component
  (heading is now an interpolated value, not a template literal); drop the now
  unused renderedValueAfter helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(runtime-context): drop permanent classification flag, document surface_kinds cache

Review findings (PR #4778 review 4493764666):
- F5: remove CHANNEL_CLASSIFICATION_AVAILABLE (permanently true after the stub
  flip) and run the lifecycle ExtensionList fetch unconditionally when a
  lifecycle facade is wired.
- F3: document AvailableExtensionPackage.surface_kinds as an intentional
  single-parse cache (re-deriving in summary() would re-run the manifest
  projection, undoing the parse-once optimization).

F1 (per-turn ExtensionList cost) accepted as-is: the fetch is spawned off the
critical path under a 500 ms budget with abort-on-drop; caching deferred to a
follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): suppress Activate for channel kinds during pairing

Review finding (PR #4778 review 4494623523, finding 1): primaryExtensionAction
only suppressed the primary Activate button for legacy wasm_channel, so a
manifest-backed kind=channel Slack card fell through to 'activate' in
pairing_required/pairing states where the dedicated pairing section already
owns the flow. Suppress the primary action for channel-surface kinds in those
states (via isChannelExtensionKind); installed channels still return 'activate'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): add channels.slack key + regression tests for surface-kind projection

Review findings (PR #4778 review 06:28):
- Add missing channels.slack i18n key to all 11 locales (legacy Slack row
  rendered the raw key because the i18n helper returns the key on miss; the
  || "Slack" fallback never fired).
- Add filesystem-path test: a /system manifest with
  product_adapter.inbound.surface_kind = external_channel projects to
  ExternalChannel surface (previously only the bundled catalog path was covered).
- Add extension_kind regression test: non-channel summaries keep their runtime
  wire kind (wasm_tool, mcp_server) while channel surfaces map to "channel".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): gate Slack catalog entry behind slack-v2-host-beta; test wasm_channel registry

Review findings (PR #4778 review 07:01):
- Gate the Slack first-party catalog entry, its only-Slack symbols (slack_package,
  slack_assets, SLACK_MANIFEST, slack_manifest_digest), the factory trust-policy
  Slack AdminEntry, and Slack-asserting tests behind slack-v2-host-beta. Without
  the feature the Slack route/runtime/WebUI mounts don't exist, so the catalog
  must not advertise an unrunnable Slack extension. Clean clippy + tests in both
  feature-on and feature-off configs.
- Add useExtensions hook test proving an uninstalled kind=wasm_channel registry
  entry lands in channelRegistry, not toolRegistry (isChannelExtensionKind covers
  both channel and wasm_channel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): consolidate Slack trust policy tests (#4778)

* feat(reborn): expose outbound delivery targets to model

---------

Co-authored-by: Henry Park <henrypark133@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): unify extension registry flow (#4900)

* fix(reborn): unify extension registry flow

* fix(reborn): address extension registry review comments

* test(reborn): cover merged extension registry flow

* fix(reborn): clean google oauth configure copy

* fix(reborn): align notion oauth setup copy

* fix(reborn): clean token setup copy

* fix(reborn): prefer GitHub extension for repository data (#4894)

* fix(reborn): prefer github extension for repository data

* refactor(reborn): centralize github http routing hint

* fix(reborn): filter PRs from GitHub issue listings (#4888)

* fix(reborn): filter pull requests from github issues

* fix(reborn): preserve github issue pagination

* fix(reborn): use issue-only github search pagination

* fix(reborn): update github wasm trace fixture

* fix: auto-generate BETTER_AUTH_SECRET in bos-dev.sh and track styles.css in git

- bos-dev.sh now generates BETTER_AUTH_SECRET via openssl when .env is created or has an empty value
- Track ui/src/styles.css in git (add ! exception in .gitignore)
- Prevents auth plugin crash and UI build failure on fresh clones

* feat(reborn): polish the Automations panel UI (#4919)

* fix(automations-ui): readable summary cards and NEXT RUN value

Reflow the summary strip to at most three cards per row so the detail
text no longer wraps one word per line, and let StatCard accept a
valueClassName override so the NEXT RUN date renders at a smaller size
instead of truncating to "Jun…". Default StatCard sizing is unchanged.

* fix(automations-ui): surface delivery save errors and gate Slack hint

The delivery-defaults panel swallowed save/clear failures and showed no
feedback; it now renders an inline error from the mutation and flashes the
"Saved" confirmation on Clear as well as Save. The "reply approve <code> in
Slack" footnote is hidden unless an external Slack-style target exists.

* fix(automations-ui): label sub-hourly cron schedules

Minute- and hour-level cadences such as "* * * * *", "*/15 * * * *", and
"0 * * * *" rendered as "Custom schedule" because they have no single clock
time. They now read as "Every minute", "Every 15 minutes", and "Hourly at
:00".

* fix(automations-ui): space the run-row action button icons

The "Open run" and "Logs" buttons in the recent-runs list rendered the
icon flush against the label because the non-primary Button variants don't
add a gap between children. Add the same icon margin the rest of the app
uses for icon+label buttons.

* test(automations): lock the panel UI fixes into the served bundle

Add static-asset assertions driving the composed router so each Automations
panel UX fix — sub-hourly cron labels, summary card reflow + smaller NEXT RUN
value, run-row icon spacing, and delivery save-error/Slack-hint gating — is
guarded against a regression that drops it from the shipped SPA source.

* fix(i18n): add automations.delivery.saveFailed to every locale pack

The new key was added only to en.js, which breaks the i18n consistency test
that requires all locale packs to share the English key set. Add it to the
ten other packs (English placeholder, matching the existing untranslated
automations strings there).

* Explicit gate-open feedback for busy threads (no parking) (#4838)

* feat(product): explicit gate-open feedback for busy threads, no parking

Design decision (supersedes the closed defer-and-drain PR #4812): a message
arriving while another run holds the thread is recorded with the honest
terminal status RejectedBusy and the user gets an explicit notice — gate-aware
("an approval gate is open on this thread — resolve it before continuing,
then resend") when the blocking run is BlockedApproval/BlockedAuth, generic
otherwise. No background resubmission: the user is the retry actor.

- threads: MessageStatus::RejectedBusy + mark_message_rejected_busy (both
  backends); DeferredBusy kept as a legacy deserialization label, no longer
  written; RejectedBusy -> Submitted allowed so resends work
- product_workflow: ThreadBusy branches mark RejectedBusy; response variant
  renamed RejectedBusy with status-derived notice field
- webui: rejected_busy ack renders the notice as a system message in chat;
  wire-shape test asserts tag + notice
- slack: copy already honest from #4811; stale #4812 revisit comment removed
- e2e: runtime-level test proves ThreadBusy -> RejectedBusy with notice,
  NO resubmission on the blocking run's terminal event, and a fresh submit
  succeeds afterward

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): make RejectedBusy terminal across replay, ack, compaction, UI

Reviewers found RejectedBusy still inheriting auto-resubmit/defer semantics,
contradicting the no-parking contract. Fixes:

- inbound_turn: from_replay_parts returns a terminal AlreadyRejected handoff
  for RejectedBusy (re-rejects, never resubmits); to_ack now emits a settled
  ProductInboundAck::RejectedBusy so transport retries get Duplicate instead
  of resubmitting. Legacy DeferredBusy rows keep the resubmit path.
- reborn_services: replayed RejectedBusy returns RebornSubmitTurnResponse::
  RejectedBusy again (idempotent re-rejection) instead of building a fresh
  submission; status-to-notice mapping locked by BlockedApproval/BlockedAuth/
  generic tests.
- compaction: RejectedBusy (and frozen legacy DeferredBusy) map to
  SkipEphemeral(StableNonModelVisible), not DeferUntilStable, so busy threads
  can still compact instead of deferring forever.
- webui useChat: always clear processing on rejected_busy, mark the optimistic
  message failed, collision-free system-message id; added hook tests.
- tests: runtime no-resubmission assertion anchored on message identity;
  mark_message_rejected_busy negative coverage; webui handler test reuses
  StubServices via a queued response.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(threads): retire mark_message_deferred_busy writer

DeferredBusy is now a read-only legacy label — production writes RejectedBusy.
Remove the live writer from the SessionThreadService trait, both backends, the
Arc forwarder, and all test fakes. Legacy DeferredBusy read/replay coverage is
preserved via a doc-hidden inject_legacy_deferred_busy_for_test back-door on the
in-memory backend (never called from production). The DeferredBusy enum variant
and all read/replay/compaction handling are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): RejectedBusy follow-ups — Slack hint, honest replay, fail-loud, gating

- slack_delivery: SlackFinalReplyDeliveryObserver now recognizes
  ProductInboundAck::RejectedBusy { active_run_id: Some(_) } and posts the
  gate-aware busy hint (was DeferredBusy-only, so Slack rejections settled
  silently); None active_run_id posts nothing. Tests added.
- reborn_services: RejectedBusy response run metadata (active_run_id, status,
  event_cursor) is now Option — fresh ThreadBusy returns Some(real values),
  idempotent replay returns None instead of fabricating a fake Running run at
  cursor 0. Wire-shape + replay tests updated.
- inbound_turn: RejectedBusy replay fails loud on a malformed stored
  turn_run_id instead of silently dropping it.
- compaction: RejectedBusy + frozen legacy DeferredBusy map to
  SkipEphemeral(StableNonModelVisible), not DeferUntilStable, so busy threads
  compact instead of deferring forever. Integration tests added.
- threads: inject_legacy_deferred_busy_for_test gated behind a test-support
  cargo feature (absent from production builds); filesystem contract coverage
  for mark_message_rejected_busy (happy + invalid transitions).
- webui useChat: always clear processing on rejected_busy, mark the optimistic
  message failed, collision-free system-message id; hook tests added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test/docs(review): RejectedBusy coverage, durable timeline rejection, contract docs

- inbound_turn: regression test for fail-loud malformed RejectedBusy turn_run_id
- product_workflow_contract: RejectedBusy(None) settles + transport retry = Duplicate
- webui wire test: assert status + event_cursor present on fresh RejectedBusy path
- thread contracts (both backends): RejectedBusy -> Submitted resend transition
- webui useChat: persisted rejected_busy/deferred_busy rows now render failed on
  history reload with durable resend copy (was a normal-looking sent message);
  history-messages tests added
- slack_delivery: shared busy-hint path renamed deferred_busy_* -> busy_hint_*
  (fns, call sites, logs, docs); DeferredBusy kept only in the legacy arm
- runtime.rs: inline arch justification above the large RejectedBusy e2e test
- docs/reborn/contracts: product-adapters.md + conversation-binding.md document
  RejectedBusy as a durable terminal outcome; DeferredBusy marked legacy

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): openai-compat build break, Slack duplicate hint, conversations doc

- openai-compat-beta: ProductInboundAck::RejectedBusy added to the four
  non-exhaustive match sites (ack_helpers, chat_workflow, responses_workflow x2),
  mapped to the same retryable 429 as DeferredBusy — fixes E0004 that broke any
  build enabling openai-compat-beta. RejectedBusy->429 tests added.
- slack_delivery: busy-hint run-id extraction now unwraps
  Duplicate { prior } recursively, so a transport retry arriving as
  Duplicate { prior: RejectedBusy { Some(run) } } still posts the busy-thread
  hint when the first was lost; the per-(conversation, run_id) throttle
  suppresses genuine repeat posts. Tests added.
- ironclaw_conversations/CLAUDE.md: the idempotency guardrail now distinguishes
  transient submit failures (retry, rotate key) from thread-busy admission
  (terminal RejectedBusy, no retry-until-submitted, user resends).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): RejectedBusy terminal at storage, compaction safety, ledger no-fake-run

- threads: RejectedBusy removed from ensure_user_accepted in both backends —
  a stored RejectedBusy row can no longer transition to Submitted; resend is a
  fresh message. Prior-pass RejectedBusy->Submitted tests inverted to assert the
  transition is now terminal (InvalidMessageTransition). DeferredBusy admission
  kept (legacy replay still resubmits).
- compaction: only RejectedBusy (terminal) is SkipEphemeral; DeferredBusy moved
  back to DeferUntilStable since legacy rows can still reach Submitted — prevents
  a summary silently omitting a message that later becomes model-visible.
- workflow ledger: RejectedBusy { active_run_id: None } maps to
  ActionDispatchKind::NoOp instead of minting a fresh TurnRunId; still settles
  durably, no fabricated run id.
- inbound_turn test: the misnamed legacy-DeferredBusy test now actually injects a
  legacy DeferredBusy row and asserts resubmission, distinct from the RejectedBusy
  re-rejection test.
- openai-compat: added cancel-path RejectedBusy->429 handler test.
- slack_delivery: comments corrected to describe the recursive Duplicate{prior}
  extraction (no longer claim Duplicate{DeferredBusy} returns None).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): non-retryable 429 for terminal RejectedBusy + busy rename + coverage

Address open PR review findings on the busy-thread rejection work:

- openai-compat: split the busy ack arm so terminal RejectedBusy maps to
  429 retryable=false (client must issue a new request), while legacy
  DeferredBusy keeps retryable=true. Covers chat create, responses create,
  and responses cancel paths; regression unit tests on the retryable flag.
- slack_delivery: rename SLACK_DEFERRED_BUSY_* constants to SLACK_BUSY_*
  (path now serves RejectedBusy + legacy DeferredBusy); refresh the stale
  "silently dropped (pending gate)" comment to cover generic RejectedBusy.
- compaction_task: correct the StableNonModelVisible doc comment — only
  terminal RejectedBusy is skipped; legacy DeferredBusy is DeferUntilStable
  (it can still transition to Submitted).
- tests: add filesystem legacy DeferredBusy on-disk round-trip coverage and
  a reborn_services RejectedBusy mark-failure reconcile-via-replay test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(test): update root busy test to terminal RejectedBusy contract

The root integration test product_workflow_retries_after_filesystem_deferred_busy_release
still asserted the legacy DeferredBusy auto-resubmit contract (busy ->
DeferredBusy -> retry resubmits, submission_count 1->2). The live product
workflow now emits terminal RejectedBusy for busy user messages; the PR
updated crate-level tests but missed this root-level one, failing the
"Reborn root tests" CI job.

Rewrite + rename to product_workflow_rejects_busy_and_does_not_resubmit_on_filesystem_replay,
mirroring crates/ironclaw_product_workflow inbound_turn_contract's
rejected_busy_replay_is_re_rejected_not_resubmitted: first ack is
RejectedBusy (submission_count == 1); a same-event replay settles via the
idempotency ledger and returns Duplicate { prior: RejectedBusy } with no
resubmission (submission_count stays 1).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): terminal RejectedBusy cut point no longer blocks compaction

Review finding (High): the compaction terminal cut-point validation rejected
every SkipEphemeral disposition with InvalidCutPoint, but RejectedBusy now
classifies as SkipEphemeral(StableNonModelVisible). So a compaction range whose
drop_through_seq landed on a RejectedBusy message hard-failed — contradicting
this PR's goal that terminal RejectedBusy must never block compaction.

Allow a stable-non-model-visible terminal (RejectedBusy) as a legal cut point:
it is excluded from the compacted output like the in-range SkipEphemeral case
and compaction proceeds. Non-User Include and RejectInvalid still error.
Regression test: compaction_port_accepts_terminal_cut_point_that_is_rejected_busy.

Also add legacy_deferred_busy_mark_failure_reconciles_via_replay covering the
reconcile branch's legacy DeferredBusy replay path (RejectedBusy was already
covered).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): narrow terminal cut-point accept to StableNonModelVisible

Review follow-up: the terminal cut-point arm accepted SkipEphemeral(_) with a
wildcard, which would also admit CapabilityDisplayPreview as a valid terminal.
Only StableNonModelVisible (terminal RejectedBusy) should qualify. Match the
explicit variant so other ephemeral skip reasons fall through to
InvalidCutPoint and fail loud.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): reconcile only RejectedBusy as terminal + coverage/naming follow-ups

Address the latest review round (post origin/main merge):

- reborn_services: the mark-failure reconcile path treated legacy DeferredBusy
  as a terminal already-settled state. DeferredBusy is non-terminal (a later
  replay treats it like Accepted and can resubmit), so claiming terminal over
  it violated the no-resubmit guarantee. Drop DeferredBusy from the reconcile
  predicate — only RejectedBusy is terminal; a DeferredBusy row now surfaces
  the mark failure (503 retryable) instead of a false-terminal RejectedBusy.
  Flipped the legacy test to assert the surfaced error.
- compaction: add regression test that a terminal CapabilityDisplayPreview cut
  point returns InvalidCutPoint (only StableNonModelVisible is a legal terminal).
- fakes: FakeProductAdapter now records RejectedBusy in accepted_envelopes
  (durable like Accepted/DeferredBusy) so fake-based tests don't undercount.
- slack_delivery: arch-exempt comment updated from deferred-busy to busy-thread
  / RejectedBusy terminology.
- reborn_services_contract: retext a busy-submit test that asserts RejectedBusy
  but still labeled the path "deferred".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(test): correct scripted-helper comments — DeferredBusy is non-terminal

Follow-up to the reconcile fix: the DeferredBusyMarkFails scripted helper and
its replay branch still documented the old behavior (DeferredBusy "settles"
reconciliation). reconcile_terminal_duplicate now accepts only RejectedBusy as
terminal, so a DeferredBusy replay surfaces the mark error (Unavailable/503)
instead of a false-terminal RejectedBusy. Comment-only; logic unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(test): align sibling probe-count comments with 3-call reconcile flow

Follow-up: the replay_call_count field doc and rejected_busy_mark_fails() doc
still described the old 2-call probe sequence. Both scripted mark-fail helpers
return None on the first two idempotency probes and Some(..) on the third
(reconcile) probe (count <= 2 guard). Comment-only; logic unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): reject empty cut-point range instead of summarizing nothing

Review finding (Medium): making a terminal RejectedBusy a legal cut point opened
an edge — a range whose only message is that rejection (or any all-skip-ephemeral
span) produced an empty validated_messages, then still ran inference on an empty
prompt and persisted a meaningless summary artifact.

Guard after the deferred-reason check: if validated_messages is empty, return
InvalidCutPoint before build_input — nothing model-visible to summarize. The
deferred-reason early-return stays first so legitimate deferrals are unaffected.
Regression test: a range whose only message is a terminal RejectedBusy returns
InvalidCutPoint and never calls inference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test/docs(review): RejectedBusy command ack, ack_helpers, 429 spec, banner

Address review test/doc gaps on the busy-rejection work:

- product_command_workflow_contract: cover the command-dispatch path where
  command_service returns RejectedBusy -> UnsupportedActionKind -> terminal
  Rejected ack (previously only user-message RejectedBusy was tested).
- ack_helpers: unit-test that internal_refs_from_ack rejects RejectedBusy with
  the internal error (no internal refs bound for a terminal busy ack).
- docs/reborn/contracts/openai-compatible-api.md: document the busy 429 split —
  terminal RejectedBusy is non-retryable (client must issue a new request),
  legacy DeferredBusy stays retryable.
- reborn_services_contract: add the missing opening separator on the Legacy
  DeferredBusy test section banner.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): trim summary span to last visible message at a busy cut point

Review finding (Medium): when the compaction cut point is a terminal RejectedBusy,
the summary span ended at drop_through_seq, covering that non-visible message. The
thread backends' context builder skips any ReplaceRangeWhenSelected summary whose
span covers a non-model-context-visible message (summary_covers_hidden_content), so
the summary was persisted but never applied — a dead artifact.

Trim end_sequence to the last model-visible (Include'd) message's sequence so the
span excludes trailing non-visible terminals; the summary then applies. Folds the
empty-range guard into the same `validated_messages.last()` match (None => empty
range => InvalidCutPoint) — no production unwrap/expect. Regression test asserts a
[visible@1, RejectedBusy@2] range compacted through seq 2 yields a summary spanning
end_sequence=1, not 2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: pin command RejectedBusy error to ProductAdapterError::Internal

Review nit: the command-RejectedBusy test used a bare expect_err (any error).
Pin the concrete public variant: ProductAdapterError::Internal — which is what
ProductWorkflowError::UnsupportedActionKind maps to at the adapter boundary. The
kind string ("unsupported action kind: ...") is wrapped in RedactedString and
not exposed via Display, so Internal is the tightest assertable pin from the
public return type.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(threads): summary may span permanently-terminal non-visible messages

Review finding (Medium, egGm): summary_covers_hidden_content blocked a
ReplaceRangeWhenSelected summary whose span covered ANY non-model-context-visible
message. Compaction legitimately spans non-visible rows it skipped from the
summary content (e.g. an interior terminal RejectedBusy, or a capability preview),
so those summaries were silently dropped — the compaction-layer trailing trim
couldn't fix an interior hole.

Block the summary only when the span covers a non-visible message that can still
RESURFACE as model-visible (Draft / Interrupted / Superseded / DeferredBusy).
Permanently-terminal non-visible rows (RejectedBusy, CapabilityDisplayPreview
kind) never resurface, so spanning them is safe — the summary content already
excludes them and they are never shown in context. Identical change in both
in_memory and filesystem backends via a shared can_resurface_as_model_visible
helper; Redacted/Deleted keep blocking. Also corrects the pre-existing
capability-preview span behavior (two tests updated). Regression tests on both
backends: interior RejectedBusy summary is applied; interior Draft is not.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* fix(reborn): allow read-only GitHub capabilities (#4893)

* fix(reborn): allow read-only github capabilities

* test(reborn): future-proof github approval boundary

* fix(reborn): keep GitHub code search gated

---------

Co-authored-by: Robert Yan <mstr.raphael@gmail.com>

* clean up

* delete

* fix: run scheduled automations and report accurate status (#4920)

* fix(automations-ui): readable summary cards and NEXT RUN value

Reflow the summary strip to at most three cards per row so the detail
text no longer wraps one word per line, and let StatCard accept a
valueClassName override so the NEXT RUN date renders at a smaller size
instead of truncating to "Jun…". Default StatCard sizing is unchanged.

* fix(automations-ui): surface delivery save errors and gate Slack hint

The delivery-defaults panel swallowed save/clear failures and showed no
feedback; it now renders an inline error from the mutation and flashes the
"Saved" confirmation on Clear as well as Save. The "reply approve <code> in
Slack" footnote is hidden unless an external Slack-style target exists.

* fix(automations-ui): label sub-hourly cron schedules

Minute- and hour-level cadences such as "* * * * *", "*/15 * * * *", and
"0 * * * *" rendered as "Custom schedule" because they have no single clock
time. They now read as "Every minute", "Every 15 minutes", and "Hourly at
:00".

* fix(automations-ui): space the run-row action button icons

The "Open run" and "Logs" buttons in the recent-runs list rendered the
icon flush against the label because the non-primary Button variants don't
add a gap between children. Add the same icon margin the rest of the app
uses for icon+label buttons.

* fix(automations-ui): consistent summary counts and next-run

The Running/Failures summary cards counted individual runs while the
matching filter tabs counted automations, so the numbers disagreed; both
now count automations. The soonest "Next run" no longer includes paused
triggers, which keep a stored slot they will never actually fire.

* test(automations): lock the panel UI fixes into the served bundle

Add static-asset assertions driving the composed router so each Automations
panel UX fix — sub-hourly cron labels, summary card reflow + smaller NEXT RUN
value, run-row icon spacing, and delivery save-error/Slack-hint gating — is
guarded against a regression that drops it from the shipped SPA source.

* feat(automations): surface scheduler-off state and run it by default on serve

Scheduled automations never fired because the trigger poller is disabled by
default and nothing told the user. The list response now carries
scheduler_enabled (sourced from runtime readiness) and the panel shows a
"scheduling is turned off" notice when it is false. The local `ironclaw-reborn
serve` surface enables the poller by default; config and env still override it.

* fix(i18n): add automations.delivery.saveFailed to every locale pack

The new key was added only to en.js, which breaks the i18n consistency test
that requires all locale packs to share the English key set. Add it to the
ten other packs (English placeholder, matching the existing untranslated
automations strings there).

* fix(automations-ui): clear stale Saved flash before a new delivery write

The save-error alert is gated on !showSaved, so a "Saved" flash still
showing from a prior success would hide the error of a new failing
save/clear. Reset the flash (and its timer) at the start of every attempt.

* fix(automations): harden next-run filter and lock scheduler_enabled on the wire

Use loose `!= null` in the soonest-next-run filter so a missing
next_run_timestamp can't slip through, and assert scheduler_enabled in the
list-automations handler contract test so a serialization drift of the new
field is caught at the wire, not just in the facade.

* fix(i18n): add automations.schedulerOff keys to every locale pack

The scheduler-off notice keys were added only to en.js, which breaks the
i18n consistency test requiring all locale packs to share the English key
set. Add both keys to the ten other packs (English placeholder).

* i18n(automations): translate schedulerOff strings in all locale packs

The scheduler-off notice was English in every non-English pack, giving
Arabic/German/Spanish/French/Hindi/Japanese/Korean/pt-BR/Ukrainian/zh-CN
users a mixed-language UI. Provide real translations.

* i18n(automations): translate delivery.saveFailed in all locale packs

The save-failed delivery error was English in every non-English pack. Provide
real translations so users don't see mixed-language UI when a save fails.

* Localize automation summary counts

* Remove duplicate automation summary locale keys

---------

Co-authored-by: Robert Yan <mstr.raphael@gmail.com>

* fix(approvals): persist "always allow" across threads — drop thread_id from persistent approval scope (#4825) (#4835)

* make 'always allow' approvals persist (tested on google suite)'

* test(approvals): lock criterion-5 backward-compat for project-scoped policies (#4825)

* reduce slop

* fix(approvals): address approval scope review

* fix(approvals): preserve legacy approval lookup

* fix(approvals): find legacy grants across threads

* fix(approvals): drop legacy approval scope compatibility

* fix(approvals): simplify threadless policy lookup

---------

Co-authored-by: Emil Bogomolov <emil.bogomolov@near.ai>
Co-authored-by: Henry Park <henrypark133@gmail.com>

* [codex] Use WebUI base URL for OAuth callback origins (#4932)

* Fix Railway WebUI OAuth callback origin

* fix(reborn-cli): address OAuth base URL review

* fix(reborn-cli): fail closed on hosted oauth base url

* fix(host-runtime): accept empty body/body_base64 in builtin.http (#4827)

* fix(host-runtime): accept empty body/body_base64 in builtin.http

The HTTP tool's `body()` validator rejected any request that carried
*both* a `body` and a `body_base64` field, even when both were empty
strings. The JSON schema lists both fields, so models routinely emit
`{"body": "", "body_base64": "", ...}` as defaults for a bodyless GET.
That tripped the mutual-exclusion check and failed the call with
`InputEncode` before it ever dispatched — on every attempt — so the
agent could never make a successful request and would retry until its
loop gave up.

Treat a null or empty-string field as absent: only a non-empty value
counts as "set", so the mutual-exclusion check fires only when the
caller genuinely supplies two competing bodies. Behavior for a real
`body`, a real `body_base64`, or a genuine both-non-empty conflict is
unchanged.

Adds unit tests for body() covering empty-both, absent, string body,
base64 body, both-non-empty (rejected), and JSON-object body.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style: rustfmt the body() unit tests

---------

Co-authored-by: Pranav Raja <pranav.raja@near.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* feat(reborn): observability seams — trajectory observer + LLM provider injection (#4588)

* feat(reborn): expose a trajectory observer hook on RebornRuntimeInput

The reborn runtime is sealed: build_reborn_runtime returns only the final
AssistantReply, and per-step capability (tool) calls + results live in internal
stores. Downstream consumers (benchmark harnesses, UI/debuggers) can't observe
the agent's trajectory.

Add `RebornTrajectoryObserver` (pub trait: on_capability_input(call_id, name,
args) / on_capability_result(call_id, output)) and
`RebornRuntimeInput::with_trajectory_observer`. The local-dev capability IO
(`LocalDevCapabilityIo`) forwards each tool call's name+args (at input staging)
and result (at result write) to the observer when present — reusing the same
data it already records for display previews. No-op when unset; best-effort.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* debug: trace observer hook firing (temporary)

* feat(reborn): trajectory observer — capability_id on result, reliable spine

Provider tool calls are staged by a lower decorator that bypasses the
LocalDevCapabilityIo input path, so on_capability_input does not fire for
them. on_capability_result fires for every completed capability — make it
carry the capability_id so consumers can reconstruct the trajectory (name +
output) from results alone. Input args capture is a follow-up.

* feat(reborn): capture capability input args at the host port chokepoint

Provider tool calls are staged by ProviderToolCallInputResolver, which keeps
args in a private map and bypasses the capability-IO input hook — so inputs
never reached the trajectory observer (only results did). Move the observer
trait down to ironclaw_loop_support (CapabilityTrajectoryObserver, re-exported
from composition as RebornTrajectoryObserver) and hook it in
HostRuntimeLoopCapabilityPort::invoke_capability right after the input
resolves — the one place the model's resolved arguments are visible. Threaded
through HostRuntimeLoopCapabilityPortFactory + the local-dev factory. Result
hook unchanged. Now name + args + output are all captured.

* feat(reborn): host LLM-provider injection seam

ResolvedRebornLlm::with_provider — drive the runtime with a caller-supplied
LlmProvider (e.g. an instrumented wrapper that counts tokens/cost and captures
reasoning) instead of always building one from config; build_llm_gateway honors
the override. The only viable observability path for reborn, whose model calls
run in spawned worker tasks a per-task tracing subscriber can't reach.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(reborn): cover trajectory observer + LLM provider override seams

Addresses Firat's two blocking review findings on #4588 (both
missing-integration-test, per AGENTS.md "test through the caller"):

1. Trajectory observer callbacks — drive the real call sites with a
   recording CapabilityTrajectoryObserver:
   - host port: invoke_capability via HostRuntimeLoopCapabilityPortFactory
     ::with_trajectory_observer asserts on_capability_input fires with the
     resolved capability id + tool-call arguments.
   - local-dev IO: register_provider_tool_call_input + write_capability_result
     assert on_capability_input and on_capability_result fire and correlate by
     input ref.

2. LLM provider override — build_llm_gateway_drives_provider_override_not_config
   injects a counting mock via ResolvedRebornLlm::with_provider, points config
   at a dead endpoint, and asserts the gateway returns the mock's sentinel
   (proving the override is driven, not a config-built chain).

Also fixes pre-existing breakage this surfaced: 5 LocalDevLoopCapabilityPort
Factory test initializers (shell_tests.rs + tests.rs) were missing the
trajectory_observer field added by this PR, so the composition crate's tests
did not compile under --features root-llm-provider.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): make trajectory observer input semantics consistent

Addresses Copilot's follow-up findings on the observer seam:

- Drop the `on_capability_input` callback from `LocalDevCapabilityIo::
  register_provider_tool_call_input`. It forwarded the raw provider tool
  name (`builtin_echo`) as the capability id — conflicting with the
  observer contract (resolved dotted `builtin.echo`) and the authoritative
  port-level hook — and `ProviderToolCallInputResolver` doesn't delegate
  here for provider tool calls, so it never fired in practice anyway.
  `HostRuntimeLoopCapabilityPort::invoke_capability` remains the single
  source of `on_capability_input` (resolved id); `LocalDevCapabilityIo`
  remains the source of `on_capability_result`.

- Clarify the trait doc: `arguments` is the raw model-emitted tool-call
  input resolved from the input ref (the callback fires before schema
  normalization), which is what the trajectory should record.

- Refocus the local-dev test on `on_capability_result` forwarding +
  correlation, and assert input staging does NOT emit `on_capability_input`
  from local-dev IO. Port-level input semantics stay covered by the
  capability_port.rs test.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* wire trajectory_observer through RefreshingLocalDevCapabilityPortConfig

Completes the main-merge conflict resolution: local_dev.rs passes
trajectory_observer into the refreshing-port config, so the config struct +
port struct must carry it and build_inner must apply it via
.with_trajectory_observer(). (Missed staging this file in the merge commit.)

* test(reborn): lock down the observability seams against regression

#4588 exposes two seams a downstream harness relies on. Add tests so a
future refactor can't silently break either:

- capability_io_forwards_result_to_trajectory_observer: drives
  write_capability_result and asserts on_capability_result fires with the
  correct (call_id, capability_id, output) — the result half of the
  trajectory observer (tool-call outputs).
- build_llm_gateway_drives_provider_override_not_config: asserts the gateway
  drives a provider injected via ResolvedRebornLlm::with_provider (config
  points at a dead endpoint), proving the provider-injection seam works —
  this is how the bench captures reasoning / tokens / cost / system-prompt /
  tool-definitions. (Restores the test dropped during the main merge.)

The input half (on_capability_input) is already covered by
invoke_capability_forwards_resolved_input_to_trajectory_observer in
ironclaw_loop_support. All three pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): drop the false-confidence result-hook test

capability_io_forwards_result_to_trajectory_observer called
write_capability_result directly, so it stayed green even though the
result hook is unreachable end-to-end while capability dispatch fails
(the LocalDevYolo InputEncode regression) — i.e. it did not fail when
the feature it claimed to cover was actually broken.…
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
* feat(reborn): declare Slack channel as extension manifest

* fix(reborn): address extension review feedback

* fix(reborn): address Slack product-adapter extension review feedback

- Reserve "slack" as host-bundled extension id (prevent filesystem shadowing)
- Add builtin_first_party_trust_policy regression test for Slack admin entry
- Activate manifest-backed channel packages from WebUI (suppress only wasm_channel)
- Preserve legacy Slack connect controls for pre-install deployments
- Project only ProductSurfaceKind::ExternalChannel to channel kind
- Parse each manifest once via ExtensionManifestRecord
- Rename product_adapter.host_beta section to stable product_adapter.inbound
- Add caller-level tests: list_extension_registry, extension_info, ChannelsTab render

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(runtime-context): enable connected-channel classification via surface_kinds

nearai#4778 lands the ProductAdapter surface projection, so the lifecycle summary
now carries surface_kinds. Flip CHANNEL_CLASSIFICATION_AVAILABLE to true and
make extension_is_channel_surface a real predicate (ExternalChannel), so
connected channel names render in the model runtime context instead of unknown.

Convert the two stubbed tests to positive cases: empty list -> Known([]),
mixed list -> only active channel-surface extensions reported.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(runtime-context): cargo fmt + drop unreachable classification branch

Remove the dead 'if !CHANNEL_CLASSIFICATION_AVAILABLE' arm inside the
Some(Ok(response)) match: lifecycle_fut only issues the ExtensionList call
when classification is enabled, so a present response always means it is on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): repair Slack extension asset + locale checks after channel refactor

- assets.rs smoke test: assert showLegacySlackConnectActions (the refactor's
  built-in Slack status path) instead of the removed slackBuiltinStatus helper
- add extensions.kind.channel to all 10 non-en locales (en gained the key with
  the new channel surface kind; locale-parity test requires all locales match)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(webui-v2): cover ExtensionCard channel overflow + localize registry heading

Review findings (PR nearai#4778 review 4493764666):
- F2: add extension-card.test.mjs proving kind=channel/wasm_channel surface
  Setup (setup_required/failed) and Reconfigure (active/ready) overflow actions
  on the real component; channels-tab.test stubbed ExtensionCard so this was
  uncovered.
- F4: render the 'Available channels' registry heading via t(channels.availableChannels)
  instead of a hardcoded literal; add the key to all 11 locales. Update the
  channels-tab test to locate the registry section by the RegistryCard component
  (heading is now an interpolated value, not a template literal); drop the now
  unused renderedValueAfter helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(runtime-context): drop permanent classification flag, document surface_kinds cache

Review findings (PR nearai#4778 review 4493764666):
- F5: remove CHANNEL_CLASSIFICATION_AVAILABLE (permanently true after the stub
  flip) and run the lifecycle ExtensionList fetch unconditionally when a
  lifecycle facade is wired.
- F3: document AvailableExtensionPackage.surface_kinds as an intentional
  single-parse cache (re-deriving in summary() would re-run the manifest
  projection, undoing the parse-once optimization).

F1 (per-turn ExtensionList cost) accepted as-is: the fetch is spawned off the
critical path under a 500 ms budget with abort-on-drop; caching deferred to a
follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): suppress Activate for channel kinds during pairing

Review finding (PR nearai#4778 review 4494623523, finding 1): primaryExtensionAction
only suppressed the primary Activate button for legacy wasm_channel, so a
manifest-backed kind=channel Slack card fell through to 'activate' in
pairing_required/pairing states where the dedicated pairing section already
owns the flow. Suppress the primary action for channel-surface kinds in those
states (via isChannelExtensionKind); installed channels still return 'activate'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): add channels.slack key + regression tests for surface-kind projection

Review findings (PR nearai#4778 review 06:28):
- Add missing channels.slack i18n key to all 11 locales (legacy Slack row
  rendered the raw key because the i18n helper returns the key on miss; the
  || "Slack" fallback never fired).
- Add filesystem-path test: a /system manifest with
  product_adapter.inbound.surface_kind = external_channel projects to
  ExternalChannel surface (previously only the bundled catalog path was covered).
- Add extension_kind regression test: non-channel summaries keep their runtime
  wire kind (wasm_tool, mcp_server) while channel surfaces map to "channel".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): gate Slack catalog entry behind slack-v2-host-beta; test wasm_channel registry

Review findings (PR nearai#4778 review 07:01):
- Gate the Slack first-party catalog entry, its only-Slack symbols (slack_package,
  slack_assets, SLACK_MANIFEST, slack_manifest_digest), the factory trust-policy
  Slack AdminEntry, and Slack-asserting tests behind slack-v2-host-beta. Without
  the feature the Slack route/runtime/WebUI mounts don't exist, so the catalog
  must not advertise an unrunnable Slack extension. Clean clippy + tests in both
  feature-on and feature-off configs.
- Add useExtensions hook test proving an uninstalled kind=wasm_channel registry
  entry lands in channelRegistry, not toolRegistry (isChannelExtensionKind covers
  both channel and wasm_channel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): consolidate Slack trust policy tests (nearai#4778)

* feat(reborn): expose outbound delivery targets to model

---------

Co-authored-by: Henry Park <henrypark133@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@coderabbitai coderabbitai Bot mentioned this pull request Jul 8, 2026
17 of 30 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: dependencies Dependency updates scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants