Skip to content

fix(extensions): chat "connect account" dead-end — already-connected signal, builtin description trust, docs - #7361

Merged
BenKurrek merged 5 commits into
mainfrom
fix/chat-connect-guidance-and-usage
Aug 7, 2026
Merged

BenKurrek merged 5 commits into
mainfrom
fix/chat-connect-guidance-and-usage

Conversation

@BenKurrek

@BenKurrek BenKurrek commented Aug 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The bug: asking the agent to connect Slack from chat dead-ends. In the 2026-08-07 QA session (thread e79a994f, run 251aec0b on ironclaw-qa-testing-libsql, deployed commit 389aa99) the tester typed "connect account" and the model answered "I can't initiate the Slack OAuth flow from here directly — that needs to be done through the web interface" — even though the trace shows the model had already run the designed path (extension_search → extension_install(slack)), the activation credential gate had passed, and the tester's personal Slack account was therefore already connected. The truthful answer was "you're already connected"; instead the user spent ~29 minutes hunting the web UI. The repo's own QA spec requires this flow to work (qa_3a_slack_connect, scripts/reborn_webui_v2_live_qa/run_live_qa.py:144: ask IronClaw "connect to Slack", go through the flow, expected: connected).

Three causes, three fixes — all in the "model is confused about extension connect" cluster:

  • The product's own docs scripted the refusal (docs/onboard.mdx, docs/channels/overview.mdx). They still said "Asking the agent to connect a channel doesn't work… It may tell you it can't help; do it yourself in the web interface" — stale since the chat-side connect gate shipped in early July, and load-bearing because self_knowledge.md (injected every turn) makes docs.ironclaw.com the model's authority on IronClaw's own capabilities. The pages conflated the operator half (registering app/bot credentials — genuinely web-only) with the per-user OAuth half that extension_install drives through the same activation credential gate as the Channels card. Both pages now distinguish the two halves.
  • Nothing ever told the model "already connected" (ironclaw_extension_manager). When install-driven activation passes because the caller's declared credential requirements are all satisfied, the response carried only conditional guidance ("If WebChat shows an account connection panel… If the user's account is already connected, continue") — the gate's verdict was never surfaced, so the model couldn't tell which branch was true. The install response now appends an explicit already-connected confirmation exactly when declared requirements were verified present for the calling user, and never for credential-free extensions.
  • A lifecycle tool's description was silently scrubbed from every prompt (ironclaw_host_runtime::surface). Host-bundled capability descriptions carried description_trust = Untrusted, so the loop-tier prompt-text denylist strict-scanned compiled-in text and dropped builtin.extension_register_hosted_mcp from the prompt's capability surface on every turn (WARN each turn: "browser authorization-code flow" matched the "authorization" credential pattern). This is the same confusion surface — the disclosure protocol tells the model to use the lifecycle tools for connect/install requests, while the scanner was quietly deleting one of them from its instructions, and any future connect-flow capability whose description mentions OAuth would meet the same fate. HostBundled — the only source eligible for effective FirstParty/System trust — now maps to VerifiedCatalog like signature/digest-verified registry installs; InstalledLocal/UserRegistered/unknown keep the strict scan.

Change Type

  • Bug fix
  • Documentation

Linked Issue

None — diagnosed from the QA "Slack issues" report (thread e79a994f-55bd-56a8-87c3-9c9032405be2, run artifact 251aec0b, Railway ironclaw-rebornneari/qa-testing-libsql trace logs).

Validation

  • cargo fmt --all -- --check
  • cargo clippy -p ironclaw_loop_contracts -p ironclaw_host_runtime -p ironclaw_extension_manager --all-targets --all-features (clean; workspace -D warnings lane in CI)
  • cargo build (via test builds)
  • Relevant tests pass: ironclaw_host_runtime (446+37), ironclaw_extension_manager (141+2), ironclaw_loop_contracts (new pins), ironclaw_architecture_tests (55 suites, 0 failures), reborn_integration_extension_runtime (22 pass; the 2 case_2 Postgres legs require a local Docker daemon — not running on this machine — and fail identically without this diff; CI covers them)
  • Manual testing: root-caused from the deployment's trace-level logs (model request surface, gate behavior, per-turn scanner WARNs)
  • review-pr / pr-shepherd

Test Strategy

User behavior: asking the agent to connect an already-connected extension in chat gets a truthful "already connected — continue" instead of a deflection to the web interface; extension_register_hosted_mcp appears in the prompt's capability surface; the docs the model self-grounds on describe what chat can actually do.

Risk areas:

  • Model behavior
  • Security or permissions
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: instruction_bundle retain/omit pins for auth-vocabulary descriptions by trust level (loop_contracts); install_response_confirms_connection_only_when_caller_credentials_were_verified (manager)
  • Reborn integration: reborn_integration_extension_runtime exercised (no assertions changed); standalone_agent_surface_exposes_extension_lifecycle_tools extended to pin VerifiedCatalog trust for all model-visible lifecycle capabilities through the real host runtime; standalone_extension_activate_hosted_mcp_stages_discovery_and_publishes_tools extended to pin the already-connected confirmation on the seeded-credential path; credential-free install pinned to NOT claim a connection
  • Recorded fixture: Not applicable — swept tests/fixtures/ for the changed strings; none pinned
  • Browser E2E: Not applicable — no frontend change; the in-chat connection-panel path (channel_connection_required) is untouched
  • Backend or runtime: covered by the manager surface test driving the real host-runtime registry; persistence unchanged
  • Live canary: Not applicable

What the tests prove: each fix fails without its change — the trust assertions were RED (Untrusted) before the surface mapping change, and the connect confirmation was RED (message absent) before the install-arm change; the untrusted-provenance strict scan is pinned so the safety posture cannot silently widen.

Commands run: cargo test -p ironclaw_host_runtime, cargo test -p ironclaw_extension_manager, cargo test -p ironclaw_loop_contracts --lib, cargo test -p ironclaw_architecture_tests, cargo test -p ironclaw_integration_tests --test reborn_integration_extension_runtime, python3 scripts/ci/docs_publication_boundary.py, cargo fmt.

Compatibility, rollback, follow-ups

Compatibility. The trust widening is scoped to HostBundled (compiled/shipped with the binary — descriptions are repo-authored constants; structural prompt limits still apply on every variant, and untrusted provenances keep the full denylist). The install message change is additive prose on the model-visible result; no wire schema change. Docs regenerate llms.txt on deploy, which also updates what the self-knowledge protocol serves the model.

Rollback. Each commit reverts independently; no persistence, schema, or config changes.

Follow-ups (out of scope, from the same diagnosis).

  • NEAR AI streamed runs record zero usage / $0 cost (stream_options.include_usage never sent) — fix is written and verified against the live endpoint, staged separately from this PR.
  • Run-artifact exports carry logs: {available: true, entries: []} — the operator in-memory tracing ring loses run lines within seconds of completion.
  • Model-gateway spans label provider_id=unconfigured on the NEAR AI backend.
  • QA env has unconfigured OAuth vendors (github, mcp-*) hitting OAuth gate vendor unavailable; falling through — qa_4b_github_connect will reproduce the same complaint class.
  • Consider the same already-connected confirmation on the API-only extension_activate continuation arm.
  • Re-run the manual qa_3a chat-connect journey on the QA deployment once this deploys.

🤖 Generated with Claude Code

…nstall confirmation

Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread
e79a994f, run 251aec0b on ironclaw-qa-testing-libsql):

1. Host-bundled capability descriptions were description_trust=Untrusted,
   so the loop-tier prompt-text denylist strict-scanned compiled-in text
   and silently omitted builtin.extension_register_hosted_mcp from every
   model prompt's capability surface ("browser authorization-code flow"
   matched the "authorization" credential pattern). HostBundled is the
   only source eligible for effective FirstParty/System trust, so its
   repo-authored descriptions now cross the verified-catalog boundary
   like signature/digest-verified registry installs. Untrusted provenance
   (InstalledLocal, UserRegistered, unknown) keeps the strict scan.

2. When install-driven activation passed the credential gate because the
   caller's declared requirements were all satisfied, the response never
   said so — the model got only conditional guidance ("If WebChat shows
   an account connection panel...") and deflected an explicit "connect
   account" request to the web interface even though the account was
   already connected. The install response now appends an explicit
   already-connected confirmation exactly when declared requirements
   were verified present for the calling user.

Regression tests: manager surface test pins VerifiedCatalog trust for all
model-visible lifecycle capabilities through the real host runtime;
instruction-bundle tests pin retain/omit behavior for auth-vocabulary
descriptions by trust; install-path tests pin the confirmation on the
seeded-credential path and its absence for credential-free extensions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@railway-app

railway-app Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-7361 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Aug 7, 2026 at 7:27 pm

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7361 August 7, 2026 16:37 Destroyed
@github-actions github-actions Bot added the scope: docs Documentation label Aug 7, 2026
@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Review was skipped due to path filters

⛔ Files ignored due to path filters (4)
  • tests/snapshots/golden_payload__context_surfacing.snap is excluded by !**/*.snap, !tests/snapshots/**
  • tests/snapshots/golden_payload__gated_turn_approve.snap is excluded by !**/*.snap, !tests/snapshots/**
  • tests/snapshots/golden_payload__parallel_tool_calls.snap is excluded by !**/*.snap, !tests/snapshots/**
  • tests/snapshots/golden_payload__tool_call.snap is excluded by !**/*.snap, !tests/snapshots/**

CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including **/dist/** will override the default block on the dist directory, by removing the pattern from both the lists.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 5ea496ba-24df-484d-a188-8b3720b5fece

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Agent-assisted channel setup can now install and activate extensions, open OAuth or pairing panels, and confirm when an account is already connected.
    • Capability descriptions now indicate whether their information comes from a verified source.
  • Bug Fixes

    • Untrusted capability descriptions are filtered from visible capability information.
  • Documentation

    • Clarified Slack and Telegram setup steps, including operator credential configuration and personal account connection requirements.

Walkthrough

Host-bundled capability descriptions now use verified catalog trust. Install activation reports an existing connection when verified credentials are already satisfied. Tests cover filtering and activation behavior. Channel documentation separates operator credential setup from personal connection flows.

Changes

Capability trust and activation guidance

Layer / File(s) Summary
Capability trust classification and filtering
crates/kernel/.../surface.rs, crates/contracts/.../instruction_bundle.rs
Host-bundled descriptions use VerifiedCatalog trust. Tests keep verified descriptions visible and omit untrusted authentication-related descriptions.
Credential-aware activation response
crates/extensions/.../extension_lifecycle_capabilities.rs
Activation passes caller-credential verification state and appends an “already connected” message only for verified credentials. Tests cover credential-free and hosted MCP activation.
Channel setup guidance
docs/channels/*, docs/onboard.mdx
Documentation separates operator credential registration from agent-assisted personal OAuth or pairing setup for Slack and Telegram.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant User
  participant ExtensionActivation
  participant CapabilitySurface
  participant OAuthOrPairingPanel
  User->>ExtensionActivation: Ask agent to connect a channel
  ExtensionActivation->>CapabilitySurface: Expose trusted activation capability
  ExtensionActivation-->>User: Report activation and connection state
  ExtensionActivation->>OAuthOrPairingPanel: Open personal OAuth or pairing when required
Loading

Possibly related PRs

  • nearai/ironclaw#7253: Both changes update hosted MCP capability descriptions and their trust or visibility semantics.
  • nearai/ironclaw#7264: The changes address extension-surface trust metadata and host-runtime capability ownership.

Suggested reviewers: thisisjoshford

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description covers the change, validation, tests, compatibility, rollback, and follow-ups, but omits several required template sections. Add explicit Security Impact, Reborn Trust-Boundary Checklist, Database Impact, Blast Radius, Review Follow-Through, and Review track sections.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title follows Conventional Commits style and accurately summarizes the connection confirmation, trust classification, and documentation changes.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Aug 7, 2026
@ironloopai

ironloopai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

🧭 IronLoop Run · Review

This comment updates in place as the Run moves through its stages.

⬛ Final result · Stopped

🟨 Queued → 🟦 Working → ⬛ Stopped

Automatic trigger · attempt 1 of 3 · stopped after 5m 56s

IronLoop stopped because the pull request target branch or head changed while this Run was active.

Run details

Run: e5af84da-e7b3-4578-8f38-8043d9a0f35b
Base: main at 81724a6
Head: fix/chat-connect-guidance-and-usage at f53ebed
Created: 2026-08-07 16:38 UTC
Updated: 2026-08-07 16:44 UTC

serrrfirat
serrrfirat previously approved these changes Aug 7, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/channels/overview.mdx`:
- Around line 96-107: Keep docs/channels/overview.mdx lines 96-107 as the
canonical distinction, update docs/channels/slack.mdx lines 45-62 to state that
WebUI registers instance credentials while chat installs or activates the
extension and completes or confirms personal OAuth, and align docs/onboard.mdx
lines 179-183 with the same two-step contract; do not claim that agent-driven
Slack connection is unsupported after operator setup.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: f935b6a8-6ea8-4028-9b4e-2a2472d6ca05

📥 Commits

Reviewing files that changed from the base of the PR and between 81724a6 and f53ebed.

📒 Files selected for processing (6)
  • crates/contracts/ironclaw_loop_contracts/src/instruction_bundle.rs
  • crates/domains/ironclaw_llm/src/nearai_chat.rs
  • crates/extensions/ironclaw_extension_manager/src/extension_lifecycle_capabilities.rs
  • crates/kernel/ironclaw_host_runtime/src/surface.rs
  • docs/channels/overview.mdx
  • docs/onboard.mdx

Comment thread docs/channels/overview.mdx
The onboarding and channels pages claimed "asking the agent to connect a
channel doesn't work" and that the agent "may tell you it can't help".
That describes only the operator half (registering app/bot credentials).
The per-user half has shipped since early July: extension_install runs the
same activation credential gate as the Channels card, raises the in-chat
OAuth connection panel when the account is unconnected, and (as of the
sibling fix) confirms when it is already connected.

The self-knowledge protocol makes these pages the model's authority on
IronClaw's own capabilities, so the stale claim scripted the exact
refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an
account that was already connected. Correct both pages to distinguish
the operator step from the chat-drivable personal connect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@BenKurrek
BenKurrek force-pushed the fix/chat-connect-guidance-and-usage branch from f53ebed to b5b1366 Compare August 7, 2026 16:44
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7361 August 7, 2026 16:44 Destroyed
@BenKurrek BenKurrek changed the title fix: chat connect-intent truthfulness, host-bundled description trust, NEAR AI streaming usage fix(extensions): chat "connect account" dead-end — already-connected signal, builtin description trust, docs Aug 7, 2026
… contract

The slack page's operator-step note and the telegram troubleshooting
accordion still carried the blanket "asking the agent to connect will not
work" claim the overview/onboarding correction removed — same drift,
different phrasing (review catch on #7361, plus one more instance found
by a broader sweep). Both now state the two-step contract: the operator
half stays in the web interface; after it, chat drives the personal half
(install/activate -> in-chat connection or pairing panel, or an
already-connected confirmation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7361 August 7, 2026 16:48 Destroyed
@BenKurrek
BenKurrek enabled auto-merge August 7, 2026 16:52
serrrfirat
serrrfirat previously approved these changes Aug 7, 2026
@BenKurrek
BenKurrek added this pull request to the merge queue Aug 7, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Aug 7, 2026
@BenKurrek
BenKurrek added this pull request to the merge queue Aug 7, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Aug 7, 2026
BenKurrek and others added 2 commits August 7, 2026 15:18
…ange

The queue run failed golden_payload because the branch predated current
main and its own surface.rs trust fix changes the surface digest. The
recaptured snapshots differ ONLY in the surface sha256 lines (verified
char-by-char) — no prompt text or capability-list changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-7361 August 7, 2026 19:19 Destroyed
@BenKurrek

Copy link
Copy Markdown
Collaborator Author

Queue failure diagnosed and fixed (pushed e541958c63):

  • The branch sat on pre-feat(reborn): enable progressive tool disclosure by default #6958/pre-fix(telegram): accept /pair as a pairing-code alias #7363 main, and this PR's own description-trust fix intentionally changes the surface digest — but the four golden_payload snapshots were never recaptured (the PR lane's affected-scope didn't select that bucket; the merge queue runs everything, so it failed there first).
  • Fix: merged current main into the branch and recaptured. Verified char-by-char that the snapshot delta is only the surface sha256: lines — no prompt text or capability-list changes. reborn_integration_golden_payload 20/20 locally on the merged tree.

Unrelated but observed in the same window: main's push run at df90072c4e failed the extension-host coverage floor while the identical commit had passed the full workflow in the merge queue ten minutes earlier — covered-line variance (~45 lines) beyond the 20-line tolerance, retry in flight. If it recurs, that floor needs a recapture/tolerance bump, not a code fix.

Needs a fresh re-queue when checks come back green.

@BenKurrek
BenKurrek added this pull request to the merge queue Aug 7, 2026
Merged via the queue into main with commit d27dba3 Aug 7, 2026
44 checks passed
@BenKurrek
BenKurrek deleted the fix/chat-connect-guidance-and-usage branch August 7, 2026 19:54
BenKurrek added a commit that referenced this pull request Aug 7, 2026
Main's #7361/#7363 landed 66 lines in
`ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked
up when folding onto main. That put the crate at 13,181 against the 13,115
ceiling re-captured earlier in this PR.

Not growth from this PR. The gate's upward check is a hard `lines > ceiling`
with no headroom — `TOLERANCE` (400) governs only the downward
ratchet-nudge — so a ceiling captured at the exact observed value reddens
every open branch the moment anyone adds a line to that crate, including
from main. Re-captured at the measured value; count read from the gate's own
failure message.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BenKurrek added a commit that referenced this pull request Aug 7, 2026
…the audit

The strongest single exhibit for shortlist item 2, contributed by the
#7157 branch steward after this audit's cutoff and verified against the
gate's code: TOLERANCE = 400 is consulted in exactly one direction (the
banked-slack check, ceiling.saturating_sub(lines) > TOLERANCE); the
growth check is a bare lines > ceiling. With the in-file 'set to
current, not padded' instruction, every ceiling is a hard cap at the
observed count — so one line landing on main in any contracts crate
reds every open branch at its next fold until someone re-captures.

Measured recurrence on #7157: loop_contracts re-captured four times,
~once per fold (14,479 -> 13,850 -> 13,949 -> 13,115 -> 13,181), the
last tripped by main's #7361/#7363 adding 66 lines to
instruction_bundle.rs — nothing the branch wrote. All four deltas were
<= 105 lines: either repair shape in §3.2 (one-line upward tolerance
using the existing constant, or mid-window pinning) would have absorbed
every one with zero red builds. This audit's own sabotage already
proved the jaws (+1 line host_api red / -1 line common red); the fold
history shows the operational cost. The repair stays a recommendation —
adding growth headroom to a ratchet is the owner's call, not this PR's.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
personal-upstream-sync Bot pushed a commit to theredspoon/ironclaw that referenced this pull request Aug 7, 2026
…ls, delivery heuristics deleted (nearai#7157)

* feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted

Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2
composition inversion, the WS6 crate renames, and the WS7 family moves), with
all 44 review comments dispositioned.

Two-lane delivery model: a run's final reply always lands in its own
conversation (lane 1); reaching any other surface is the model's explicit
`builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target
per call, synchronous through the DeliveryCoordinator, provider-issued
message refs as evidence). Background-run notices fan out to a user-configured
notification-channel set (new record field + read-side legacy migration,
`builtin.notification_channels_set`, first-approve-wins, WebUI multi-select).
The stored delivery heuristics are deleted (route_current, builtin:web_app,
outbound_delivery_target_set, per-trigger delivery_target_id + precedence
chains + four-slot preference fallback), with an idempotent boot migration of
stored trigger targets into explicit prompt steps and the retired vocabulary
pinned in reborn_retired_taxonomy.rs.

Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition,
ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages,
run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→
ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire.
The model-delivery implementation moved extension_host→assistant
(CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host
naming product types; the deferred-slot registration and post-coordinator bind now
live in composition's production assembly, mirroring TriggeredRunDeliveryDriver.

Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary
re-imported from ironclaw_extension_contracts, product-adapter/inbound
vocabulary from host_api/product_contracts, module-charter map's outbound
row renamed to the two-lane vocabulary. Post-branch CI gates adapted in
the same change: skills/ classified in the PR test planner (test-first,
sabotage-verified), panic baseline ratcheted down, nested test fixtures
renamed to the scanner-sanctioned support_tests.rs shape, composition's
inline trigger-migration tests split to tests.rs (mass budget green with
no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21:
the ActivePreferenceTargetCodecs port), loop_contracts ceiling
re-captured down 14479 -> 13850 after the delivery-vocabulary deletion.

Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging
framework): the two-lane guidance moved into the canonical messaging core
prompt (host_api prompts/messaging/send_message.core.md), now naming
builtin__outbound_deliver with the arrive-twice and trigger caveats for
every messaging extension; slack vendor addendum/manifest taken as nearai#6831
shipped them; ceiling-table union (host_api 18570 beside this PR's two
re-captures); retired slack schema embed and deleted preferences
capability stay deleted; golden context-surfacing snapshot regenerated
(one surface-hash line).

Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch +
sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230
beside this PR's re-captures) and main's tracing-target syntax sweep
(target = -> target:, gate-enforced) applied over this PR's kept lines;
deleted delivery-heuristic code stays deleted.

Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep):
zero conflicts; guidance/doc-pointer changes auto-merged over this delta.

Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing
is prompt-owned with a pinned source-surface default (bare "send me" =
the surface you asked from; web app = no delivery step; explicit
destinations override, one delivery step each) — iterated against live
recordings until a real model followed it, with two live-recorded QA
fixtures (bare-webui, multi-channel) plus contracts and replays. The
automations-page panel is retained as the notification-channel selector
(notices only); the conversational notification_channels_set tool writes
the same validated set.

Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow
on main): mark_terminal reports whether the durable write committed and
a confirmed send whose Delivered row failed to commit returns
DeliveredUnconfirmed (refs retained, durably_recorded: false), never a
fabricated Delivered — regression-tested and sabotage-verified. Plus a
CodeRabbit triage batch: correctable coordinator errors stay
model-visible, omitted target_ids no longer clears the set, the success
schema requires evidence, the composition outbound facade is dissolved,
and guidance/contract docs are aligned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: harden channel delivery routines and replay

* fix: preserve automation loading and identity freshness

* fix: close channel delivery review defects

* fix: skip paused routine catch-up slots

* ci: record channel delivery composition budget

* test: align composition baseline with channel delivery

* fix(auth): survive interrupted OAuth callbacks

* fix(auth): keep callback coordination panic-free

* fix(ci): reconcile channel delivery merge seams

* fix(ci): recapture merged contracts ceiling

* fix(delivery): close review findings across delivery, migration, and guidance

Fixes the findings from the multi-agent review of this PR. Every behavioral
fix ships with a regression test that fails before it.

CI (red on this head)
- `standalone_yolo_notification_channels_set_bypasses_approval_gate`
  expected the shared "invalid outbound delivery request" summary for a
  `builtin__notification_channels_set` call. Production deliberately
  specializes that message per operation and pins it with
  `notification_channel_failure_names_the_operation_the_model_can_correct`;
  the assertion was the stale side.

Delivery evidence (kernel + assistant + outbound)
- `AlreadyDelivered` replays reported `delivered: false` "unverified"
  because the ledger row retains no provider refs, inviting the duplicate
  resend the at-most-once claim exists to prevent. Evidence gained
  `already_delivered`; a replay now reads as delivered with an honest
  "not resent" summary. The classification suite had no `AlreadyDelivered`
  case at all, which is why this shipped.
- `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually
  sent something, but `delivered_messages_from_outcome` dropped its refs,
  so gate reply-routes went unrecorded and a live OAuth prompt could never
  be retracted on that path.
- `content` is now rejected when empty: the input schema advertised
  minLength 1 and nothing enforced it, so empty content reached the channel
  as an empty part and returned an opaque provider error.

Background-run notifier (assistant)
- One run legitimately emits several `RunBlocked` notices (re-auth
  stand-in, unserviceable-auth cancellation, run failure), but all three
  derived the same projection ref, and the delivery id hashes it. The
  second notice to a target came back `AlreadyDelivered`, was treated as
  success, and was never sent — a user could be told a routine needed
  re-authorization and never told it then failed. Notices carry a
  discriminator; once-per-run kinds keep their historical id shape, so
  existing delivery identities are unchanged.
- When every catalog lookup failed, the empty result was recorded as
  `NoDefaultConfigured`, reporting a backend outage as the benign "user
  configured nothing" state. It now records `Failed`.

Boot migration (composition)
- The retired `builtin:web_app` target meant "no external delivery". It was
  being rewritten into a delivery step to an id nothing can resolve,
  inverting the stored intent on every later fire. It now clears without
  adding a step.
- One unmigratable row aborted the entire composition boot, with the error
  telling the operator to shorten a prompt through the UI that no longer
  starts. It now pauses its own routine — a paused trigger cannot fire, so
  "never fire unrouted" still holds per record — and boot continues. Only a
  systemic store failure stays boot-fatal. A row deleted during the CAS
  retry ends that record instead of failing boot.
- The CAS retry loop, its bounded exhaustion, and the vanished-row arm had
  no caller-level coverage; adds a delegating repository double that forces
  CAS misses. The prior fail-closed test is rewritten to pin the invariant
  it documented (route survives, record not half-migrated) under the new
  per-record mechanism.

Model-visible messages (composition)
- The targets-list denial said "not permitted to change the outbound
  delivery target" for a read-only call, and the lease denial named the
  retired delivery-target concept on the notification-channel path that is
  its only production caller. Both are now operation-specific and pinned.

WebUI (frontend)
- `setNotificationChannels()` with no argument posted `target_ids: []`,
  turning an omitted argument into a destructive clear-all and defeating
  the backend contract that deliberately rejects an omitted field.
- The notification-channels panel stayed editable after a failed read, so
  toggling one row full-replaced the stored set from an empty baseline and
  silently dropped every channel the user never saw. Editing is now locked
  on a failed read, with a rendered explanation.
- Adds the missing `tools.description.builtin.notification_channels_set`
  key to all 11 locales, plus save-failure coverage for the hook (which was
  correct, but untested) and locale-parity tests.

Guidance
- The new `.claude/rules/tools.md` was ported from a pre-restructure branch:
  it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never
  existed), and its review command grepped three paths removed by WS6/WS7.
  Its `paths:` frontmatter also never matched the product/composition
  callers its rules govern, so the rule never loaded for them.
- `ironclaw_loop_contracts` now records both embedded prompt assets; this
  PR added a second one while the crate's Known-debt entry still said one.
- Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in
  this PR which both bumped) and fixes a pre-rename path in the
  extension-runtime checklist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): unify the DM-rule enforcement point and re-ratchet composition

CI (composition mass budget, red on the previous head): the per-record
migration quarantine pushed composition 6 LOC over its absolute ceiling.
Resolved by the reduction the budget file itself blesses rather than a
raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split
verbatim into `runtime/approval/tests.rs` (the gate excludes test-only
files but counts inline test modules). Composition is now 40,432 LOC,
smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the
arch-test record are re-captured together at the measured value per the
gate's one-directional ratchet rule.

The codec-scan that decodes a binding and enforces "an OAuth authorization
URL only ever lands in a personal DM" existed twice — once in
`TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` —
with both copies commented as "the single enforcement point". They are now
one implementation, shared by the notifier and `builtin.outbound_deliver`,
with a context label so each path keeps its own diagnostic.

That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`)
failed no test in the crate. The vendor codecs pin the predicate in
isolation and the coordinator test pins rejection handling with a double
that decides the verdict itself, so nothing covered the wiring that joins
them. Adds a contract test driving the real resolver through
`DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target
never reaches the vendor adapter. Sabotage-verified: the test fails with
the rule disabled and passes with it restored.

Smaller findings: the notification-channel schema cap now derives from
`ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring
`8`; `triggered_run_delivery`'s module and trait docs described the retired
result-push model this PR deletes; the two new notification strings used a
different brand spelling and dash style from the nine siblings in their own
module; and several new comments navigated by pre-rename paths
(`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a
citation of a test symbol that does not exist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): scope the delivery catalog to the authenticated actor

`builtin.outbound_deliver` resolved its destination catalog under
`ResourceScope.user_id` while performing the send as
`authenticated_actor_user_id`. Those are the same user on a personal thread
and on an automation fire, but they diverge on a shared-route channel
conversation: the scope user is the route's SUBJECT
(`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the
message. Any participant of such a channel could therefore name the
subject's target ids and push bot-identity content into the subject's own
destinations — their personal DM included — from a conversation the subject
may never read.

The catalog now follows the actor, so a caller stays inside their own
connected surfaces on every path and an unfamiliar target simply does not
resolve. Behavior is unchanged wherever owner and actor already agree,
which is every non-shared-route path.

Regression test drives the divergent case through the port (participant
denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a
control proving the owner's own delivery still works. Sabotage-verified:
restoring owner-scoping fails it.

NOT changed here, and flagged for a product decision: the sibling
`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derive their caller from the same
owner-preferring `effective_user_id`, so on a shared route a participant
can still enumerate — and, with the approval gate auto-approved, rewrite —
the subject's notification channels. That helper also scopes approval
gates and capability leases, so flipping its precedence risks breaking
approval raise/resume matching in a path no test covers; it needs its own
change with that coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): stop rewriting the DM-target row on every message

The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert`
for every admitted inbound direct message, and the store unconditionally
wrote a fresh row. After the first message the stored record is already
correct, so the steady state was one durable backend write per DM message,
forever, whose only effect was a new `updated_at` — and each message's
reply-delivery observation was serialized behind it. An unchanged record
now short-circuits; the existing row is loaded here anyway to preserve
`created_at`, so the comparison costs nothing.

Also adds the regression test the `NoDefaultConfigured` -> `Failed`
classification fix landed without: the notifier's `SkipEntry` lookup lane
had no coverage at all (no test ever made a catalog lookup error), so
neither the skip nor the all-failed arm was exercised. The triggered
harness gains an injectable catalog provider for it. Sabotage-verified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(loop): bound the connected-channels line so the runtime slice fits

Confirmed live, not theoretical: a worst-case runtime context renders 4,391
bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that
`instruction_bundle::push_runtime_context` validates the whole slice on —
and exceeding it is a run-ending error on EVERY prompt build for that user,
not a one-off.

This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a
previously-fitting context over. The individual parts are each bounded
(location 200 chars at its producer, locale 35, per-label safe-text
validation), but nothing bounded their SUM, and the connected-channels line
is the one part that grows without limit: up to 20 entries whose names and
presentation hints are only individually capped.

That line now renders as many channels as fit a 1 KiB budget and folds the
rest into the "+N more" counter it already carried, so the fixed guidance
can never be squeezed out by variable content. The worst-case test that
found this stays as the pin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): resolve outbound capabilities as the acting user

`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derived their caller from
`effective_user_id`, which prefers the thread owner over the actor. Those
agree on a direct message and on an automation fire, but diverge on a
shared-route channel conversation, where the owner is the route's
configured subject — the deployment operator by default
(`channel_workflow.rs`) — and the actor is whoever posted.

So any participant of a shared channel could enumerate the operator's
connected destinations and rewrite the operator's notification-channel set,
which is where approval prompts, re-auth prompts and failure notices are
delivered. The caller now follows the acting user, matching the fix already
applied to `builtin.outbound_deliver`.

This deliberately REVERSES a previously pinned preference. Two tests
asserted the owner won when the two differ; that pin predates shared-route
subjects defaulting to the operator, and it contradicts the rule that a run
acts as whoever invoked it. Both are updated to pin the actor, with the
reversal recorded at each site rather than silently relaxed, and the
notification-channel write is now asserted to land under the acting user
with the thread owner's own set left untouched.

INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run`
still follow the owner, because they scope the approval-gate raise and the
capability lease and those must stay matched between raise and resume.
Unifying them belongs with the follow-up that removes shared-route subject
binding entirely so a shared channel runs wholly as its invoker; that needs
approval raise/resume coverage which does not exist yet. A new test pins
the split so the interim state is explicit rather than accidental.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep loop_contracts under its size ceiling and ratchet down

The runtime-context byte-budget fix and its worst-case pin pushed
`ironclaw_loop_contracts` to 14,032 production lines against a 13,949
ceiling. Resolved by the reduction the gate prefers over a raise:
`runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim
into a `runtime_context/tests.rs` sibling, which `production_rust_files`
excludes (an inline test module inside a production file is counted; a
test-only file is not).

The crate now measures 13,115 — 834 lines below the previous ceiling and
smaller than before this review round — so the ceiling is re-captured
downward at the measured value rather than raised, per the gate's
one-directional ratchet. Count read from the gate's own failure message,
not by eye.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): chart the notification-channel handlers in the WebUI charter map

`handlers_module_charter` failed: `get_notification_channels` and
`set_notification_channels` — handlers this PR adds — had no sub-owner row
in `CONTRACT.md`'s enforced charter map, and the row they belong to still
named `get_outbound_preferences`, `set_outbound_preferences` and
`outbound_preferences_activity_id`, all deleted by this PR.

Both halves are fixed together because the gate checks both in one test:
unclaimed items first, then entries naming items that no longer exist. Only
the first had fired, so the stale half was still latent behind it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): re-capture the loop_contracts ceiling after main's merge

Main's nearai#7361/nearai#7363 landed 66 lines in
`ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked
up when folding onto main. That put the crate at 13,181 against the 13,115
ceiling re-captured earlier in this PR.

Not growth from this PR. The gate's upward check is a hard `lines > ceiling`
with no headroom — `TOLERANCE` (400) governs only the downward
ratchet-nudge — so a ceiling captured at the exact observed value reddens
every open branch the moment anyone adds a line to that crate, including
from main. Re-captured at the measured value; count read from the gate's own
failure message.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
personal-upstream-sync Bot pushed a commit to theredspoon/ironclaw that referenced this pull request Aug 8, 2026
…mpt-surface PRs (nearai#7371)

* fix(ci): schedule the golden lane for prompt-surface production changes

A production change to the model-visible prompt surface (the capability
surface digest, the instruction bundle, the communication-context
renderer, or a shipped loop-tier prompt asset) ran only crate buckets on
the PR lane, so stale golden_payload snapshots surfaced first as a
merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface
digest; the PR lane never ran the golden bucket). Add a curated
prompt-surface owner table that ADDITIONALLY schedules the golden
integration lane without consuming the path's normal package
classification. Self-tested per entry plus a negative control pinning
that ordinary production changes keep the narrow plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): recapture the extension_host coverage floor with measured wobble tolerance

The 2026-08-04 entry was an adjustment whose own text ordered the next
real recapture. Since then code left the crate (denominator 24529 ->
24415) while the ratio ROSE to 88.12%, and the same commit df90072
measured >=21560 covered lines in the merge-queue lane but 21515 twice
on the push lane — a >=45-line same-commit spread over a 20-line
tolerance, redding main on noise (run 31208592262). Recapture both
fields from that run's own gate output and size tolerance_lines to the
measured cross-lane wobble.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): cover the host-managed prompt composer and derive per-entry tests

Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives
the InstructionBundleBuilder, model.rs shapes the pinned request) was
missing from the prompt-surface table — the exact gap class the mapping
exists to close. And the self-test enumerated entries by hand, so a new
entry could ship untested. Add the host_managed_ports prefix and derive
the positive cases from the tables themselves; an entry whose crate the
fixture lacks now fails the suite explicitly instead of skipping.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@serrrfirat serrrfirat mentioned this pull request Aug 10, 2026
29 tasks
personal-upstream-sync Bot pushed a commit to theredspoon/ironclaw that referenced this pull request Aug 12, 2026
…ead gates deleted (nearai#7373)

* test(architecture): drop the dead ironclaw_storage row and arm the substrate list

Gate-audit finding (open-and-shut): SUBSTRATE_CRATES in
reborn_composition_boundaries.rs carried three rows of rot, all invisible
because the loop's `let Some(..) else { continue }` silently skipped any
entry that resolves to no workspace package:

- "ironclaw_storage": no such package exists (verified against
  `cargo metadata --no-deps`; the only MISSING name of the 29 listed).
- "ironclaw_approvals" and "ironclaw_assistant" were each listed twice.

The silent skip is replaced with a panic naming the stale entry, so the
list can no longer rot invisibly. Verified by sabotage: adding a bogus
"ironclaw_zzz_probe" row now fails the test with
"is listed in SUBSTRATE_CRATES but is not a workspace package"; the
clean list passes (23/23).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): prune the dead sanctioned path from the specificity gate

Gate-audit finding (open-and-shut): SANCTIONED_PATHS in
reborn_extension_specificity.rs still exempted
`extension_host/extension_installation_store.rs` — a file deleted by
nearai#6430. No scanned path matches the fragment (verified with rg across
crates/), so the entry exempted nothing; it is also the one exclusion
surface in this gate with no staleness check, which is how it outlived
its file. Full specificity suite green after removal (8/8).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): drop the v1 ironclaw_gateway/static exclusions from the telegram gates

Gate-audit finding (open-and-shut): both cross-tree scans in
telegram_extension_gates.rs still carved out `ironclaw_gateway/static`
— the v1 monolith's embedded UI, whose crate was deleted with the src/
monolith (no crates/*/ironclaw_gateway directory exists). The exclusions
matched nothing; scans now cover the whole tree with no dead carve-outs.
Suite green after removal (12/12).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): make the dto-collapse gate's header describe the gate that exists

Gate-audit finding (open-and-shut doc rot): the module doc still
described the pre-nearai#6447 freeze design — a dangling doc-link to
FROZEN_COLLAPSE_DTOS (renamed RETIRED_COLLAPSE_DTOS in nearai#6447), a
promised delete-without-trimming failure and an empty-allowlist
assertion that do not exist in the file, and a named owner for a
collapse that completed. The mechanism itself is armed and untouched;
the header now describes the permanent zero-gate it became, and records
the two originally-frozen names that deliberately left governance
(CapabilityOutcome via nearai#6299 deletion, CapabilityDispatchRequest blessed
as the canonical port type). Suite green (2/2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): repoint the manifest-reparse allowlist note at the colocated asset

Gate-audit finding (open-and-shut doc rot): the BundledAsset allowlist
entry's justification still cited include_str! of
assets/memory_native/manifest.toml — a path retired when WS2 (nearai#7037)
colocated packages; the live include in memory_native_extension.rs
reaches crates/extensions/packages/memory-native/manifest.toml. Comment
only; the gate's mechanism and counts are untouched. Suite green (2/2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): give the memory-vocabulary gate the partial-tree floor its twin has

Gate-audit finding: reborn_memory_retired_vocabulary.rs had no
MIN_SCANNED_FILES floor, unlike its explicit twin
reborn_retired_taxonomy.rs — so a partially-moved tree (the CHECKLIST
WS0 / nearai#6963 'green while measuring nothing' shape) would scan a
fraction of the files and still report the vocabulary clean. The gate
was in fact born with an already-dead sanctioned path (its own header
records this), so the rot class is not hypothetical for this file.

Adds the same 500-file floor (real count ~4000), asserts it in the main
gate, and pins the premise on a fixture: a 10-file partial tree scans
clean and is rejected by the floor. Suite green (4/4); clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): close the transport gate's nested-use-group fail-open

Gate-audit finding (sabotage-verified): product_symbols_in's braced-group
branch closed at the FIRST '}' (group.find('}')), so a nested group —
use ironclaw_assistant::{m::{X}}; — truncated mid-element and recorded
zero symbols. Probed live before the fix: appending
use ironclaw_assistant::{zzz_audit::{ZzzProbe}}; to webui's lib.rs left
transports_name_only_the_frozen_residue_of_product_symbols GREEN, while
the plain-path spelling of the same import correctly failed. The same
truncation dropped qualified elements inside flat groups
({qualified_module::X} recorded nothing).

The group branch now does a balanced-brace walk, splits elements at
depth-0 commas only, and records a qualified/nested element's leading
path segment — the same key the single-path branch records for
ironclaw_assistant::module::X. Flat-element semantics are byte-for-byte
unchanged, so the frozen 100-row webui inventory is untouched (suite
green 6/6 on the live tree). Regression fixtures added to
import_scanner_reads_symbols_out_of_real_use_shapes; the original
sabotage now fails with the gate's own message (re-verified).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: delete check-e2e-matrix-files.sh — a gate for a workflow that no longer exists

Gate-audit finding (provably inert): the script's default target is
.github/workflows/e2e.yml, deleted when the v1 e2e suites were retired
(git log --diff-filter=D shows the removing commit); no workflow, script,
hook, doc, or guidance file references check-e2e-matrix-files.sh
(verified with rg across the repo including .github and .githooks).
A checker nothing runs, pointed at a file nothing provides, is dead
weight that reads as coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: delete the measured-broken check-boundaries.sh and its guidance references

Gate-audit finding (provably inert, previously measured): crates/AGENTS.md
recorded on 2026-08-05 that the script fails on a clean tree (check 5
false-positives on live test files) and that checks 1/2/3/6 target the
deleted v1 src/ tree, passing vacuously. No workflow or hook runs it; its
only callers were guidance files, two of which claimed it 'enforces'
root-tests feature gating — an enforcement claim the skill-maintainer
rules forbid for a check nothing executes.

Removed the script and every live reference: the crates/AGENTS.md warning
row becomes a tombstone note; the testing skill + exemplar reference drop
the false enforcement parenthetical; the architecture-review skill's
Verify line drops the dead command; deslop-reborn's allowed-tools drops
the permission; .coderabbit.yaml's driver-leak instruction now points at
the live enforcement (reborn_persistence_driver_boundary). Two dated
docs/internal/ plan snapshots keep their historical mentions.

Verified: python3 scripts/ci/check-guidance.py OK (2084 path references)
and its self-test OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(product): stop hardcoding charter sub-owner counts in the family map

Gate-audit finding (stale prose): crates/product/AGENTS.md said
'19-sub-owner reborn_services charter map' — the enforced map has had 20
sub-owners since nearai#7235 added the inspector row (counted from the live
table). Rather than chase the number, drop both inline counts: the
owning maps and their gates are authoritative, and the re-verify
commands are already inline (skill-maintainer rule: no counts without a
regeneration recipe). check-guidance.py OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): correct the scanner-fixture file's name-filter claim

Gate-audit finding (doc rot with a false coverage claim): the header
said naming the FILE reborn_* makes code_style.yml's
'cargo test -p ironclaw_architecture_tests reborn' see it — but that
argument is a test-NAME filter (the measurement is documented in
reborn_contracts_vendor_census.rs), and none of this file's test fns
contains the substring, so that smoke lane runs 0 of them (11 collected
by the full plan). Comment-only; the note now records the real semantics
so file names are not trusted for lane coverage. Suite green (11/11).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): gate & ratchet audit report + proposed preflight gauntlet

The audit the owner asked for after PR nearai#7157 went red six times across
four gates: every architecture-test gate, module charter, CI script, and
committed baseline inventoried with a verdict and evidence; the handful
worth acting on ranked by friction x weakness; the CI-ergonomics analysis
(why failures surface one per ~1h round-trip: no --no-fail-fast anywhere
in CI, cancel-in-progress on push, sequential fast-checks steps —
measured: two broken gates report 1 failure in 18s under the CI shape vs
both in 211s with --no-fail-fast); and the sabotage log for every probe.

scripts/preflight-gates.sh is the concrete pre-push proposal: the
deterministic-gate classes only (script gates ~10s + architecture suite
--no-fail-fast + changed-crate charter tests), covering all four nearai#7157
gate classes locally in one command. Unwired — nothing invokes it.
Validated end-to-end on this branch: exit 0, 'every deterministic gate
green', 402.8s including gate-binary recompiles.

Placement verified: python3 scripts/ci/docs_publication_boundary.py OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(planner): classify preflight-gates.sh and the deleted check-boundaries.sh

The gate audit's own PR hit the planner's fail-closed arm — 'unmapped
test or CI path: scripts/check-boundaries.sh' — exactly the class the
arm exists to force a decision on (and the audit's report documents).
Per the PR_STATIC_CONTROL_PATHS membership rule (no Reborn test lane
exercises either file):

- scripts/preflight-gates.sh — the audit's proposed local pre-push
  gauntlet; referenced by no workflow.
- scripts/check-boundaries.sh — deleted by the audit; the entry lets the
  deletion diff (and any revert) classify instead of failing every
  downstream Reborn lane.

Verified: the planner now produces mode=selected with the
architecture-misc bucket for this branch's diff, and
python3 scripts/ci/test_reborn_pr_test_plan.py is OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): add the fold-tripped asymmetric-tolerance exhibit to the audit

The strongest single exhibit for shortlist item 2, contributed by the
nearai#7157 branch steward after this audit's cutoff and verified against the
gate's code: TOLERANCE = 400 is consulted in exactly one direction (the
banked-slack check, ceiling.saturating_sub(lines) > TOLERANCE); the
growth check is a bare lines > ceiling. With the in-file 'set to
current, not padded' instruction, every ceiling is a hard cap at the
observed count — so one line landing on main in any contracts crate
reds every open branch at its next fold until someone re-captures.

Measured recurrence on nearai#7157: loop_contracts re-captured four times,
~once per fold (14,479 -> 13,850 -> 13,949 -> 13,115 -> 13,181), the
last tripped by main's nearai#7361/nearai#7363 adding 66 lines to
instruction_bundle.rs — nothing the branch wrote. All four deltas were
<= 105 lines: either repair shape in §3.2 (one-line upward tolerance
using the existing constant, or mid-window pinning) would have absorbed
every one with zero red builds. This audit's own sabotage already
proved the jaws (+1 line host_api red / -1 line common red); the fold
history shows the operational cost. The repair stays a recommendation —
adding growth headroom to a ratchet is the owner's call, not this PR's.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): give the contracts size ceiling upward working slack

Owner-directed repair of the audit's sharpest finding (report §3.2): the
gate's TOLERANCE = 400 was consulted in exactly one direction — the
banked-slack check — while the growth check was a bare lines > ceiling.
Combined with 'set to current, not padded' pins, every ceiling was a hard
cap at the exact observed count, so one line landing on main in any
contracts crate redded every open branch at its next fold until someone
re-captured. Measured on nearai#7157: four loop_contracts re-captures, roughly
once per fold, every delta <= 105 lines — the gate generating its own
busywork.

The growth check now allows GROWTH_TOLERANCE = 150 of working slack
above each pin (sized to composition-budget precedent; the reviewed
raises this gate has caught were +1,069 and +1,214 lines, far above it),
and all six ceilings are re-pinned to the counts the test itself
reported with every ceiling at 0 — which also removes the +400 seed
padding on common/loop_contracts/prompt_envelope that contradicted the
capture rule and put those crates one deleted line from the banked jaw.

Sabotage-verified both ways: +1 line in host_api and -1 line in common —
both red before this change — now pass; a +151-line probe still fails
with the effective-ceiling arithmetic in the message. Full
reborn_dependency_boundaries binary green (41/41); clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(budget): re-equalize composition pins to observed — restore the working window

Owner-directed companion to the contracts-ceiling repair (same annoying
class, other mass gate): merged main-side growth since the 2026-08-05
equalization had drifted +101 LOC and +5 Arc<dyn> sites through the
tolerance windows, leaving 49 LOC / 10 sites of live headroom — the next
routine composition PR would have gone red on wiring alone (the gate
audit measured this the same day it was pinned).

Per the TOML's own maintenance instructions: loc_ceiling/loc_observed
40423 -> 40524 and arc_dyn 814 -> 819, measured with the gate's --print,
set to current not padded, dated notes appended (not overwritten), and
the arch-test record (COMPOSITION_ABSOLUTE_SRC_LOC) moved in the same
commit as its file requires. ceiling_bp stays 658 — the WS0 floor is
deliberately not re-set.

Verified: check-composition-budget.sh OK; its 76-case self-test green;
reborn_restructure_baselines green; probe +100 LOC now passes (was red
at 49 headroom), probe +160 LOC still fails.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): record the landed zero-slack repairs in the audit report

The §3.2 repair moved from recommendation to landed at owner direction;
the report's answer, inventory rows, and §7 ledger now say so, with the
counting-rule fix promoted to the top remaining recommendation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* gates: pin the ceiling-window arithmetic; fail preflight discovery closed

Two review-round hardenings (the open CodeRabbit Majors):

- reborn_dependency_boundaries.rs: extract the size-ceiling comparison into
  contracts_ceiling_verdict() and pin its four window edges with a committed
  regression test (contracts_size_ceiling_window_edges_hold) — accept at
  ceiling+GROWTH_TOLERANCE, reject one line past, accept at
  ceiling-TOLERANCE, reject one banked line further, and a zero-measure scan
  reads Banked, never a silent pass. The pre-repair asymmetry (tolerance
  consulted only downward) can no longer return silently. Live-gate behavior
  re-probed unchanged after the rewiring: +1 line to host_api passes, +151
  fails with the same effective-ceiling message.
- preflight-gates.sh: setup and changed-file discovery now fail closed — a
  missing repo root exits 2, and a failed merge-base/diff widens the charter
  run to all five crates instead of silently skipping them (the same
  fallback the missing-base branch already used). A broken setup may cost
  compile time, never a silent skip.

Full boundary binary 42/42 green; clippy clean; preflight-gates.sh
end-to-end green on this tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…signal, builtin description trust, docs (nearai#7361)

* fix(extensions): host-bundled description trust + already-connected install confirmation

Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread
e79a994f, run 251aec0b on ironclaw-qa-testing-libsql):

1. Host-bundled capability descriptions were description_trust=Untrusted,
   so the loop-tier prompt-text denylist strict-scanned compiled-in text
   and silently omitted builtin.extension_register_hosted_mcp from every
   model prompt's capability surface ("browser authorization-code flow"
   matched the "authorization" credential pattern). HostBundled is the
   only source eligible for effective FirstParty/System trust, so its
   repo-authored descriptions now cross the verified-catalog boundary
   like signature/digest-verified registry installs. Untrusted provenance
   (InstalledLocal, UserRegistered, unknown) keeps the strict scan.

2. When install-driven activation passed the credential gate because the
   caller's declared requirements were all satisfied, the response never
   said so — the model got only conditional guidance ("If WebChat shows
   an account connection panel...") and deflected an explicit "connect
   account" request to the web interface even though the account was
   already connected. The install response now appends an explicit
   already-connected confirmation exactly when declared requirements
   were verified present for the calling user.

Regression tests: manager surface test pins VerifiedCatalog trust for all
model-visible lifecycle capabilities through the real host runtime;
instruction-bundle tests pin retain/omit behavior for auth-vocabulary
descriptions by trust; install-path tests pin the confirmation on the
seeded-credential path and its absence for credential-free extensions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): chat can drive the personal half of channel connect

The onboarding and channels pages claimed "asking the agent to connect a
channel doesn't work" and that the agent "may tell you it can't help".
That describes only the operator half (registering app/bot credentials).
The per-user half has shipped since early July: extension_install runs the
same activation credential gate as the Channels card, raises the in-chat
OAuth connection panel when the account is unconnected, and (as of the
sibling fix) confirms when it is already connected.

The self-knowledge protocol makes these pages the model's authority on
IronClaw's own capabilities, so the stale claim scripted the exact
refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an
account that was already connected. Correct both pages to distinguish
the operator step from the chat-drivable personal connect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): align slack and telegram setup notes with the connect contract

The slack page's operator-step note and the telegram troubleshooting
accordion still carried the blanket "asking the agent to connect will not
work" claim the overview/onboarding correction removed — same drift,
different phrasing (review catch on nearai#7361, plus one more instance found
by a broader sweep). Both now state the two-step contract: the operator
half stays in the web interface; after it, chat drives the personal half
(install/activate -> in-chat connection or pairing panel, or an
already-connected confirmation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(golden): recapture surface digests over the description-trust change

The queue run failed golden_payload because the branch predated current
main and its own surface.rs trust fix changes the surface digest. The
recaptured snapshots differ ONLY in the surface sha256 lines (verified
char-by-char) — no prompt text or capability-list changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…ls, delivery heuristics deleted (nearai#7157)

* feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted

Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2
composition inversion, the WS6 crate renames, and the WS7 family moves), with
all 44 review comments dispositioned.

Two-lane delivery model: a run's final reply always lands in its own
conversation (lane 1); reaching any other surface is the model's explicit
`builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target
per call, synchronous through the DeliveryCoordinator, provider-issued
message refs as evidence). Background-run notices fan out to a user-configured
notification-channel set (new record field + read-side legacy migration,
`builtin.notification_channels_set`, first-approve-wins, WebUI multi-select).
The stored delivery heuristics are deleted (route_current, builtin:web_app,
outbound_delivery_target_set, per-trigger delivery_target_id + precedence
chains + four-slot preference fallback), with an idempotent boot migration of
stored trigger targets into explicit prompt steps and the retired vocabulary
pinned in reborn_retired_taxonomy.rs.

Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition,
ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages,
run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→
ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire.
The model-delivery implementation moved extension_host→assistant
(CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host
naming product types; the deferred-slot registration and post-coordinator bind now
live in composition's production assembly, mirroring TriggeredRunDeliveryDriver.

Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary
re-imported from ironclaw_extension_contracts, product-adapter/inbound
vocabulary from host_api/product_contracts, module-charter map's outbound
row renamed to the two-lane vocabulary. Post-branch CI gates adapted in
the same change: skills/ classified in the PR test planner (test-first,
sabotage-verified), panic baseline ratcheted down, nested test fixtures
renamed to the scanner-sanctioned support_tests.rs shape, composition's
inline trigger-migration tests split to tests.rs (mass budget green with
no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21:
the ActivePreferenceTargetCodecs port), loop_contracts ceiling
re-captured down 14479 -> 13850 after the delivery-vocabulary deletion.

Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging
framework): the two-lane guidance moved into the canonical messaging core
prompt (host_api prompts/messaging/send_message.core.md), now naming
builtin__outbound_deliver with the arrive-twice and trigger caveats for
every messaging extension; slack vendor addendum/manifest taken as nearai#6831
shipped them; ceiling-table union (host_api 18570 beside this PR's two
re-captures); retired slack schema embed and deleted preferences
capability stay deleted; golden context-surfacing snapshot regenerated
(one surface-hash line).

Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch +
sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230
beside this PR's re-captures) and main's tracing-target syntax sweep
(target = -> target:, gate-enforced) applied over this PR's kept lines;
deleted delivery-heuristic code stays deleted.

Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep):
zero conflicts; guidance/doc-pointer changes auto-merged over this delta.

Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing
is prompt-owned with a pinned source-surface default (bare "send me" =
the surface you asked from; web app = no delivery step; explicit
destinations override, one delivery step each) — iterated against live
recordings until a real model followed it, with two live-recorded QA
fixtures (bare-webui, multi-channel) plus contracts and replays. The
automations-page panel is retained as the notification-channel selector
(notices only); the conversational notification_channels_set tool writes
the same validated set.

Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow
on main): mark_terminal reports whether the durable write committed and
a confirmed send whose Delivered row failed to commit returns
DeliveredUnconfirmed (refs retained, durably_recorded: false), never a
fabricated Delivered — regression-tested and sabotage-verified. Plus a
CodeRabbit triage batch: correctable coordinator errors stay
model-visible, omitted target_ids no longer clears the set, the success
schema requires evidence, the composition outbound facade is dissolved,
and guidance/contract docs are aligned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: harden channel delivery routines and replay

* fix: preserve automation loading and identity freshness

* fix: close channel delivery review defects

* fix: skip paused routine catch-up slots

* ci: record channel delivery composition budget

* test: align composition baseline with channel delivery

* fix(auth): survive interrupted OAuth callbacks

* fix(auth): keep callback coordination panic-free

* fix(ci): reconcile channel delivery merge seams

* fix(ci): recapture merged contracts ceiling

* fix(delivery): close review findings across delivery, migration, and guidance

Fixes the findings from the multi-agent review of this PR. Every behavioral
fix ships with a regression test that fails before it.

CI (red on this head)
- `standalone_yolo_notification_channels_set_bypasses_approval_gate`
  expected the shared "invalid outbound delivery request" summary for a
  `builtin__notification_channels_set` call. Production deliberately
  specializes that message per operation and pins it with
  `notification_channel_failure_names_the_operation_the_model_can_correct`;
  the assertion was the stale side.

Delivery evidence (kernel + assistant + outbound)
- `AlreadyDelivered` replays reported `delivered: false` "unverified"
  because the ledger row retains no provider refs, inviting the duplicate
  resend the at-most-once claim exists to prevent. Evidence gained
  `already_delivered`; a replay now reads as delivered with an honest
  "not resent" summary. The classification suite had no `AlreadyDelivered`
  case at all, which is why this shipped.
- `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually
  sent something, but `delivered_messages_from_outcome` dropped its refs,
  so gate reply-routes went unrecorded and a live OAuth prompt could never
  be retracted on that path.
- `content` is now rejected when empty: the input schema advertised
  minLength 1 and nothing enforced it, so empty content reached the channel
  as an empty part and returned an opaque provider error.

Background-run notifier (assistant)
- One run legitimately emits several `RunBlocked` notices (re-auth
  stand-in, unserviceable-auth cancellation, run failure), but all three
  derived the same projection ref, and the delivery id hashes it. The
  second notice to a target came back `AlreadyDelivered`, was treated as
  success, and was never sent — a user could be told a routine needed
  re-authorization and never told it then failed. Notices carry a
  discriminator; once-per-run kinds keep their historical id shape, so
  existing delivery identities are unchanged.
- When every catalog lookup failed, the empty result was recorded as
  `NoDefaultConfigured`, reporting a backend outage as the benign "user
  configured nothing" state. It now records `Failed`.

Boot migration (composition)
- The retired `builtin:web_app` target meant "no external delivery". It was
  being rewritten into a delivery step to an id nothing can resolve,
  inverting the stored intent on every later fire. It now clears without
  adding a step.
- One unmigratable row aborted the entire composition boot, with the error
  telling the operator to shorten a prompt through the UI that no longer
  starts. It now pauses its own routine — a paused trigger cannot fire, so
  "never fire unrouted" still holds per record — and boot continues. Only a
  systemic store failure stays boot-fatal. A row deleted during the CAS
  retry ends that record instead of failing boot.
- The CAS retry loop, its bounded exhaustion, and the vanished-row arm had
  no caller-level coverage; adds a delegating repository double that forces
  CAS misses. The prior fail-closed test is rewritten to pin the invariant
  it documented (route survives, record not half-migrated) under the new
  per-record mechanism.

Model-visible messages (composition)
- The targets-list denial said "not permitted to change the outbound
  delivery target" for a read-only call, and the lease denial named the
  retired delivery-target concept on the notification-channel path that is
  its only production caller. Both are now operation-specific and pinned.

WebUI (frontend)
- `setNotificationChannels()` with no argument posted `target_ids: []`,
  turning an omitted argument into a destructive clear-all and defeating
  the backend contract that deliberately rejects an omitted field.
- The notification-channels panel stayed editable after a failed read, so
  toggling one row full-replaced the stored set from an empty baseline and
  silently dropped every channel the user never saw. Editing is now locked
  on a failed read, with a rendered explanation.
- Adds the missing `tools.description.builtin.notification_channels_set`
  key to all 11 locales, plus save-failure coverage for the hook (which was
  correct, but untested) and locale-parity tests.

Guidance
- The new `.claude/rules/tools.md` was ported from a pre-restructure branch:
  it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never
  existed), and its review command grepped three paths removed by WS6/WS7.
  Its `paths:` frontmatter also never matched the product/composition
  callers its rules govern, so the rule never loaded for them.
- `ironclaw_loop_contracts` now records both embedded prompt assets; this
  PR added a second one while the crate's Known-debt entry still said one.
- Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in
  this PR which both bumped) and fixes a pre-rename path in the
  extension-runtime checklist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): unify the DM-rule enforcement point and re-ratchet composition

CI (composition mass budget, red on the previous head): the per-record
migration quarantine pushed composition 6 LOC over its absolute ceiling.
Resolved by the reduction the budget file itself blesses rather than a
raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split
verbatim into `runtime/approval/tests.rs` (the gate excludes test-only
files but counts inline test modules). Composition is now 40,432 LOC,
smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the
arch-test record are re-captured together at the measured value per the
gate's one-directional ratchet rule.

The codec-scan that decodes a binding and enforces "an OAuth authorization
URL only ever lands in a personal DM" existed twice — once in
`TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` —
with both copies commented as "the single enforcement point". They are now
one implementation, shared by the notifier and `builtin.outbound_deliver`,
with a context label so each path keeps its own diagnostic.

That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`)
failed no test in the crate. The vendor codecs pin the predicate in
isolation and the coordinator test pins rejection handling with a double
that decides the verdict itself, so nothing covered the wiring that joins
them. Adds a contract test driving the real resolver through
`DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target
never reaches the vendor adapter. Sabotage-verified: the test fails with
the rule disabled and passes with it restored.

Smaller findings: the notification-channel schema cap now derives from
`ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring
`8`; `triggered_run_delivery`'s module and trait docs described the retired
result-push model this PR deletes; the two new notification strings used a
different brand spelling and dash style from the nine siblings in their own
module; and several new comments navigated by pre-rename paths
(`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a
citation of a test symbol that does not exist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): scope the delivery catalog to the authenticated actor

`builtin.outbound_deliver` resolved its destination catalog under
`ResourceScope.user_id` while performing the send as
`authenticated_actor_user_id`. Those are the same user on a personal thread
and on an automation fire, but they diverge on a shared-route channel
conversation: the scope user is the route's SUBJECT
(`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the
message. Any participant of such a channel could therefore name the
subject's target ids and push bot-identity content into the subject's own
destinations — their personal DM included — from a conversation the subject
may never read.

The catalog now follows the actor, so a caller stays inside their own
connected surfaces on every path and an unfamiliar target simply does not
resolve. Behavior is unchanged wherever owner and actor already agree,
which is every non-shared-route path.

Regression test drives the divergent case through the port (participant
denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a
control proving the owner's own delivery still works. Sabotage-verified:
restoring owner-scoping fails it.

NOT changed here, and flagged for a product decision: the sibling
`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derive their caller from the same
owner-preferring `effective_user_id`, so on a shared route a participant
can still enumerate — and, with the approval gate auto-approved, rewrite —
the subject's notification channels. That helper also scopes approval
gates and capability leases, so flipping its precedence risks breaking
approval raise/resume matching in a path no test covers; it needs its own
change with that coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): stop rewriting the DM-target row on every message

The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert`
for every admitted inbound direct message, and the store unconditionally
wrote a fresh row. After the first message the stored record is already
correct, so the steady state was one durable backend write per DM message,
forever, whose only effect was a new `updated_at` — and each message's
reply-delivery observation was serialized behind it. An unchanged record
now short-circuits; the existing row is loaded here anyway to preserve
`created_at`, so the comparison costs nothing.

Also adds the regression test the `NoDefaultConfigured` -> `Failed`
classification fix landed without: the notifier's `SkipEntry` lookup lane
had no coverage at all (no test ever made a catalog lookup error), so
neither the skip nor the all-failed arm was exercised. The triggered
harness gains an injectable catalog provider for it. Sabotage-verified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(loop): bound the connected-channels line so the runtime slice fits

Confirmed live, not theoretical: a worst-case runtime context renders 4,391
bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that
`instruction_bundle::push_runtime_context` validates the whole slice on —
and exceeding it is a run-ending error on EVERY prompt build for that user,
not a one-off.

This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a
previously-fitting context over. The individual parts are each bounded
(location 200 chars at its producer, locale 35, per-label safe-text
validation), but nothing bounded their SUM, and the connected-channels line
is the one part that grows without limit: up to 20 entries whose names and
presentation hints are only individually capped.

That line now renders as many channels as fit a 1 KiB budget and folds the
rest into the "+N more" counter it already carried, so the fixed guidance
can never be squeezed out by variable content. The worst-case test that
found this stays as the pin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): resolve outbound capabilities as the acting user

`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derived their caller from
`effective_user_id`, which prefers the thread owner over the actor. Those
agree on a direct message and on an automation fire, but diverge on a
shared-route channel conversation, where the owner is the route's
configured subject — the deployment operator by default
(`channel_workflow.rs`) — and the actor is whoever posted.

So any participant of a shared channel could enumerate the operator's
connected destinations and rewrite the operator's notification-channel set,
which is where approval prompts, re-auth prompts and failure notices are
delivered. The caller now follows the acting user, matching the fix already
applied to `builtin.outbound_deliver`.

This deliberately REVERSES a previously pinned preference. Two tests
asserted the owner won when the two differ; that pin predates shared-route
subjects defaulting to the operator, and it contradicts the rule that a run
acts as whoever invoked it. Both are updated to pin the actor, with the
reversal recorded at each site rather than silently relaxed, and the
notification-channel write is now asserted to land under the acting user
with the thread owner's own set left untouched.

INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run`
still follow the owner, because they scope the approval-gate raise and the
capability lease and those must stay matched between raise and resume.
Unifying them belongs with the follow-up that removes shared-route subject
binding entirely so a shared channel runs wholly as its invoker; that needs
approval raise/resume coverage which does not exist yet. A new test pins
the split so the interim state is explicit rather than accidental.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep loop_contracts under its size ceiling and ratchet down

The runtime-context byte-budget fix and its worst-case pin pushed
`ironclaw_loop_contracts` to 14,032 production lines against a 13,949
ceiling. Resolved by the reduction the gate prefers over a raise:
`runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim
into a `runtime_context/tests.rs` sibling, which `production_rust_files`
excludes (an inline test module inside a production file is counted; a
test-only file is not).

The crate now measures 13,115 — 834 lines below the previous ceiling and
smaller than before this review round — so the ceiling is re-captured
downward at the measured value rather than raised, per the gate's
one-directional ratchet. Count read from the gate's own failure message,
not by eye.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): chart the notification-channel handlers in the WebUI charter map

`handlers_module_charter` failed: `get_notification_channels` and
`set_notification_channels` — handlers this PR adds — had no sub-owner row
in `CONTRACT.md`'s enforced charter map, and the row they belong to still
named `get_outbound_preferences`, `set_outbound_preferences` and
`outbound_preferences_activity_id`, all deleted by this PR.

Both halves are fixed together because the gate checks both in one test:
unclaimed items first, then entries naming items that no longer exist. Only
the first had fired, so the stale half was still latent behind it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): re-capture the loop_contracts ceiling after main's merge

Main's nearai#7361/nearai#7363 landed 66 lines in
`ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked
up when folding onto main. That put the crate at 13,181 against the 13,115
ceiling re-captured earlier in this PR.

Not growth from this PR. The gate's upward check is a hard `lines > ceiling`
with no headroom — `TOLERANCE` (400) governs only the downward
ratchet-nudge — so a ceiling captured at the exact observed value reddens
every open branch the moment anyone adds a line to that crate, including
from main. Re-captured at the measured value; count read from the gate's own
failure message.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…mpt-surface PRs (nearai#7371)

* fix(ci): schedule the golden lane for prompt-surface production changes

A production change to the model-visible prompt surface (the capability
surface digest, the instruction bundle, the communication-context
renderer, or a shipped loop-tier prompt asset) ran only crate buckets on
the PR lane, so stale golden_payload snapshots surfaced first as a
merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface
digest; the PR lane never ran the golden bucket). Add a curated
prompt-surface owner table that ADDITIONALLY schedules the golden
integration lane without consuming the path's normal package
classification. Self-tested per entry plus a negative control pinning
that ordinary production changes keep the narrow plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): recapture the extension_host coverage floor with measured wobble tolerance

The 2026-08-04 entry was an adjustment whose own text ordered the next
real recapture. Since then code left the crate (denominator 24529 ->
24415) while the ratio ROSE to 88.12%, and the same commit df90072
measured >=21560 covered lines in the merge-queue lane but 21515 twice
on the push lane — a >=45-line same-commit spread over a 20-line
tolerance, redding main on noise (run 31208592262). Recapture both
fields from that run's own gate output and size tolerance_lines to the
measured cross-lane wobble.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): cover the host-managed prompt composer and derive per-entry tests

Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives
the InstructionBundleBuilder, model.rs shapes the pinned request) was
missing from the prompt-surface table — the exact gap class the mapping
exists to close. And the self-test enumerated entries by hand, so a new
entry could ship untested. Add the host_managed_ports prefix and derive
the positive cases from the tables themselves; an entry whose crate the
fixture lacks now fails the suite explicitly instead of skipping.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…signal, builtin description trust, docs (nearai#7361)

* fix(extensions): host-bundled description trust + already-connected install confirmation

Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread
e79a994f, run 251aec0b on ironclaw-qa-testing-libsql):

1. Host-bundled capability descriptions were description_trust=Untrusted,
   so the loop-tier prompt-text denylist strict-scanned compiled-in text
   and silently omitted builtin.extension_register_hosted_mcp from every
   model prompt's capability surface ("browser authorization-code flow"
   matched the "authorization" credential pattern). HostBundled is the
   only source eligible for effective FirstParty/System trust, so its
   repo-authored descriptions now cross the verified-catalog boundary
   like signature/digest-verified registry installs. Untrusted provenance
   (InstalledLocal, UserRegistered, unknown) keeps the strict scan.

2. When install-driven activation passed the credential gate because the
   caller's declared requirements were all satisfied, the response never
   said so — the model got only conditional guidance ("If WebChat shows
   an account connection panel...") and deflected an explicit "connect
   account" request to the web interface even though the account was
   already connected. The install response now appends an explicit
   already-connected confirmation exactly when declared requirements
   were verified present for the calling user.

Regression tests: manager surface test pins VerifiedCatalog trust for all
model-visible lifecycle capabilities through the real host runtime;
instruction-bundle tests pin retain/omit behavior for auth-vocabulary
descriptions by trust; install-path tests pin the confirmation on the
seeded-credential path and its absence for credential-free extensions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): chat can drive the personal half of channel connect

The onboarding and channels pages claimed "asking the agent to connect a
channel doesn't work" and that the agent "may tell you it can't help".
That describes only the operator half (registering app/bot credentials).
The per-user half has shipped since early July: extension_install runs the
same activation credential gate as the Channels card, raises the in-chat
OAuth connection panel when the account is unconnected, and (as of the
sibling fix) confirms when it is already connected.

The self-knowledge protocol makes these pages the model's authority on
IronClaw's own capabilities, so the stale claim scripted the exact
refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an
account that was already connected. Correct both pages to distinguish
the operator step from the chat-drivable personal connect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): align slack and telegram setup notes with the connect contract

The slack page's operator-step note and the telegram troubleshooting
accordion still carried the blanket "asking the agent to connect will not
work" claim the overview/onboarding correction removed — same drift,
different phrasing (review catch on nearai#7361, plus one more instance found
by a broader sweep). Both now state the two-step contract: the operator
half stays in the web interface; after it, chat drives the personal half
(install/activate -> in-chat connection or pairing panel, or an
already-connected confirmation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(golden): recapture surface digests over the description-trust change

The queue run failed golden_payload because the branch predated current
main and its own surface.rs trust fix changes the surface digest. The
recaptured snapshots differ ONLY in the surface sha256 lines (verified
char-by-char) — no prompt text or capability-list changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…ls, delivery heuristics deleted (nearai#7157)

* feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted

Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2
composition inversion, the WS6 crate renames, and the WS7 family moves), with
all 44 review comments dispositioned.

Two-lane delivery model: a run's final reply always lands in its own
conversation (lane 1); reaching any other surface is the model's explicit
`builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target
per call, synchronous through the DeliveryCoordinator, provider-issued
message refs as evidence). Background-run notices fan out to a user-configured
notification-channel set (new record field + read-side legacy migration,
`builtin.notification_channels_set`, first-approve-wins, WebUI multi-select).
The stored delivery heuristics are deleted (route_current, builtin:web_app,
outbound_delivery_target_set, per-trigger delivery_target_id + precedence
chains + four-slot preference fallback), with an idempotent boot migration of
stored trigger targets into explicit prompt steps and the retired vocabulary
pinned in reborn_retired_taxonomy.rs.

Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition,
ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages,
run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→
ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire.
The model-delivery implementation moved extension_host→assistant
(CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host
naming product types; the deferred-slot registration and post-coordinator bind now
live in composition's production assembly, mirroring TriggeredRunDeliveryDriver.

Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary
re-imported from ironclaw_extension_contracts, product-adapter/inbound
vocabulary from host_api/product_contracts, module-charter map's outbound
row renamed to the two-lane vocabulary. Post-branch CI gates adapted in
the same change: skills/ classified in the PR test planner (test-first,
sabotage-verified), panic baseline ratcheted down, nested test fixtures
renamed to the scanner-sanctioned support_tests.rs shape, composition's
inline trigger-migration tests split to tests.rs (mass budget green with
no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21:
the ActivePreferenceTargetCodecs port), loop_contracts ceiling
re-captured down 14479 -> 13850 after the delivery-vocabulary deletion.

Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging
framework): the two-lane guidance moved into the canonical messaging core
prompt (host_api prompts/messaging/send_message.core.md), now naming
builtin__outbound_deliver with the arrive-twice and trigger caveats for
every messaging extension; slack vendor addendum/manifest taken as nearai#6831
shipped them; ceiling-table union (host_api 18570 beside this PR's two
re-captures); retired slack schema embed and deleted preferences
capability stay deleted; golden context-surfacing snapshot regenerated
(one surface-hash line).

Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch +
sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230
beside this PR's re-captures) and main's tracing-target syntax sweep
(target = -> target:, gate-enforced) applied over this PR's kept lines;
deleted delivery-heuristic code stays deleted.

Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep):
zero conflicts; guidance/doc-pointer changes auto-merged over this delta.

Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing
is prompt-owned with a pinned source-surface default (bare "send me" =
the surface you asked from; web app = no delivery step; explicit
destinations override, one delivery step each) — iterated against live
recordings until a real model followed it, with two live-recorded QA
fixtures (bare-webui, multi-channel) plus contracts and replays. The
automations-page panel is retained as the notification-channel selector
(notices only); the conversational notification_channels_set tool writes
the same validated set.

Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow
on main): mark_terminal reports whether the durable write committed and
a confirmed send whose Delivered row failed to commit returns
DeliveredUnconfirmed (refs retained, durably_recorded: false), never a
fabricated Delivered — regression-tested and sabotage-verified. Plus a
CodeRabbit triage batch: correctable coordinator errors stay
model-visible, omitted target_ids no longer clears the set, the success
schema requires evidence, the composition outbound facade is dissolved,
and guidance/contract docs are aligned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: harden channel delivery routines and replay

* fix: preserve automation loading and identity freshness

* fix: close channel delivery review defects

* fix: skip paused routine catch-up slots

* ci: record channel delivery composition budget

* test: align composition baseline with channel delivery

* fix(auth): survive interrupted OAuth callbacks

* fix(auth): keep callback coordination panic-free

* fix(ci): reconcile channel delivery merge seams

* fix(ci): recapture merged contracts ceiling

* fix(delivery): close review findings across delivery, migration, and guidance

Fixes the findings from the multi-agent review of this PR. Every behavioral
fix ships with a regression test that fails before it.

CI (red on this head)
- `standalone_yolo_notification_channels_set_bypasses_approval_gate`
  expected the shared "invalid outbound delivery request" summary for a
  `builtin__notification_channels_set` call. Production deliberately
  specializes that message per operation and pins it with
  `notification_channel_failure_names_the_operation_the_model_can_correct`;
  the assertion was the stale side.

Delivery evidence (kernel + assistant + outbound)
- `AlreadyDelivered` replays reported `delivered: false` "unverified"
  because the ledger row retains no provider refs, inviting the duplicate
  resend the at-most-once claim exists to prevent. Evidence gained
  `already_delivered`; a replay now reads as delivered with an honest
  "not resent" summary. The classification suite had no `AlreadyDelivered`
  case at all, which is why this shipped.
- `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually
  sent something, but `delivered_messages_from_outcome` dropped its refs,
  so gate reply-routes went unrecorded and a live OAuth prompt could never
  be retracted on that path.
- `content` is now rejected when empty: the input schema advertised
  minLength 1 and nothing enforced it, so empty content reached the channel
  as an empty part and returned an opaque provider error.

Background-run notifier (assistant)
- One run legitimately emits several `RunBlocked` notices (re-auth
  stand-in, unserviceable-auth cancellation, run failure), but all three
  derived the same projection ref, and the delivery id hashes it. The
  second notice to a target came back `AlreadyDelivered`, was treated as
  success, and was never sent — a user could be told a routine needed
  re-authorization and never told it then failed. Notices carry a
  discriminator; once-per-run kinds keep their historical id shape, so
  existing delivery identities are unchanged.
- When every catalog lookup failed, the empty result was recorded as
  `NoDefaultConfigured`, reporting a backend outage as the benign "user
  configured nothing" state. It now records `Failed`.

Boot migration (composition)
- The retired `builtin:web_app` target meant "no external delivery". It was
  being rewritten into a delivery step to an id nothing can resolve,
  inverting the stored intent on every later fire. It now clears without
  adding a step.
- One unmigratable row aborted the entire composition boot, with the error
  telling the operator to shorten a prompt through the UI that no longer
  starts. It now pauses its own routine — a paused trigger cannot fire, so
  "never fire unrouted" still holds per record — and boot continues. Only a
  systemic store failure stays boot-fatal. A row deleted during the CAS
  retry ends that record instead of failing boot.
- The CAS retry loop, its bounded exhaustion, and the vanished-row arm had
  no caller-level coverage; adds a delegating repository double that forces
  CAS misses. The prior fail-closed test is rewritten to pin the invariant
  it documented (route survives, record not half-migrated) under the new
  per-record mechanism.

Model-visible messages (composition)
- The targets-list denial said "not permitted to change the outbound
  delivery target" for a read-only call, and the lease denial named the
  retired delivery-target concept on the notification-channel path that is
  its only production caller. Both are now operation-specific and pinned.

WebUI (frontend)
- `setNotificationChannels()` with no argument posted `target_ids: []`,
  turning an omitted argument into a destructive clear-all and defeating
  the backend contract that deliberately rejects an omitted field.
- The notification-channels panel stayed editable after a failed read, so
  toggling one row full-replaced the stored set from an empty baseline and
  silently dropped every channel the user never saw. Editing is now locked
  on a failed read, with a rendered explanation.
- Adds the missing `tools.description.builtin.notification_channels_set`
  key to all 11 locales, plus save-failure coverage for the hook (which was
  correct, but untested) and locale-parity tests.

Guidance
- The new `.claude/rules/tools.md` was ported from a pre-restructure branch:
  it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never
  existed), and its review command grepped three paths removed by WS6/WS7.
  Its `paths:` frontmatter also never matched the product/composition
  callers its rules govern, so the rule never loaded for them.
- `ironclaw_loop_contracts` now records both embedded prompt assets; this
  PR added a second one while the crate's Known-debt entry still said one.
- Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in
  this PR which both bumped) and fixes a pre-rename path in the
  extension-runtime checklist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): unify the DM-rule enforcement point and re-ratchet composition

CI (composition mass budget, red on the previous head): the per-record
migration quarantine pushed composition 6 LOC over its absolute ceiling.
Resolved by the reduction the budget file itself blesses rather than a
raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split
verbatim into `runtime/approval/tests.rs` (the gate excludes test-only
files but counts inline test modules). Composition is now 40,432 LOC,
smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the
arch-test record are re-captured together at the measured value per the
gate's one-directional ratchet rule.

The codec-scan that decodes a binding and enforces "an OAuth authorization
URL only ever lands in a personal DM" existed twice — once in
`TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` —
with both copies commented as "the single enforcement point". They are now
one implementation, shared by the notifier and `builtin.outbound_deliver`,
with a context label so each path keeps its own diagnostic.

That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`)
failed no test in the crate. The vendor codecs pin the predicate in
isolation and the coordinator test pins rejection handling with a double
that decides the verdict itself, so nothing covered the wiring that joins
them. Adds a contract test driving the real resolver through
`DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target
never reaches the vendor adapter. Sabotage-verified: the test fails with
the rule disabled and passes with it restored.

Smaller findings: the notification-channel schema cap now derives from
`ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring
`8`; `triggered_run_delivery`'s module and trait docs described the retired
result-push model this PR deletes; the two new notification strings used a
different brand spelling and dash style from the nine siblings in their own
module; and several new comments navigated by pre-rename paths
(`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a
citation of a test symbol that does not exist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): scope the delivery catalog to the authenticated actor

`builtin.outbound_deliver` resolved its destination catalog under
`ResourceScope.user_id` while performing the send as
`authenticated_actor_user_id`. Those are the same user on a personal thread
and on an automation fire, but they diverge on a shared-route channel
conversation: the scope user is the route's SUBJECT
(`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the
message. Any participant of such a channel could therefore name the
subject's target ids and push bot-identity content into the subject's own
destinations — their personal DM included — from a conversation the subject
may never read.

The catalog now follows the actor, so a caller stays inside their own
connected surfaces on every path and an unfamiliar target simply does not
resolve. Behavior is unchanged wherever owner and actor already agree,
which is every non-shared-route path.

Regression test drives the divergent case through the port (participant
denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a
control proving the owner's own delivery still works. Sabotage-verified:
restoring owner-scoping fails it.

NOT changed here, and flagged for a product decision: the sibling
`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derive their caller from the same
owner-preferring `effective_user_id`, so on a shared route a participant
can still enumerate — and, with the approval gate auto-approved, rewrite —
the subject's notification channels. That helper also scopes approval
gates and capability leases, so flipping its precedence risks breaking
approval raise/resume matching in a path no test covers; it needs its own
change with that coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): stop rewriting the DM-target row on every message

The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert`
for every admitted inbound direct message, and the store unconditionally
wrote a fresh row. After the first message the stored record is already
correct, so the steady state was one durable backend write per DM message,
forever, whose only effect was a new `updated_at` — and each message's
reply-delivery observation was serialized behind it. An unchanged record
now short-circuits; the existing row is loaded here anyway to preserve
`created_at`, so the comparison costs nothing.

Also adds the regression test the `NoDefaultConfigured` -> `Failed`
classification fix landed without: the notifier's `SkipEntry` lookup lane
had no coverage at all (no test ever made a catalog lookup error), so
neither the skip nor the all-failed arm was exercised. The triggered
harness gains an injectable catalog provider for it. Sabotage-verified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(loop): bound the connected-channels line so the runtime slice fits

Confirmed live, not theoretical: a worst-case runtime context renders 4,391
bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that
`instruction_bundle::push_runtime_context` validates the whole slice on —
and exceeding it is a run-ending error on EVERY prompt build for that user,
not a one-off.

This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a
previously-fitting context over. The individual parts are each bounded
(location 200 chars at its producer, locale 35, per-label safe-text
validation), but nothing bounded their SUM, and the connected-channels line
is the one part that grows without limit: up to 20 entries whose names and
presentation hints are only individually capped.

That line now renders as many channels as fit a 1 KiB budget and folds the
rest into the "+N more" counter it already carried, so the fixed guidance
can never be squeezed out by variable content. The worst-case test that
found this stays as the pin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): resolve outbound capabilities as the acting user

`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derived their caller from
`effective_user_id`, which prefers the thread owner over the actor. Those
agree on a direct message and on an automation fire, but diverge on a
shared-route channel conversation, where the owner is the route's
configured subject — the deployment operator by default
(`channel_workflow.rs`) — and the actor is whoever posted.

So any participant of a shared channel could enumerate the operator's
connected destinations and rewrite the operator's notification-channel set,
which is where approval prompts, re-auth prompts and failure notices are
delivered. The caller now follows the acting user, matching the fix already
applied to `builtin.outbound_deliver`.

This deliberately REVERSES a previously pinned preference. Two tests
asserted the owner won when the two differ; that pin predates shared-route
subjects defaulting to the operator, and it contradicts the rule that a run
acts as whoever invoked it. Both are updated to pin the actor, with the
reversal recorded at each site rather than silently relaxed, and the
notification-channel write is now asserted to land under the acting user
with the thread owner's own set left untouched.

INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run`
still follow the owner, because they scope the approval-gate raise and the
capability lease and those must stay matched between raise and resume.
Unifying them belongs with the follow-up that removes shared-route subject
binding entirely so a shared channel runs wholly as its invoker; that needs
approval raise/resume coverage which does not exist yet. A new test pins
the split so the interim state is explicit rather than accidental.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep loop_contracts under its size ceiling and ratchet down

The runtime-context byte-budget fix and its worst-case pin pushed
`ironclaw_loop_contracts` to 14,032 production lines against a 13,949
ceiling. Resolved by the reduction the gate prefers over a raise:
`runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim
into a `runtime_context/tests.rs` sibling, which `production_rust_files`
excludes (an inline test module inside a production file is counted; a
test-only file is not).

The crate now measures 13,115 — 834 lines below the previous ceiling and
smaller than before this review round — so the ceiling is re-captured
downward at the measured value rather than raised, per the gate's
one-directional ratchet. Count read from the gate's own failure message,
not by eye.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): chart the notification-channel handlers in the WebUI charter map

`handlers_module_charter` failed: `get_notification_channels` and
`set_notification_channels` — handlers this PR adds — had no sub-owner row
in `CONTRACT.md`'s enforced charter map, and the row they belong to still
named `get_outbound_preferences`, `set_outbound_preferences` and
`outbound_preferences_activity_id`, all deleted by this PR.

Both halves are fixed together because the gate checks both in one test:
unclaimed items first, then entries naming items that no longer exist. Only
the first had fired, so the stale half was still latent behind it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): re-capture the loop_contracts ceiling after main's merge

Main's nearai#7361/nearai#7363 landed 66 lines in
`ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked
up when folding onto main. That put the crate at 13,181 against the 13,115
ceiling re-captured earlier in this PR.

Not growth from this PR. The gate's upward check is a hard `lines > ceiling`
with no headroom — `TOLERANCE` (400) governs only the downward
ratchet-nudge — so a ceiling captured at the exact observed value reddens
every open branch the moment anyone adds a line to that crate, including
from main. Re-captured at the measured value; count read from the gate's own
failure message.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…mpt-surface PRs (nearai#7371)

* fix(ci): schedule the golden lane for prompt-surface production changes

A production change to the model-visible prompt surface (the capability
surface digest, the instruction bundle, the communication-context
renderer, or a shipped loop-tier prompt asset) ran only crate buckets on
the PR lane, so stale golden_payload snapshots surfaced first as a
merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface
digest; the PR lane never ran the golden bucket). Add a curated
prompt-surface owner table that ADDITIONALLY schedules the golden
integration lane without consuming the path's normal package
classification. Self-tested per entry plus a negative control pinning
that ordinary production changes keep the narrow plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): recapture the extension_host coverage floor with measured wobble tolerance

The 2026-08-04 entry was an adjustment whose own text ordered the next
real recapture. Since then code left the crate (denominator 24529 ->
24415) while the ratio ROSE to 88.12%, and the same commit df90072
measured >=21560 covered lines in the merge-queue lane but 21515 twice
on the push lane — a >=45-line same-commit spread over a 20-line
tolerance, redding main on noise (run 31208592262). Recapture both
fields from that run's own gate output and size tolerance_lines to the
measured cross-lane wobble.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): cover the host-managed prompt composer and derive per-entry tests

Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives
the InstructionBundleBuilder, model.rs shapes the pinned request) was
missing from the prompt-surface table — the exact gap class the mapping
exists to close. And the self-test enumerated entries by hand, so a new
entry could ship untested. Add the host_managed_ports prefix and derive
the positive cases from the tables themselves; an entry whose crate the
fixture lacks now fails the suite explicitly instead of skipping.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…signal, builtin description trust, docs (nearai#7361)

* fix(extensions): host-bundled description trust + already-connected install confirmation

Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread
e79a994f, run 251aec0b on ironclaw-qa-testing-libsql):

1. Host-bundled capability descriptions were description_trust=Untrusted,
   so the loop-tier prompt-text denylist strict-scanned compiled-in text
   and silently omitted builtin.extension_register_hosted_mcp from every
   model prompt's capability surface ("browser authorization-code flow"
   matched the "authorization" credential pattern). HostBundled is the
   only source eligible for effective FirstParty/System trust, so its
   repo-authored descriptions now cross the verified-catalog boundary
   like signature/digest-verified registry installs. Untrusted provenance
   (InstalledLocal, UserRegistered, unknown) keeps the strict scan.

2. When install-driven activation passed the credential gate because the
   caller's declared requirements were all satisfied, the response never
   said so — the model got only conditional guidance ("If WebChat shows
   an account connection panel...") and deflected an explicit "connect
   account" request to the web interface even though the account was
   already connected. The install response now appends an explicit
   already-connected confirmation exactly when declared requirements
   were verified present for the calling user.

Regression tests: manager surface test pins VerifiedCatalog trust for all
model-visible lifecycle capabilities through the real host runtime;
instruction-bundle tests pin retain/omit behavior for auth-vocabulary
descriptions by trust; install-path tests pin the confirmation on the
seeded-credential path and its absence for credential-free extensions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): chat can drive the personal half of channel connect

The onboarding and channels pages claimed "asking the agent to connect a
channel doesn't work" and that the agent "may tell you it can't help".
That describes only the operator half (registering app/bot credentials).
The per-user half has shipped since early July: extension_install runs the
same activation credential gate as the Channels card, raises the in-chat
OAuth connection panel when the account is unconnected, and (as of the
sibling fix) confirms when it is already connected.

The self-knowledge protocol makes these pages the model's authority on
IronClaw's own capabilities, so the stale claim scripted the exact
refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an
account that was already connected. Correct both pages to distinguish
the operator step from the chat-drivable personal connect.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channels): align slack and telegram setup notes with the connect contract

The slack page's operator-step note and the telegram troubleshooting
accordion still carried the blanket "asking the agent to connect will not
work" claim the overview/onboarding correction removed — same drift,
different phrasing (review catch on nearai#7361, plus one more instance found
by a broader sweep). Both now state the two-step contract: the operator
half stays in the web interface; after it, chat drives the personal half
(install/activate -> in-chat connection or pairing panel, or an
already-connected confirmation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(golden): recapture surface digests over the description-trust change

The queue run failed golden_payload because the branch predated current
main and its own surface.rs trust fix changes the surface digest. The
recaptured snapshots differ ONLY in the surface sha256 lines (verified
char-by-char) — no prompt text or capability-list changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…ls, delivery heuristics deleted (nearai#7157)

* feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted

Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2
composition inversion, the WS6 crate renames, and the WS7 family moves), with
all 44 review comments dispositioned.

Two-lane delivery model: a run's final reply always lands in its own
conversation (lane 1); reaching any other surface is the model's explicit
`builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target
per call, synchronous through the DeliveryCoordinator, provider-issued
message refs as evidence). Background-run notices fan out to a user-configured
notification-channel set (new record field + read-side legacy migration,
`builtin.notification_channels_set`, first-approve-wins, WebUI multi-select).
The stored delivery heuristics are deleted (route_current, builtin:web_app,
outbound_delivery_target_set, per-trigger delivery_target_id + precedence
chains + four-slot preference fallback), with an idempotent boot migration of
stored trigger targets into explicit prompt steps and the retired vocabulary
pinned in reborn_retired_taxonomy.rs.

Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition,
ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages,
run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→
ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire.
The model-delivery implementation moved extension_host→assistant
(CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host
naming product types; the deferred-slot registration and post-coordinator bind now
live in composition's production assembly, mirroring TriggeredRunDeliveryDriver.

Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary
re-imported from ironclaw_extension_contracts, product-adapter/inbound
vocabulary from host_api/product_contracts, module-charter map's outbound
row renamed to the two-lane vocabulary. Post-branch CI gates adapted in
the same change: skills/ classified in the PR test planner (test-first,
sabotage-verified), panic baseline ratcheted down, nested test fixtures
renamed to the scanner-sanctioned support_tests.rs shape, composition's
inline trigger-migration tests split to tests.rs (mass budget green with
no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21:
the ActivePreferenceTargetCodecs port), loop_contracts ceiling
re-captured down 14479 -> 13850 after the delivery-vocabulary deletion.

Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging
framework): the two-lane guidance moved into the canonical messaging core
prompt (host_api prompts/messaging/send_message.core.md), now naming
builtin__outbound_deliver with the arrive-twice and trigger caveats for
every messaging extension; slack vendor addendum/manifest taken as nearai#6831
shipped them; ceiling-table union (host_api 18570 beside this PR's two
re-captures); retired slack schema embed and deleted preferences
capability stay deleted; golden context-surfacing snapshot regenerated
(one surface-hash line).

Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch +
sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230
beside this PR's re-captures) and main's tracing-target syntax sweep
(target = -> target:, gate-enforced) applied over this PR's kept lines;
deleted delivery-heuristic code stays deleted.

Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep):
zero conflicts; guidance/doc-pointer changes auto-merged over this delta.

Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing
is prompt-owned with a pinned source-surface default (bare "send me" =
the surface you asked from; web app = no delivery step; explicit
destinations override, one delivery step each) — iterated against live
recordings until a real model followed it, with two live-recorded QA
fixtures (bare-webui, multi-channel) plus contracts and replays. The
automations-page panel is retained as the notification-channel selector
(notices only); the conversational notification_channels_set tool writes
the same validated set.

Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow
on main): mark_terminal reports whether the durable write committed and
a confirmed send whose Delivered row failed to commit returns
DeliveredUnconfirmed (refs retained, durably_recorded: false), never a
fabricated Delivered — regression-tested and sabotage-verified. Plus a
CodeRabbit triage batch: correctable coordinator errors stay
model-visible, omitted target_ids no longer clears the set, the success
schema requires evidence, the composition outbound facade is dissolved,
and guidance/contract docs are aligned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: harden channel delivery routines and replay

* fix: preserve automation loading and identity freshness

* fix: close channel delivery review defects

* fix: skip paused routine catch-up slots

* ci: record channel delivery composition budget

* test: align composition baseline with channel delivery

* fix(auth): survive interrupted OAuth callbacks

* fix(auth): keep callback coordination panic-free

* fix(ci): reconcile channel delivery merge seams

* fix(ci): recapture merged contracts ceiling

* fix(delivery): close review findings across delivery, migration, and guidance

Fixes the findings from the multi-agent review of this PR. Every behavioral
fix ships with a regression test that fails before it.

CI (red on this head)
- `standalone_yolo_notification_channels_set_bypasses_approval_gate`
  expected the shared "invalid outbound delivery request" summary for a
  `builtin__notification_channels_set` call. Production deliberately
  specializes that message per operation and pins it with
  `notification_channel_failure_names_the_operation_the_model_can_correct`;
  the assertion was the stale side.

Delivery evidence (kernel + assistant + outbound)
- `AlreadyDelivered` replays reported `delivered: false` "unverified"
  because the ledger row retains no provider refs, inviting the duplicate
  resend the at-most-once claim exists to prevent. Evidence gained
  `already_delivered`; a replay now reads as delivered with an honest
  "not resent" summary. The classification suite had no `AlreadyDelivered`
  case at all, which is why this shipped.
- `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually
  sent something, but `delivered_messages_from_outcome` dropped its refs,
  so gate reply-routes went unrecorded and a live OAuth prompt could never
  be retracted on that path.
- `content` is now rejected when empty: the input schema advertised
  minLength 1 and nothing enforced it, so empty content reached the channel
  as an empty part and returned an opaque provider error.

Background-run notifier (assistant)
- One run legitimately emits several `RunBlocked` notices (re-auth
  stand-in, unserviceable-auth cancellation, run failure), but all three
  derived the same projection ref, and the delivery id hashes it. The
  second notice to a target came back `AlreadyDelivered`, was treated as
  success, and was never sent — a user could be told a routine needed
  re-authorization and never told it then failed. Notices carry a
  discriminator; once-per-run kinds keep their historical id shape, so
  existing delivery identities are unchanged.
- When every catalog lookup failed, the empty result was recorded as
  `NoDefaultConfigured`, reporting a backend outage as the benign "user
  configured nothing" state. It now records `Failed`.

Boot migration (composition)
- The retired `builtin:web_app` target meant "no external delivery". It was
  being rewritten into a delivery step to an id nothing can resolve,
  inverting the stored intent on every later fire. It now clears without
  adding a step.
- One unmigratable row aborted the entire composition boot, with the error
  telling the operator to shorten a prompt through the UI that no longer
  starts. It now pauses its own routine — a paused trigger cannot fire, so
  "never fire unrouted" still holds per record — and boot continues. Only a
  systemic store failure stays boot-fatal. A row deleted during the CAS
  retry ends that record instead of failing boot.
- The CAS retry loop, its bounded exhaustion, and the vanished-row arm had
  no caller-level coverage; adds a delegating repository double that forces
  CAS misses. The prior fail-closed test is rewritten to pin the invariant
  it documented (route survives, record not half-migrated) under the new
  per-record mechanism.

Model-visible messages (composition)
- The targets-list denial said "not permitted to change the outbound
  delivery target" for a read-only call, and the lease denial named the
  retired delivery-target concept on the notification-channel path that is
  its only production caller. Both are now operation-specific and pinned.

WebUI (frontend)
- `setNotificationChannels()` with no argument posted `target_ids: []`,
  turning an omitted argument into a destructive clear-all and defeating
  the backend contract that deliberately rejects an omitted field.
- The notification-channels panel stayed editable after a failed read, so
  toggling one row full-replaced the stored set from an empty baseline and
  silently dropped every channel the user never saw. Editing is now locked
  on a failed read, with a rendered explanation.
- Adds the missing `tools.description.builtin.notification_channels_set`
  key to all 11 locales, plus save-failure coverage for the hook (which was
  correct, but untested) and locale-parity tests.

Guidance
- The new `.claude/rules/tools.md` was ported from a pre-restructure branch:
  it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never
  existed), and its review command grepped three paths removed by WS6/WS7.
  Its `paths:` frontmatter also never matched the product/composition
  callers its rules govern, so the rule never loaded for them.
- `ironclaw_loop_contracts` now records both embedded prompt assets; this
  PR added a second one while the crate's Known-debt entry still said one.
- Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in
  this PR which both bumped) and fixes a pre-rename path in the
  extension-runtime checklist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): unify the DM-rule enforcement point and re-ratchet composition

CI (composition mass budget, red on the previous head): the per-record
migration quarantine pushed composition 6 LOC over its absolute ceiling.
Resolved by the reduction the budget file itself blesses rather than a
raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split
verbatim into `runtime/approval/tests.rs` (the gate excludes test-only
files but counts inline test modules). Composition is now 40,432 LOC,
smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the
arch-test record are re-captured together at the measured value per the
gate's one-directional ratchet rule.

The codec-scan that decodes a binding and enforces "an OAuth authorization
URL only ever lands in a personal DM" existed twice — once in
`TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` —
with both copies commented as "the single enforcement point". They are now
one implementation, shared by the notifier and `builtin.outbound_deliver`,
with a context label so each path keeps its own diagnostic.

That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`)
failed no test in the crate. The vendor codecs pin the predicate in
isolation and the coordinator test pins rejection handling with a double
that decides the verdict itself, so nothing covered the wiring that joins
them. Adds a contract test driving the real resolver through
`DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target
never reaches the vendor adapter. Sabotage-verified: the test fails with
the rule disabled and passes with it restored.

Smaller findings: the notification-channel schema cap now derives from
`ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring
`8`; `triggered_run_delivery`'s module and trait docs described the retired
result-push model this PR deletes; the two new notification strings used a
different brand spelling and dash style from the nine siblings in their own
module; and several new comments navigated by pre-rename paths
(`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a
citation of a test symbol that does not exist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): scope the delivery catalog to the authenticated actor

`builtin.outbound_deliver` resolved its destination catalog under
`ResourceScope.user_id` while performing the send as
`authenticated_actor_user_id`. Those are the same user on a personal thread
and on an automation fire, but they diverge on a shared-route channel
conversation: the scope user is the route's SUBJECT
(`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the
message. Any participant of such a channel could therefore name the
subject's target ids and push bot-identity content into the subject's own
destinations — their personal DM included — from a conversation the subject
may never read.

The catalog now follows the actor, so a caller stays inside their own
connected surfaces on every path and an unfamiliar target simply does not
resolve. Behavior is unchanged wherever owner and actor already agree,
which is every non-shared-route path.

Regression test drives the divergent case through the port (participant
denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a
control proving the owner's own delivery still works. Sabotage-verified:
restoring owner-scoping fails it.

NOT changed here, and flagged for a product decision: the sibling
`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derive their caller from the same
owner-preferring `effective_user_id`, so on a shared route a participant
can still enumerate — and, with the approval gate auto-approved, rewrite —
the subject's notification channels. That helper also scopes approval
gates and capability leases, so flipping its precedence risks breaking
approval raise/resume matching in a path no test covers; it needs its own
change with that coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(delivery): stop rewriting the DM-target row on every message

The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert`
for every admitted inbound direct message, and the store unconditionally
wrote a fresh row. After the first message the stored record is already
correct, so the steady state was one durable backend write per DM message,
forever, whose only effect was a new `updated_at` — and each message's
reply-delivery observation was serialized behind it. An unchanged record
now short-circuits; the existing row is loaded here anyway to preserve
`created_at`, so the comparison costs nothing.

Also adds the regression test the `NoDefaultConfigured` -> `Failed`
classification fix landed without: the notifier's `SkipEntry` lookup lane
had no coverage at all (no test ever made a catalog lookup error), so
neither the skip nor the all-failed arm was exercised. The triggered
harness gains an injectable catalog provider for it. Sabotage-verified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(loop): bound the connected-channels line so the runtime slice fits

Confirmed live, not theoretical: a worst-case runtime context renders 4,391
bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that
`instruction_bundle::push_runtime_context` validates the whole slice on —
and exceeding it is a run-ending error on EVERY prompt build for that user,
not a one-off.

This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a
previously-fitting context over. The individual parts are each bounded
(location 200 chars at its producer, locale 35, per-label safe-text
validation), but nothing bounded their SUM, and the connected-channels line
is the one part that grows without limit: up to 20 entries whose names and
presentation hints are only individually capped.

That line now renders as many channels as fit a 1 KiB budget and folds the
rest into the "+N more" counter it already carried, so the fixed guidance
can never be squeezed out by variable content. The worst-case test that
found this stays as the pin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(outbound): resolve outbound capabilities as the acting user

`builtin.outbound_delivery_targets_list` and
`builtin.notification_channels_set` derived their caller from
`effective_user_id`, which prefers the thread owner over the actor. Those
agree on a direct message and on an automation fire, but diverge on a
shared-route channel conversation, where the owner is the route's
configured subject — the deployment operator by default
(`channel_workflow.rs`) — and the actor is whoever posted.

So any participant of a shared channel could enumerate the operator's
connected destinations and rewrite the operator's notification-channel set,
which is where approval prompts, re-auth prompts and failure notices are
delivered. The caller now follows the acting user, matching the fix already
applied to `builtin.outbound_deliver`.

This deliberately REVERSES a previously pinned preference. Two tests
asserted the owner won when the two differ; that pin predates shared-route
subjects defaulting to the operator, and it contradicts the rule that a run
acts as whoever invoked it. Both are updated to pin the actor, with the
reversal recorded at each site rather than silently relaxed, and the
notification-channel write is now asserted to land under the acting user
with the thread owner's own set left untouched.

INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run`
still follow the owner, because they scope the approval-gate raise and the
capability lease and those must stay matched between raise and resume.
Unifying them belongs with the follow-up that removes shared-route subject
binding entirely so a shared channel runs wholly as its invoker; that needs
approval raise/resume coverage which does not exist yet. A new test pins
the split so the interim state is explicit rather than accidental.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep loop_contracts under its size ceiling and ratchet down

The runtime-context byte-budget fix and its worst-case pin pushed
`ironclaw_loop_contracts` to 14,032 production lines against a 13,949
ceiling. Resolved by the reduction the gate prefers over a raise:
`runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim
into a `runtime_context/tests.rs` sibling, which `production_rust_files`
excludes (an inline test module inside a production file is counted; a
test-only file is not).

The crate now measures 13,115 — 834 lines below the previous ceiling and
smaller than before this review round — so the ceiling is re-captured
downward at the measured value rather than raised, per the gate's
one-directional ratchet. Count read from the gate's own failure message,
not by eye.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): chart the notification-channel handlers in the WebUI charter map

`handlers_module_charter` failed: `get_notification_channels` and
`set_notification_channels` — handlers this PR adds — had no sub-owner row
in `CONTRACT.md`'s enforced charter map, and the row they belong to still
named `get_outbound_preferences`, `set_outbound_preferences` and
`outbound_preferences_activity_id`, all deleted by this PR.

Both halves are fixed together because the gate checks both in one test:
unclaimed items first, then entries naming items that no longer exist. Only
the first had fired, so the stale half was still latent behind it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): re-capture the loop_contracts ceiling after main's merge

Main's nearai#7361/nearai#7363 landed 66 lines in
`ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked
up when folding onto main. That put the crate at 13,181 against the 13,115
ceiling re-captured earlier in this PR.

Not growth from this PR. The gate's upward check is a hard `lines > ceiling`
with no headroom — `TOLERANCE` (400) governs only the downward
ratchet-nudge — so a ceiling captured at the exact observed value reddens
every open branch the moment anyone adds a line to that crate, including
from main. Re-captured at the measured value; count read from the gate's own
failure message.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…mpt-surface PRs (nearai#7371)

* fix(ci): schedule the golden lane for prompt-surface production changes

A production change to the model-visible prompt surface (the capability
surface digest, the instruction bundle, the communication-context
renderer, or a shipped loop-tier prompt asset) ran only crate buckets on
the PR lane, so stale golden_payload snapshots surfaced first as a
merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface
digest; the PR lane never ran the golden bucket). Add a curated
prompt-surface owner table that ADDITIONALLY schedules the golden
integration lane without consuming the path's normal package
classification. Self-tested per entry plus a negative control pinning
that ordinary production changes keep the narrow plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): recapture the extension_host coverage floor with measured wobble tolerance

The 2026-08-04 entry was an adjustment whose own text ordered the next
real recapture. Since then code left the crate (denominator 24529 ->
24415) while the ratio ROSE to 88.12%, and the same commit df90072
measured >=21560 covered lines in the merge-queue lane but 21515 twice
on the push lane — a >=45-line same-commit spread over a 20-line
tolerance, redding main on noise (run 31208592262). Recapture both
fields from that run's own gate output and size tolerance_lines to the
measured cross-lane wobble.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): cover the host-managed prompt composer and derive per-entry tests

Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives
the InstructionBundleBuilder, model.rs shapes the pinned request) was
missing from the prompt-surface table — the exact gap class the mapping
exists to close. And the self-test enumerated entries by hand, so a new
entry could ship untested. Add the host_managed_ports prefix and derive
the positive cases from the tables themselves; an entry whose crate the
fixture lacks now fails the suite explicitly instead of skipping.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…ead gates deleted (nearai#7373)

* test(architecture): drop the dead ironclaw_storage row and arm the substrate list

Gate-audit finding (open-and-shut): SUBSTRATE_CRATES in
reborn_composition_boundaries.rs carried three rows of rot, all invisible
because the loop's `let Some(..) else { continue }` silently skipped any
entry that resolves to no workspace package:

- "ironclaw_storage": no such package exists (verified against
  `cargo metadata --no-deps`; the only MISSING name of the 29 listed).
- "ironclaw_approvals" and "ironclaw_assistant" were each listed twice.

The silent skip is replaced with a panic naming the stale entry, so the
list can no longer rot invisibly. Verified by sabotage: adding a bogus
"ironclaw_zzz_probe" row now fails the test with
"is listed in SUBSTRATE_CRATES but is not a workspace package"; the
clean list passes (23/23).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): prune the dead sanctioned path from the specificity gate

Gate-audit finding (open-and-shut): SANCTIONED_PATHS in
reborn_extension_specificity.rs still exempted
`extension_host/extension_installation_store.rs` — a file deleted by
nearai#6430. No scanned path matches the fragment (verified with rg across
crates/), so the entry exempted nothing; it is also the one exclusion
surface in this gate with no staleness check, which is how it outlived
its file. Full specificity suite green after removal (8/8).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): drop the v1 ironclaw_gateway/static exclusions from the telegram gates

Gate-audit finding (open-and-shut): both cross-tree scans in
telegram_extension_gates.rs still carved out `ironclaw_gateway/static`
— the v1 monolith's embedded UI, whose crate was deleted with the src/
monolith (no crates/*/ironclaw_gateway directory exists). The exclusions
matched nothing; scans now cover the whole tree with no dead carve-outs.
Suite green after removal (12/12).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): make the dto-collapse gate's header describe the gate that exists

Gate-audit finding (open-and-shut doc rot): the module doc still
described the pre-nearai#6447 freeze design — a dangling doc-link to
FROZEN_COLLAPSE_DTOS (renamed RETIRED_COLLAPSE_DTOS in nearai#6447), a
promised delete-without-trimming failure and an empty-allowlist
assertion that do not exist in the file, and a named owner for a
collapse that completed. The mechanism itself is armed and untouched;
the header now describes the permanent zero-gate it became, and records
the two originally-frozen names that deliberately left governance
(CapabilityOutcome via nearai#6299 deletion, CapabilityDispatchRequest blessed
as the canonical port type). Suite green (2/2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): repoint the manifest-reparse allowlist note at the colocated asset

Gate-audit finding (open-and-shut doc rot): the BundledAsset allowlist
entry's justification still cited include_str! of
assets/memory_native/manifest.toml — a path retired when WS2 (nearai#7037)
colocated packages; the live include in memory_native_extension.rs
reaches crates/extensions/packages/memory-native/manifest.toml. Comment
only; the gate's mechanism and counts are untouched. Suite green (2/2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): give the memory-vocabulary gate the partial-tree floor its twin has

Gate-audit finding: reborn_memory_retired_vocabulary.rs had no
MIN_SCANNED_FILES floor, unlike its explicit twin
reborn_retired_taxonomy.rs — so a partially-moved tree (the CHECKLIST
WS0 / nearai#6963 'green while measuring nothing' shape) would scan a
fraction of the files and still report the vocabulary clean. The gate
was in fact born with an already-dead sanctioned path (its own header
records this), so the rot class is not hypothetical for this file.

Adds the same 500-file floor (real count ~4000), asserts it in the main
gate, and pins the premise on a fixture: a 10-file partial tree scans
clean and is rejected by the floor. Suite green (4/4); clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): close the transport gate's nested-use-group fail-open

Gate-audit finding (sabotage-verified): product_symbols_in's braced-group
branch closed at the FIRST '}' (group.find('}')), so a nested group —
use ironclaw_assistant::{m::{X}}; — truncated mid-element and recorded
zero symbols. Probed live before the fix: appending
use ironclaw_assistant::{zzz_audit::{ZzzProbe}}; to webui's lib.rs left
transports_name_only_the_frozen_residue_of_product_symbols GREEN, while
the plain-path spelling of the same import correctly failed. The same
truncation dropped qualified elements inside flat groups
({qualified_module::X} recorded nothing).

The group branch now does a balanced-brace walk, splits elements at
depth-0 commas only, and records a qualified/nested element's leading
path segment — the same key the single-path branch records for
ironclaw_assistant::module::X. Flat-element semantics are byte-for-byte
unchanged, so the frozen 100-row webui inventory is untouched (suite
green 6/6 on the live tree). Regression fixtures added to
import_scanner_reads_symbols_out_of_real_use_shapes; the original
sabotage now fails with the gate's own message (re-verified).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: delete check-e2e-matrix-files.sh — a gate for a workflow that no longer exists

Gate-audit finding (provably inert): the script's default target is
.github/workflows/e2e.yml, deleted when the v1 e2e suites were retired
(git log --diff-filter=D shows the removing commit); no workflow, script,
hook, doc, or guidance file references check-e2e-matrix-files.sh
(verified with rg across the repo including .github and .githooks).
A checker nothing runs, pointed at a file nothing provides, is dead
weight that reads as coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: delete the measured-broken check-boundaries.sh and its guidance references

Gate-audit finding (provably inert, previously measured): crates/AGENTS.md
recorded on 2026-08-05 that the script fails on a clean tree (check 5
false-positives on live test files) and that checks 1/2/3/6 target the
deleted v1 src/ tree, passing vacuously. No workflow or hook runs it; its
only callers were guidance files, two of which claimed it 'enforces'
root-tests feature gating — an enforcement claim the skill-maintainer
rules forbid for a check nothing executes.

Removed the script and every live reference: the crates/AGENTS.md warning
row becomes a tombstone note; the testing skill + exemplar reference drop
the false enforcement parenthetical; the architecture-review skill's
Verify line drops the dead command; deslop-reborn's allowed-tools drops
the permission; .coderabbit.yaml's driver-leak instruction now points at
the live enforcement (reborn_persistence_driver_boundary). Two dated
docs/internal/ plan snapshots keep their historical mentions.

Verified: python3 scripts/ci/check-guidance.py OK (2084 path references)
and its self-test OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(product): stop hardcoding charter sub-owner counts in the family map

Gate-audit finding (stale prose): crates/product/AGENTS.md said
'19-sub-owner reborn_services charter map' — the enforced map has had 20
sub-owners since nearai#7235 added the inspector row (counted from the live
table). Rather than chase the number, drop both inline counts: the
owning maps and their gates are authoritative, and the re-verify
commands are already inline (skill-maintainer rule: no counts without a
regeneration recipe). check-guidance.py OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): correct the scanner-fixture file's name-filter claim

Gate-audit finding (doc rot with a false coverage claim): the header
said naming the FILE reborn_* makes code_style.yml's
'cargo test -p ironclaw_architecture_tests reborn' see it — but that
argument is a test-NAME filter (the measurement is documented in
reborn_contracts_vendor_census.rs), and none of this file's test fns
contains the substring, so that smoke lane runs 0 of them (11 collected
by the full plan). Comment-only; the note now records the real semantics
so file names are not trusted for lane coverage. Suite green (11/11).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): gate & ratchet audit report + proposed preflight gauntlet

The audit the owner asked for after PR nearai#7157 went red six times across
four gates: every architecture-test gate, module charter, CI script, and
committed baseline inventoried with a verdict and evidence; the handful
worth acting on ranked by friction x weakness; the CI-ergonomics analysis
(why failures surface one per ~1h round-trip: no --no-fail-fast anywhere
in CI, cancel-in-progress on push, sequential fast-checks steps —
measured: two broken gates report 1 failure in 18s under the CI shape vs
both in 211s with --no-fail-fast); and the sabotage log for every probe.

scripts/preflight-gates.sh is the concrete pre-push proposal: the
deterministic-gate classes only (script gates ~10s + architecture suite
--no-fail-fast + changed-crate charter tests), covering all four nearai#7157
gate classes locally in one command. Unwired — nothing invokes it.
Validated end-to-end on this branch: exit 0, 'every deterministic gate
green', 402.8s including gate-binary recompiles.

Placement verified: python3 scripts/ci/docs_publication_boundary.py OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(planner): classify preflight-gates.sh and the deleted check-boundaries.sh

The gate audit's own PR hit the planner's fail-closed arm — 'unmapped
test or CI path: scripts/check-boundaries.sh' — exactly the class the
arm exists to force a decision on (and the audit's report documents).
Per the PR_STATIC_CONTROL_PATHS membership rule (no Reborn test lane
exercises either file):

- scripts/preflight-gates.sh — the audit's proposed local pre-push
  gauntlet; referenced by no workflow.
- scripts/check-boundaries.sh — deleted by the audit; the entry lets the
  deletion diff (and any revert) classify instead of failing every
  downstream Reborn lane.

Verified: the planner now produces mode=selected with the
architecture-misc bucket for this branch's diff, and
python3 scripts/ci/test_reborn_pr_test_plan.py is OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): add the fold-tripped asymmetric-tolerance exhibit to the audit

The strongest single exhibit for shortlist item 2, contributed by the
nearai#7157 branch steward after this audit's cutoff and verified against the
gate's code: TOLERANCE = 400 is consulted in exactly one direction (the
banked-slack check, ceiling.saturating_sub(lines) > TOLERANCE); the
growth check is a bare lines > ceiling. With the in-file 'set to
current, not padded' instruction, every ceiling is a hard cap at the
observed count — so one line landing on main in any contracts crate
reds every open branch at its next fold until someone re-captures.

Measured recurrence on nearai#7157: loop_contracts re-captured four times,
~once per fold (14,479 -> 13,850 -> 13,949 -> 13,115 -> 13,181), the
last tripped by main's nearai#7361/nearai#7363 adding 66 lines to
instruction_bundle.rs — nothing the branch wrote. All four deltas were
<= 105 lines: either repair shape in §3.2 (one-line upward tolerance
using the existing constant, or mid-window pinning) would have absorbed
every one with zero red builds. This audit's own sabotage already
proved the jaws (+1 line host_api red / -1 line common red); the fold
history shows the operational cost. The repair stays a recommendation —
adding growth headroom to a ratchet is the owner's call, not this PR's.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): give the contracts size ceiling upward working slack

Owner-directed repair of the audit's sharpest finding (report §3.2): the
gate's TOLERANCE = 400 was consulted in exactly one direction — the
banked-slack check — while the growth check was a bare lines > ceiling.
Combined with 'set to current, not padded' pins, every ceiling was a hard
cap at the exact observed count, so one line landing on main in any
contracts crate redded every open branch at its next fold until someone
re-captured. Measured on nearai#7157: four loop_contracts re-captures, roughly
once per fold, every delta <= 105 lines — the gate generating its own
busywork.

The growth check now allows GROWTH_TOLERANCE = 150 of working slack
above each pin (sized to composition-budget precedent; the reviewed
raises this gate has caught were +1,069 and +1,214 lines, far above it),
and all six ceilings are re-pinned to the counts the test itself
reported with every ceiling at 0 — which also removes the +400 seed
padding on common/loop_contracts/prompt_envelope that contradicted the
capture rule and put those crates one deleted line from the banked jaw.

Sabotage-verified both ways: +1 line in host_api and -1 line in common —
both red before this change — now pass; a +151-line probe still fails
with the effective-ceiling arithmetic in the message. Full
reborn_dependency_boundaries binary green (41/41); clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(budget): re-equalize composition pins to observed — restore the working window

Owner-directed companion to the contracts-ceiling repair (same annoying
class, other mass gate): merged main-side growth since the 2026-08-05
equalization had drifted +101 LOC and +5 Arc<dyn> sites through the
tolerance windows, leaving 49 LOC / 10 sites of live headroom — the next
routine composition PR would have gone red on wiring alone (the gate
audit measured this the same day it was pinned).

Per the TOML's own maintenance instructions: loc_ceiling/loc_observed
40423 -> 40524 and arc_dyn 814 -> 819, measured with the gate's --print,
set to current not padded, dated notes appended (not overwritten), and
the arch-test record (COMPOSITION_ABSOLUTE_SRC_LOC) moved in the same
commit as its file requires. ceiling_bp stays 658 — the WS0 floor is
deliberately not re-set.

Verified: check-composition-budget.sh OK; its 76-case self-test green;
reborn_restructure_baselines green; probe +100 LOC now passes (was red
at 49 headroom), probe +160 LOC still fails.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): record the landed zero-slack repairs in the audit report

The §3.2 repair moved from recommendation to landed at owner direction;
the report's answer, inventory rows, and §7 ledger now say so, with the
counting-rule fix promoted to the top remaining recommendation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* gates: pin the ceiling-window arithmetic; fail preflight discovery closed

Two review-round hardenings (the open CodeRabbit Majors):

- reborn_dependency_boundaries.rs: extract the size-ceiling comparison into
  contracts_ceiling_verdict() and pin its four window edges with a committed
  regression test (contracts_size_ceiling_window_edges_hold) — accept at
  ceiling+GROWTH_TOLERANCE, reject one line past, accept at
  ceiling-TOLERANCE, reject one banked line further, and a zero-measure scan
  reads Banked, never a silent pass. The pre-repair asymmetry (tolerance
  consulted only downward) can no longer return silently. Live-gate behavior
  re-probed unchanged after the rewiring: +1 line to host_api passes, +151
  fails with the same effective-ceiling message.
- preflight-gates.sh: setup and changed-file discovery now fail closed — a
  missing repo root exits 2, and a failed merge-base/diff widens the charter
  run to all five crates instead of silently skipping them (the same
  fallback the missing-base branch already used). A broken setup may cost
  compile time, never a silent skip.

Full boundary binary 42/42 green; clippy clean; preflight-gates.sh
end-to-end green on this tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-7361 — e541958c Deployed Aug 7, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants