fix(extensions): chat "connect account" dead-end — already-connected signal, builtin description trust, docs - #7361
Conversation
…nstall confirmation
Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread
e79a994f, run 251aec0b on ironclaw-qa-testing-libsql):
1. Host-bundled capability descriptions were description_trust=Untrusted,
so the loop-tier prompt-text denylist strict-scanned compiled-in text
and silently omitted builtin.extension_register_hosted_mcp from every
model prompt's capability surface ("browser authorization-code flow"
matched the "authorization" credential pattern). HostBundled is the
only source eligible for effective FirstParty/System trust, so its
repo-authored descriptions now cross the verified-catalog boundary
like signature/digest-verified registry installs. Untrusted provenance
(InstalledLocal, UserRegistered, unknown) keeps the strict scan.
2. When install-driven activation passed the credential gate because the
caller's declared requirements were all satisfied, the response never
said so — the model got only conditional guidance ("If WebChat shows
an account connection panel...") and deflected an explicit "connect
account" request to the web interface even though the account was
already connected. The install response now appends an explicit
already-connected confirmation exactly when declared requirements
were verified present for the calling user.
Regression tests: manager surface test pins VerifiedCatalog trust for all
model-visible lifecycle capabilities through the real host runtime;
instruction-bundle tests pin retain/omit behavior for auth-vocabulary
descriptions by trust; install-path tests pin the confirmation on the
seeded-credential path and its absence for credential-free extensions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
🚅 Deployed to the ironclaw-pr-7361 environment in ironclaw-ci-preview
|
|
Important Review skippedReview was skipped due to path filters ⛔ Files ignored due to path filters (4)
CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
📝 WalkthroughSummary by CodeRabbit
WalkthroughHost-bundled capability descriptions now use verified catalog trust. Install activation reports an existing connection when verified credentials are already satisfied. Tests cover filtering and activation behavior. Channel documentation separates operator credential setup from personal connection flows. ChangesCapability trust and activation guidance
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant User
participant ExtensionActivation
participant CapabilitySurface
participant OAuthOrPairingPanel
User->>ExtensionActivation: Ask agent to connect a channel
ExtensionActivation->>CapabilitySurface: Expose trusted activation capability
ExtensionActivation-->>User: Report activation and connection state
ExtensionActivation->>OAuthOrPairingPanel: Open personal OAuth or pairing when required
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
🧭 IronLoop Run · ReviewThis comment updates in place as the Run moves through its stages. ⬛ Final result · Stopped
Automatic trigger · attempt 1 of 3 · stopped after 5m 56s IronLoop stopped because the pull request target branch or head changed while this Run was active. |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/channels/overview.mdx`:
- Around line 96-107: Keep docs/channels/overview.mdx lines 96-107 as the
canonical distinction, update docs/channels/slack.mdx lines 45-62 to state that
WebUI registers instance credentials while chat installs or activates the
extension and completes or confirms personal OAuth, and align docs/onboard.mdx
lines 179-183 with the same two-step contract; do not claim that agent-driven
Slack connection is unsupported after operator setup.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: f935b6a8-6ea8-4028-9b4e-2a2472d6ca05
📒 Files selected for processing (6)
crates/contracts/ironclaw_loop_contracts/src/instruction_bundle.rscrates/domains/ironclaw_llm/src/nearai_chat.rscrates/extensions/ironclaw_extension_manager/src/extension_lifecycle_capabilities.rscrates/kernel/ironclaw_host_runtime/src/surface.rsdocs/channels/overview.mdxdocs/onboard.mdx
The onboarding and channels pages claimed "asking the agent to connect a
channel doesn't work" and that the agent "may tell you it can't help".
That describes only the operator half (registering app/bot credentials).
The per-user half has shipped since early July: extension_install runs the
same activation credential gate as the Channels card, raises the in-chat
OAuth connection panel when the account is unconnected, and (as of the
sibling fix) confirms when it is already connected.
The self-knowledge protocol makes these pages the model's authority on
IronClaw's own capabilities, so the stale claim scripted the exact
refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an
account that was already connected. Correct both pages to distinguish
the operator step from the chat-drivable personal connect.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
f53ebed to
b5b1366
Compare
… contract The slack page's operator-step note and the telegram troubleshooting accordion still carried the blanket "asking the agent to connect will not work" claim the overview/onboarding correction removed — same drift, different phrasing (review catch on #7361, plus one more instance found by a broader sweep). Both now state the two-step contract: the operator half stays in the web interface; after it, chat drives the personal half (install/activate -> in-chat connection or pairing panel, or an already-connected confirmation). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ange The queue run failed golden_payload because the branch predated current main and its own surface.rs trust fix changes the surface digest. The recaptured snapshots differ ONLY in the surface sha256 lines (verified char-by-char) — no prompt text or capability-list changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Queue failure diagnosed and fixed (pushed
Unrelated but observed in the same window: main's push run at Needs a fresh re-queue when checks come back green. |
Main's #7361/#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…the audit The strongest single exhibit for shortlist item 2, contributed by the #7157 branch steward after this audit's cutoff and verified against the gate's code: TOLERANCE = 400 is consulted in exactly one direction (the banked-slack check, ceiling.saturating_sub(lines) > TOLERANCE); the growth check is a bare lines > ceiling. With the in-file 'set to current, not padded' instruction, every ceiling is a hard cap at the observed count — so one line landing on main in any contracts crate reds every open branch at its next fold until someone re-captures. Measured recurrence on #7157: loop_contracts re-captured four times, ~once per fold (14,479 -> 13,850 -> 13,949 -> 13,115 -> 13,181), the last tripped by main's #7361/#7363 adding 66 lines to instruction_bundle.rs — nothing the branch wrote. All four deltas were <= 105 lines: either repair shape in §3.2 (one-line upward tolerance using the existing constant, or mid-window pinning) would have absorbed every one with zero red builds. This audit's own sabotage already proved the jaws (+1 line host_api red / -1 line common red); the fold history shows the operational cost. The repair stays a recommendation — adding growth headroom to a ratchet is the owner's call, not this PR's. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…mpt-surface PRs (nearai#7371) * fix(ci): schedule the golden lane for prompt-surface production changes A production change to the model-visible prompt surface (the capability surface digest, the instruction bundle, the communication-context renderer, or a shipped loop-tier prompt asset) ran only crate buckets on the PR lane, so stale golden_payload snapshots surfaced first as a merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface digest; the PR lane never ran the golden bucket). Add a curated prompt-surface owner table that ADDITIONALLY schedules the golden integration lane without consuming the path's normal package classification. Self-tested per entry plus a negative control pinning that ordinary production changes keep the narrow plan. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): recapture the extension_host coverage floor with measured wobble tolerance The 2026-08-04 entry was an adjustment whose own text ordered the next real recapture. Since then code left the crate (denominator 24529 -> 24415) while the ratio ROSE to 88.12%, and the same commit df90072 measured >=21560 covered lines in the merge-queue lane but 21515 twice on the push lane — a >=45-line same-commit spread over a 20-line tolerance, redding main on noise (run 31208592262). Recapture both fields from that run's own gate output and size tolerance_lines to the measured cross-lane wobble. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): cover the host-managed prompt composer and derive per-entry tests Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives the InstructionBundleBuilder, model.rs shapes the pinned request) was missing from the prompt-surface table — the exact gap class the mapping exists to close. And the self-test enumerated entries by hand, so a new entry could ship untested. Add the host_managed_ports prefix and derive the positive cases from the tables themselves; an entry whose crate the fixture lacks now fails the suite explicitly instead of skipping. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ead gates deleted (nearai#7373) * test(architecture): drop the dead ironclaw_storage row and arm the substrate list Gate-audit finding (open-and-shut): SUBSTRATE_CRATES in reborn_composition_boundaries.rs carried three rows of rot, all invisible because the loop's `let Some(..) else { continue }` silently skipped any entry that resolves to no workspace package: - "ironclaw_storage": no such package exists (verified against `cargo metadata --no-deps`; the only MISSING name of the 29 listed). - "ironclaw_approvals" and "ironclaw_assistant" were each listed twice. The silent skip is replaced with a panic naming the stale entry, so the list can no longer rot invisibly. Verified by sabotage: adding a bogus "ironclaw_zzz_probe" row now fails the test with "is listed in SUBSTRATE_CRATES but is not a workspace package"; the clean list passes (23/23). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): prune the dead sanctioned path from the specificity gate Gate-audit finding (open-and-shut): SANCTIONED_PATHS in reborn_extension_specificity.rs still exempted `extension_host/extension_installation_store.rs` — a file deleted by nearai#6430. No scanned path matches the fragment (verified with rg across crates/), so the entry exempted nothing; it is also the one exclusion surface in this gate with no staleness check, which is how it outlived its file. Full specificity suite green after removal (8/8). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): drop the v1 ironclaw_gateway/static exclusions from the telegram gates Gate-audit finding (open-and-shut): both cross-tree scans in telegram_extension_gates.rs still carved out `ironclaw_gateway/static` — the v1 monolith's embedded UI, whose crate was deleted with the src/ monolith (no crates/*/ironclaw_gateway directory exists). The exclusions matched nothing; scans now cover the whole tree with no dead carve-outs. Suite green after removal (12/12). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): make the dto-collapse gate's header describe the gate that exists Gate-audit finding (open-and-shut doc rot): the module doc still described the pre-nearai#6447 freeze design — a dangling doc-link to FROZEN_COLLAPSE_DTOS (renamed RETIRED_COLLAPSE_DTOS in nearai#6447), a promised delete-without-trimming failure and an empty-allowlist assertion that do not exist in the file, and a named owner for a collapse that completed. The mechanism itself is armed and untouched; the header now describes the permanent zero-gate it became, and records the two originally-frozen names that deliberately left governance (CapabilityOutcome via nearai#6299 deletion, CapabilityDispatchRequest blessed as the canonical port type). Suite green (2/2). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): repoint the manifest-reparse allowlist note at the colocated asset Gate-audit finding (open-and-shut doc rot): the BundledAsset allowlist entry's justification still cited include_str! of assets/memory_native/manifest.toml — a path retired when WS2 (nearai#7037) colocated packages; the live include in memory_native_extension.rs reaches crates/extensions/packages/memory-native/manifest.toml. Comment only; the gate's mechanism and counts are untouched. Suite green (2/2). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): give the memory-vocabulary gate the partial-tree floor its twin has Gate-audit finding: reborn_memory_retired_vocabulary.rs had no MIN_SCANNED_FILES floor, unlike its explicit twin reborn_retired_taxonomy.rs — so a partially-moved tree (the CHECKLIST WS0 / nearai#6963 'green while measuring nothing' shape) would scan a fraction of the files and still report the vocabulary clean. The gate was in fact born with an already-dead sanctioned path (its own header records this), so the rot class is not hypothetical for this file. Adds the same 500-file floor (real count ~4000), asserts it in the main gate, and pins the premise on a fixture: a 10-file partial tree scans clean and is rejected by the floor. Suite green (4/4); clippy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): close the transport gate's nested-use-group fail-open Gate-audit finding (sabotage-verified): product_symbols_in's braced-group branch closed at the FIRST '}' (group.find('}')), so a nested group — use ironclaw_assistant::{m::{X}}; — truncated mid-element and recorded zero symbols. Probed live before the fix: appending use ironclaw_assistant::{zzz_audit::{ZzzProbe}}; to webui's lib.rs left transports_name_only_the_frozen_residue_of_product_symbols GREEN, while the plain-path spelling of the same import correctly failed. The same truncation dropped qualified elements inside flat groups ({qualified_module::X} recorded nothing). The group branch now does a balanced-brace walk, splits elements at depth-0 commas only, and records a qualified/nested element's leading path segment — the same key the single-path branch records for ironclaw_assistant::module::X. Flat-element semantics are byte-for-byte unchanged, so the frozen 100-row webui inventory is untouched (suite green 6/6 on the live tree). Regression fixtures added to import_scanner_reads_symbols_out_of_real_use_shapes; the original sabotage now fails with the gate's own message (re-verified). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: delete check-e2e-matrix-files.sh — a gate for a workflow that no longer exists Gate-audit finding (provably inert): the script's default target is .github/workflows/e2e.yml, deleted when the v1 e2e suites were retired (git log --diff-filter=D shows the removing commit); no workflow, script, hook, doc, or guidance file references check-e2e-matrix-files.sh (verified with rg across the repo including .github and .githooks). A checker nothing runs, pointed at a file nothing provides, is dead weight that reads as coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: delete the measured-broken check-boundaries.sh and its guidance references Gate-audit finding (provably inert, previously measured): crates/AGENTS.md recorded on 2026-08-05 that the script fails on a clean tree (check 5 false-positives on live test files) and that checks 1/2/3/6 target the deleted v1 src/ tree, passing vacuously. No workflow or hook runs it; its only callers were guidance files, two of which claimed it 'enforces' root-tests feature gating — an enforcement claim the skill-maintainer rules forbid for a check nothing executes. Removed the script and every live reference: the crates/AGENTS.md warning row becomes a tombstone note; the testing skill + exemplar reference drop the false enforcement parenthetical; the architecture-review skill's Verify line drops the dead command; deslop-reborn's allowed-tools drops the permission; .coderabbit.yaml's driver-leak instruction now points at the live enforcement (reborn_persistence_driver_boundary). Two dated docs/internal/ plan snapshots keep their historical mentions. Verified: python3 scripts/ci/check-guidance.py OK (2084 path references) and its self-test OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(product): stop hardcoding charter sub-owner counts in the family map Gate-audit finding (stale prose): crates/product/AGENTS.md said '19-sub-owner reborn_services charter map' — the enforced map has had 20 sub-owners since nearai#7235 added the inspector row (counted from the live table). Rather than chase the number, drop both inline counts: the owning maps and their gates are authoritative, and the re-verify commands are already inline (skill-maintainer rule: no counts without a regeneration recipe). check-guidance.py OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): correct the scanner-fixture file's name-filter claim Gate-audit finding (doc rot with a false coverage claim): the header said naming the FILE reborn_* makes code_style.yml's 'cargo test -p ironclaw_architecture_tests reborn' see it — but that argument is a test-NAME filter (the measurement is documented in reborn_contracts_vendor_census.rs), and none of this file's test fns contains the substring, so that smoke lane runs 0 of them (11 collected by the full plan). Comment-only; the note now records the real semantics so file names are not trusted for lane coverage. Suite green (11/11). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(internal): gate & ratchet audit report + proposed preflight gauntlet The audit the owner asked for after PR nearai#7157 went red six times across four gates: every architecture-test gate, module charter, CI script, and committed baseline inventoried with a verdict and evidence; the handful worth acting on ranked by friction x weakness; the CI-ergonomics analysis (why failures surface one per ~1h round-trip: no --no-fail-fast anywhere in CI, cancel-in-progress on push, sequential fast-checks steps — measured: two broken gates report 1 failure in 18s under the CI shape vs both in 211s with --no-fail-fast); and the sabotage log for every probe. scripts/preflight-gates.sh is the concrete pre-push proposal: the deterministic-gate classes only (script gates ~10s + architecture suite --no-fail-fast + changed-crate charter tests), covering all four nearai#7157 gate classes locally in one command. Unwired — nothing invokes it. Validated end-to-end on this branch: exit 0, 'every deterministic gate green', 402.8s including gate-binary recompiles. Placement verified: python3 scripts/ci/docs_publication_boundary.py OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(planner): classify preflight-gates.sh and the deleted check-boundaries.sh The gate audit's own PR hit the planner's fail-closed arm — 'unmapped test or CI path: scripts/check-boundaries.sh' — exactly the class the arm exists to force a decision on (and the audit's report documents). Per the PR_STATIC_CONTROL_PATHS membership rule (no Reborn test lane exercises either file): - scripts/preflight-gates.sh — the audit's proposed local pre-push gauntlet; referenced by no workflow. - scripts/check-boundaries.sh — deleted by the audit; the entry lets the deletion diff (and any revert) classify instead of failing every downstream Reborn lane. Verified: the planner now produces mode=selected with the architecture-misc bucket for this branch's diff, and python3 scripts/ci/test_reborn_pr_test_plan.py is OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(internal): add the fold-tripped asymmetric-tolerance exhibit to the audit The strongest single exhibit for shortlist item 2, contributed by the nearai#7157 branch steward after this audit's cutoff and verified against the gate's code: TOLERANCE = 400 is consulted in exactly one direction (the banked-slack check, ceiling.saturating_sub(lines) > TOLERANCE); the growth check is a bare lines > ceiling. With the in-file 'set to current, not padded' instruction, every ceiling is a hard cap at the observed count — so one line landing on main in any contracts crate reds every open branch at its next fold until someone re-captures. Measured recurrence on nearai#7157: loop_contracts re-captured four times, ~once per fold (14,479 -> 13,850 -> 13,949 -> 13,115 -> 13,181), the last tripped by main's nearai#7361/nearai#7363 adding 66 lines to instruction_bundle.rs — nothing the branch wrote. All four deltas were <= 105 lines: either repair shape in §3.2 (one-line upward tolerance using the existing constant, or mid-window pinning) would have absorbed every one with zero red builds. This audit's own sabotage already proved the jaws (+1 line host_api red / -1 line common red); the fold history shows the operational cost. The repair stays a recommendation — adding growth headroom to a ratchet is the owner's call, not this PR's. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): give the contracts size ceiling upward working slack Owner-directed repair of the audit's sharpest finding (report §3.2): the gate's TOLERANCE = 400 was consulted in exactly one direction — the banked-slack check — while the growth check was a bare lines > ceiling. Combined with 'set to current, not padded' pins, every ceiling was a hard cap at the exact observed count, so one line landing on main in any contracts crate redded every open branch at its next fold until someone re-captured. Measured on nearai#7157: four loop_contracts re-captures, roughly once per fold, every delta <= 105 lines — the gate generating its own busywork. The growth check now allows GROWTH_TOLERANCE = 150 of working slack above each pin (sized to composition-budget precedent; the reviewed raises this gate has caught were +1,069 and +1,214 lines, far above it), and all six ceilings are re-pinned to the counts the test itself reported with every ceiling at 0 — which also removes the +400 seed padding on common/loop_contracts/prompt_envelope that contradicted the capture rule and put those crates one deleted line from the banked jaw. Sabotage-verified both ways: +1 line in host_api and -1 line in common — both red before this change — now pass; a +151-line probe still fails with the effective-ceiling arithmetic in the message. Full reborn_dependency_boundaries binary green (41/41); clippy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(budget): re-equalize composition pins to observed — restore the working window Owner-directed companion to the contracts-ceiling repair (same annoying class, other mass gate): merged main-side growth since the 2026-08-05 equalization had drifted +101 LOC and +5 Arc<dyn> sites through the tolerance windows, leaving 49 LOC / 10 sites of live headroom — the next routine composition PR would have gone red on wiring alone (the gate audit measured this the same day it was pinned). Per the TOML's own maintenance instructions: loc_ceiling/loc_observed 40423 -> 40524 and arc_dyn 814 -> 819, measured with the gate's --print, set to current not padded, dated notes appended (not overwritten), and the arch-test record (COMPOSITION_ABSOLUTE_SRC_LOC) moved in the same commit as its file requires. ceiling_bp stays 658 — the WS0 floor is deliberately not re-set. Verified: check-composition-budget.sh OK; its 76-case self-test green; reborn_restructure_baselines green; probe +100 LOC now passes (was red at 49 headroom), probe +160 LOC still fails. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(internal): record the landed zero-slack repairs in the audit report The §3.2 repair moved from recommendation to landed at owner direction; the report's answer, inventory rows, and §7 ledger now say so, with the counting-rule fix promoted to the top remaining recommendation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * gates: pin the ceiling-window arithmetic; fail preflight discovery closed Two review-round hardenings (the open CodeRabbit Majors): - reborn_dependency_boundaries.rs: extract the size-ceiling comparison into contracts_ceiling_verdict() and pin its four window edges with a committed regression test (contracts_size_ceiling_window_edges_hold) — accept at ceiling+GROWTH_TOLERANCE, reject one line past, accept at ceiling-TOLERANCE, reject one banked line further, and a zero-measure scan reads Banked, never a silent pass. The pre-repair asymmetry (tolerance consulted only downward) can no longer return silently. Live-gate behavior re-probed unchanged after the rewiring: +1 line to host_api passes, +151 fails with the same effective-ceiling message. - preflight-gates.sh: setup and changed-file discovery now fail closed — a missing repo root exits 2, and a failed merge-base/diff widens the charter run to all five crates instead of silently skipping them (the same fallback the missing-base branch already used). A broken setup may cost compile time, never a silent skip. Full boundary binary 42/42 green; clippy clean; preflight-gates.sh end-to-end green on this tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…signal, builtin description trust, docs (nearai#7361) * fix(extensions): host-bundled description trust + already-connected install confirmation Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread e79a994f, run 251aec0b on ironclaw-qa-testing-libsql): 1. Host-bundled capability descriptions were description_trust=Untrusted, so the loop-tier prompt-text denylist strict-scanned compiled-in text and silently omitted builtin.extension_register_hosted_mcp from every model prompt's capability surface ("browser authorization-code flow" matched the "authorization" credential pattern). HostBundled is the only source eligible for effective FirstParty/System trust, so its repo-authored descriptions now cross the verified-catalog boundary like signature/digest-verified registry installs. Untrusted provenance (InstalledLocal, UserRegistered, unknown) keeps the strict scan. 2. When install-driven activation passed the credential gate because the caller's declared requirements were all satisfied, the response never said so — the model got only conditional guidance ("If WebChat shows an account connection panel...") and deflected an explicit "connect account" request to the web interface even though the account was already connected. The install response now appends an explicit already-connected confirmation exactly when declared requirements were verified present for the calling user. Regression tests: manager surface test pins VerifiedCatalog trust for all model-visible lifecycle capabilities through the real host runtime; instruction-bundle tests pin retain/omit behavior for auth-vocabulary descriptions by trust; install-path tests pin the confirmation on the seeded-credential path and its absence for credential-free extensions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(channels): chat can drive the personal half of channel connect The onboarding and channels pages claimed "asking the agent to connect a channel doesn't work" and that the agent "may tell you it can't help". That describes only the operator half (registering app/bot credentials). The per-user half has shipped since early July: extension_install runs the same activation credential gate as the Channels card, raises the in-chat OAuth connection panel when the account is unconnected, and (as of the sibling fix) confirms when it is already connected. The self-knowledge protocol makes these pages the model's authority on IronClaw's own capabilities, so the stale claim scripted the exact refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an account that was already connected. Correct both pages to distinguish the operator step from the chat-drivable personal connect. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(channels): align slack and telegram setup notes with the connect contract The slack page's operator-step note and the telegram troubleshooting accordion still carried the blanket "asking the agent to connect will not work" claim the overview/onboarding correction removed — same drift, different phrasing (review catch on nearai#7361, plus one more instance found by a broader sweep). Both now state the two-step contract: the operator half stays in the web interface; after it, chat drives the personal half (install/activate -> in-chat connection or pairing panel, or an already-connected confirmation). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(golden): recapture surface digests over the description-trust change The queue run failed golden_payload because the branch predated current main and its own surface.rs trust fix changes the surface digest. The recaptured snapshots differ ONLY in the surface sha256 lines (verified char-by-char) — no prompt text or capability-list changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…mpt-surface PRs (nearai#7371) * fix(ci): schedule the golden lane for prompt-surface production changes A production change to the model-visible prompt surface (the capability surface digest, the instruction bundle, the communication-context renderer, or a shipped loop-tier prompt asset) ran only crate buckets on the PR lane, so stale golden_payload snapshots surfaced first as a merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface digest; the PR lane never ran the golden bucket). Add a curated prompt-surface owner table that ADDITIONALLY schedules the golden integration lane without consuming the path's normal package classification. Self-tested per entry plus a negative control pinning that ordinary production changes keep the narrow plan. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): recapture the extension_host coverage floor with measured wobble tolerance The 2026-08-04 entry was an adjustment whose own text ordered the next real recapture. Since then code left the crate (denominator 24529 -> 24415) while the ratio ROSE to 88.12%, and the same commit df90072 measured >=21560 covered lines in the merge-queue lane but 21515 twice on the push lane — a >=45-line same-commit spread over a 20-line tolerance, redding main on noise (run 31208592262). Recapture both fields from that run's own gate output and size tolerance_lines to the measured cross-lane wobble. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): cover the host-managed prompt composer and derive per-entry tests Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives the InstructionBundleBuilder, model.rs shapes the pinned request) was missing from the prompt-surface table — the exact gap class the mapping exists to close. And the self-test enumerated entries by hand, so a new entry could ship untested. Add the host_managed_ports prefix and derive the positive cases from the tables themselves; an entry whose crate the fixture lacks now fails the suite explicitly instead of skipping. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…signal, builtin description trust, docs (nearai#7361) * fix(extensions): host-bundled description trust + already-connected install confirmation Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread e79a994f, run 251aec0b on ironclaw-qa-testing-libsql): 1. Host-bundled capability descriptions were description_trust=Untrusted, so the loop-tier prompt-text denylist strict-scanned compiled-in text and silently omitted builtin.extension_register_hosted_mcp from every model prompt's capability surface ("browser authorization-code flow" matched the "authorization" credential pattern). HostBundled is the only source eligible for effective FirstParty/System trust, so its repo-authored descriptions now cross the verified-catalog boundary like signature/digest-verified registry installs. Untrusted provenance (InstalledLocal, UserRegistered, unknown) keeps the strict scan. 2. When install-driven activation passed the credential gate because the caller's declared requirements were all satisfied, the response never said so — the model got only conditional guidance ("If WebChat shows an account connection panel...") and deflected an explicit "connect account" request to the web interface even though the account was already connected. The install response now appends an explicit already-connected confirmation exactly when declared requirements were verified present for the calling user. Regression tests: manager surface test pins VerifiedCatalog trust for all model-visible lifecycle capabilities through the real host runtime; instruction-bundle tests pin retain/omit behavior for auth-vocabulary descriptions by trust; install-path tests pin the confirmation on the seeded-credential path and its absence for credential-free extensions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(channels): chat can drive the personal half of channel connect The onboarding and channels pages claimed "asking the agent to connect a channel doesn't work" and that the agent "may tell you it can't help". That describes only the operator half (registering app/bot credentials). The per-user half has shipped since early July: extension_install runs the same activation credential gate as the Channels card, raises the in-chat OAuth connection panel when the account is unconnected, and (as of the sibling fix) confirms when it is already connected. The self-knowledge protocol makes these pages the model's authority on IronClaw's own capabilities, so the stale claim scripted the exact refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an account that was already connected. Correct both pages to distinguish the operator step from the chat-drivable personal connect. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(channels): align slack and telegram setup notes with the connect contract The slack page's operator-step note and the telegram troubleshooting accordion still carried the blanket "asking the agent to connect will not work" claim the overview/onboarding correction removed — same drift, different phrasing (review catch on nearai#7361, plus one more instance found by a broader sweep). Both now state the two-step contract: the operator half stays in the web interface; after it, chat drives the personal half (install/activate -> in-chat connection or pairing panel, or an already-connected confirmation). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(golden): recapture surface digests over the description-trust change The queue run failed golden_payload because the branch predated current main and its own surface.rs trust fix changes the surface digest. The recaptured snapshots differ ONLY in the surface sha256 lines (verified char-by-char) — no prompt text or capability-list changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…mpt-surface PRs (nearai#7371) * fix(ci): schedule the golden lane for prompt-surface production changes A production change to the model-visible prompt surface (the capability surface digest, the instruction bundle, the communication-context renderer, or a shipped loop-tier prompt asset) ran only crate buckets on the PR lane, so stale golden_payload snapshots surfaced first as a merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface digest; the PR lane never ran the golden bucket). Add a curated prompt-surface owner table that ADDITIONALLY schedules the golden integration lane without consuming the path's normal package classification. Self-tested per entry plus a negative control pinning that ordinary production changes keep the narrow plan. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): recapture the extension_host coverage floor with measured wobble tolerance The 2026-08-04 entry was an adjustment whose own text ordered the next real recapture. Since then code left the crate (denominator 24529 -> 24415) while the ratio ROSE to 88.12%, and the same commit df90072 measured >=21560 covered lines in the merge-queue lane but 21515 twice on the push lane — a >=45-line same-commit spread over a 20-line tolerance, redding main on noise (run 31208592262). Recapture both fields from that run's own gate output and size tolerance_lines to the measured cross-lane wobble. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): cover the host-managed prompt composer and derive per-entry tests Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives the InstructionBundleBuilder, model.rs shapes the pinned request) was missing from the prompt-surface table — the exact gap class the mapping exists to close. And the self-test enumerated entries by hand, so a new entry could ship untested. Add the host_managed_ports prefix and derive the positive cases from the tables themselves; an entry whose crate the fixture lacks now fails the suite explicitly instead of skipping. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…signal, builtin description trust, docs (nearai#7361) * fix(extensions): host-bundled description trust + already-connected install confirmation Two chat-side dead-ends from the 2026-08-07 Slack QA session (thread e79a994f, run 251aec0b on ironclaw-qa-testing-libsql): 1. Host-bundled capability descriptions were description_trust=Untrusted, so the loop-tier prompt-text denylist strict-scanned compiled-in text and silently omitted builtin.extension_register_hosted_mcp from every model prompt's capability surface ("browser authorization-code flow" matched the "authorization" credential pattern). HostBundled is the only source eligible for effective FirstParty/System trust, so its repo-authored descriptions now cross the verified-catalog boundary like signature/digest-verified registry installs. Untrusted provenance (InstalledLocal, UserRegistered, unknown) keeps the strict scan. 2. When install-driven activation passed the credential gate because the caller's declared requirements were all satisfied, the response never said so — the model got only conditional guidance ("If WebChat shows an account connection panel...") and deflected an explicit "connect account" request to the web interface even though the account was already connected. The install response now appends an explicit already-connected confirmation exactly when declared requirements were verified present for the calling user. Regression tests: manager surface test pins VerifiedCatalog trust for all model-visible lifecycle capabilities through the real host runtime; instruction-bundle tests pin retain/omit behavior for auth-vocabulary descriptions by trust; install-path tests pin the confirmation on the seeded-credential path and its absence for credential-free extensions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(channels): chat can drive the personal half of channel connect The onboarding and channels pages claimed "asking the agent to connect a channel doesn't work" and that the agent "may tell you it can't help". That describes only the operator half (registering app/bot credentials). The per-user half has shipped since early July: extension_install runs the same activation credential gate as the Channels card, raises the in-chat OAuth connection panel when the account is unconnected, and (as of the sibling fix) confirms when it is already connected. The self-knowledge protocol makes these pages the model's authority on IronClaw's own capabilities, so the stale claim scripted the exact refusal QA hit ("I can't initiate the Slack OAuth flow from here") on an account that was already connected. Correct both pages to distinguish the operator step from the chat-drivable personal connect. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(channels): align slack and telegram setup notes with the connect contract The slack page's operator-step note and the telegram troubleshooting accordion still carried the blanket "asking the agent to connect will not work" claim the overview/onboarding correction removed — same drift, different phrasing (review catch on nearai#7361, plus one more instance found by a broader sweep). Both now state the two-step contract: the operator half stays in the web interface; after it, chat drives the personal half (install/activate -> in-chat connection or pairing panel, or an already-connected confirmation). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(golden): recapture surface digests over the description-trust change The queue run failed golden_payload because the branch predated current main and its own surface.rs trust fix changes the surface digest. The recaptured snapshots differ ONLY in the surface sha256 lines (verified char-by-char) — no prompt text or capability-list changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…mpt-surface PRs (nearai#7371) * fix(ci): schedule the golden lane for prompt-surface production changes A production change to the model-visible prompt surface (the capability surface digest, the instruction bundle, the communication-context renderer, or a shipped loop-tier prompt asset) ran only crate buckets on the PR lane, so stale golden_payload snapshots surfaced first as a merge-queue bounce (nearai#7361, 2026-08-07: surface.rs changed the surface digest; the PR lane never ran the golden bucket). Add a curated prompt-surface owner table that ADDITIONALLY schedules the golden integration lane without consuming the path's normal package classification. Self-tested per entry plus a negative control pinning that ordinary production changes keep the narrow plan. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): recapture the extension_host coverage floor with measured wobble tolerance The 2026-08-04 entry was an adjustment whose own text ordered the next real recapture. Since then code left the crate (denominator 24529 -> 24415) while the ratio ROSE to 88.12%, and the same commit df90072 measured >=21560 covered lines in the merge-queue lane but 21515 twice on the push lane — a >=45-line same-commit spread over a 20-line tolerance, redding main on noise (run 31208592262). Recapture both fields from that run's own gate output and size tolerance_lines to the measured cross-lane wobble. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): cover the host-managed prompt composer and derive per-entry tests Review findings on nearai#7371: the host_managed_ports pair (prompt.rs drives the InstructionBundleBuilder, model.rs shapes the pinned request) was missing from the prompt-surface table — the exact gap class the mapping exists to close. And the self-test enumerated entries by hand, so a new entry could ship untested. Add the host_managed_ports prefix and derive the positive cases from the tables themselves; an entry whose crate the fixture lacks now fails the suite explicitly instead of skipping. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ead gates deleted (nearai#7373) * test(architecture): drop the dead ironclaw_storage row and arm the substrate list Gate-audit finding (open-and-shut): SUBSTRATE_CRATES in reborn_composition_boundaries.rs carried three rows of rot, all invisible because the loop's `let Some(..) else { continue }` silently skipped any entry that resolves to no workspace package: - "ironclaw_storage": no such package exists (verified against `cargo metadata --no-deps`; the only MISSING name of the 29 listed). - "ironclaw_approvals" and "ironclaw_assistant" were each listed twice. The silent skip is replaced with a panic naming the stale entry, so the list can no longer rot invisibly. Verified by sabotage: adding a bogus "ironclaw_zzz_probe" row now fails the test with "is listed in SUBSTRATE_CRATES but is not a workspace package"; the clean list passes (23/23). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): prune the dead sanctioned path from the specificity gate Gate-audit finding (open-and-shut): SANCTIONED_PATHS in reborn_extension_specificity.rs still exempted `extension_host/extension_installation_store.rs` — a file deleted by nearai#6430. No scanned path matches the fragment (verified with rg across crates/), so the entry exempted nothing; it is also the one exclusion surface in this gate with no staleness check, which is how it outlived its file. Full specificity suite green after removal (8/8). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): drop the v1 ironclaw_gateway/static exclusions from the telegram gates Gate-audit finding (open-and-shut): both cross-tree scans in telegram_extension_gates.rs still carved out `ironclaw_gateway/static` — the v1 monolith's embedded UI, whose crate was deleted with the src/ monolith (no crates/*/ironclaw_gateway directory exists). The exclusions matched nothing; scans now cover the whole tree with no dead carve-outs. Suite green after removal (12/12). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): make the dto-collapse gate's header describe the gate that exists Gate-audit finding (open-and-shut doc rot): the module doc still described the pre-nearai#6447 freeze design — a dangling doc-link to FROZEN_COLLAPSE_DTOS (renamed RETIRED_COLLAPSE_DTOS in nearai#6447), a promised delete-without-trimming failure and an empty-allowlist assertion that do not exist in the file, and a named owner for a collapse that completed. The mechanism itself is armed and untouched; the header now describes the permanent zero-gate it became, and records the two originally-frozen names that deliberately left governance (CapabilityOutcome via nearai#6299 deletion, CapabilityDispatchRequest blessed as the canonical port type). Suite green (2/2). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): repoint the manifest-reparse allowlist note at the colocated asset Gate-audit finding (open-and-shut doc rot): the BundledAsset allowlist entry's justification still cited include_str! of assets/memory_native/manifest.toml — a path retired when WS2 (nearai#7037) colocated packages; the live include in memory_native_extension.rs reaches crates/extensions/packages/memory-native/manifest.toml. Comment only; the gate's mechanism and counts are untouched. Suite green (2/2). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): give the memory-vocabulary gate the partial-tree floor its twin has Gate-audit finding: reborn_memory_retired_vocabulary.rs had no MIN_SCANNED_FILES floor, unlike its explicit twin reborn_retired_taxonomy.rs — so a partially-moved tree (the CHECKLIST WS0 / nearai#6963 'green while measuring nothing' shape) would scan a fraction of the files and still report the vocabulary clean. The gate was in fact born with an already-dead sanctioned path (its own header records this), so the rot class is not hypothetical for this file. Adds the same 500-file floor (real count ~4000), asserts it in the main gate, and pins the premise on a fixture: a 10-file partial tree scans clean and is rejected by the floor. Suite green (4/4); clippy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): close the transport gate's nested-use-group fail-open Gate-audit finding (sabotage-verified): product_symbols_in's braced-group branch closed at the FIRST '}' (group.find('}')), so a nested group — use ironclaw_assistant::{m::{X}}; — truncated mid-element and recorded zero symbols. Probed live before the fix: appending use ironclaw_assistant::{zzz_audit::{ZzzProbe}}; to webui's lib.rs left transports_name_only_the_frozen_residue_of_product_symbols GREEN, while the plain-path spelling of the same import correctly failed. The same truncation dropped qualified elements inside flat groups ({qualified_module::X} recorded nothing). The group branch now does a balanced-brace walk, splits elements at depth-0 commas only, and records a qualified/nested element's leading path segment — the same key the single-path branch records for ironclaw_assistant::module::X. Flat-element semantics are byte-for-byte unchanged, so the frozen 100-row webui inventory is untouched (suite green 6/6 on the live tree). Regression fixtures added to import_scanner_reads_symbols_out_of_real_use_shapes; the original sabotage now fails with the gate's own message (re-verified). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: delete check-e2e-matrix-files.sh — a gate for a workflow that no longer exists Gate-audit finding (provably inert): the script's default target is .github/workflows/e2e.yml, deleted when the v1 e2e suites were retired (git log --diff-filter=D shows the removing commit); no workflow, script, hook, doc, or guidance file references check-e2e-matrix-files.sh (verified with rg across the repo including .github and .githooks). A checker nothing runs, pointed at a file nothing provides, is dead weight that reads as coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: delete the measured-broken check-boundaries.sh and its guidance references Gate-audit finding (provably inert, previously measured): crates/AGENTS.md recorded on 2026-08-05 that the script fails on a clean tree (check 5 false-positives on live test files) and that checks 1/2/3/6 target the deleted v1 src/ tree, passing vacuously. No workflow or hook runs it; its only callers were guidance files, two of which claimed it 'enforces' root-tests feature gating — an enforcement claim the skill-maintainer rules forbid for a check nothing executes. Removed the script and every live reference: the crates/AGENTS.md warning row becomes a tombstone note; the testing skill + exemplar reference drop the false enforcement parenthetical; the architecture-review skill's Verify line drops the dead command; deslop-reborn's allowed-tools drops the permission; .coderabbit.yaml's driver-leak instruction now points at the live enforcement (reborn_persistence_driver_boundary). Two dated docs/internal/ plan snapshots keep their historical mentions. Verified: python3 scripts/ci/check-guidance.py OK (2084 path references) and its self-test OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(product): stop hardcoding charter sub-owner counts in the family map Gate-audit finding (stale prose): crates/product/AGENTS.md said '19-sub-owner reborn_services charter map' — the enforced map has had 20 sub-owners since nearai#7235 added the inspector row (counted from the live table). Rather than chase the number, drop both inline counts: the owning maps and their gates are authoritative, and the re-verify commands are already inline (skill-maintainer rule: no counts without a regeneration recipe). check-guidance.py OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): correct the scanner-fixture file's name-filter claim Gate-audit finding (doc rot with a false coverage claim): the header said naming the FILE reborn_* makes code_style.yml's 'cargo test -p ironclaw_architecture_tests reborn' see it — but that argument is a test-NAME filter (the measurement is documented in reborn_contracts_vendor_census.rs), and none of this file's test fns contains the substring, so that smoke lane runs 0 of them (11 collected by the full plan). Comment-only; the note now records the real semantics so file names are not trusted for lane coverage. Suite green (11/11). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(internal): gate & ratchet audit report + proposed preflight gauntlet The audit the owner asked for after PR nearai#7157 went red six times across four gates: every architecture-test gate, module charter, CI script, and committed baseline inventoried with a verdict and evidence; the handful worth acting on ranked by friction x weakness; the CI-ergonomics analysis (why failures surface one per ~1h round-trip: no --no-fail-fast anywhere in CI, cancel-in-progress on push, sequential fast-checks steps — measured: two broken gates report 1 failure in 18s under the CI shape vs both in 211s with --no-fail-fast); and the sabotage log for every probe. scripts/preflight-gates.sh is the concrete pre-push proposal: the deterministic-gate classes only (script gates ~10s + architecture suite --no-fail-fast + changed-crate charter tests), covering all four nearai#7157 gate classes locally in one command. Unwired — nothing invokes it. Validated end-to-end on this branch: exit 0, 'every deterministic gate green', 402.8s including gate-binary recompiles. Placement verified: python3 scripts/ci/docs_publication_boundary.py OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(planner): classify preflight-gates.sh and the deleted check-boundaries.sh The gate audit's own PR hit the planner's fail-closed arm — 'unmapped test or CI path: scripts/check-boundaries.sh' — exactly the class the arm exists to force a decision on (and the audit's report documents). Per the PR_STATIC_CONTROL_PATHS membership rule (no Reborn test lane exercises either file): - scripts/preflight-gates.sh — the audit's proposed local pre-push gauntlet; referenced by no workflow. - scripts/check-boundaries.sh — deleted by the audit; the entry lets the deletion diff (and any revert) classify instead of failing every downstream Reborn lane. Verified: the planner now produces mode=selected with the architecture-misc bucket for this branch's diff, and python3 scripts/ci/test_reborn_pr_test_plan.py is OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(internal): add the fold-tripped asymmetric-tolerance exhibit to the audit The strongest single exhibit for shortlist item 2, contributed by the nearai#7157 branch steward after this audit's cutoff and verified against the gate's code: TOLERANCE = 400 is consulted in exactly one direction (the banked-slack check, ceiling.saturating_sub(lines) > TOLERANCE); the growth check is a bare lines > ceiling. With the in-file 'set to current, not padded' instruction, every ceiling is a hard cap at the observed count — so one line landing on main in any contracts crate reds every open branch at its next fold until someone re-captures. Measured recurrence on nearai#7157: loop_contracts re-captured four times, ~once per fold (14,479 -> 13,850 -> 13,949 -> 13,115 -> 13,181), the last tripped by main's nearai#7361/nearai#7363 adding 66 lines to instruction_bundle.rs — nothing the branch wrote. All four deltas were <= 105 lines: either repair shape in §3.2 (one-line upward tolerance using the existing constant, or mid-window pinning) would have absorbed every one with zero red builds. This audit's own sabotage already proved the jaws (+1 line host_api red / -1 line common red); the fold history shows the operational cost. The repair stays a recommendation — adding growth headroom to a ratchet is the owner's call, not this PR's. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(architecture): give the contracts size ceiling upward working slack Owner-directed repair of the audit's sharpest finding (report §3.2): the gate's TOLERANCE = 400 was consulted in exactly one direction — the banked-slack check — while the growth check was a bare lines > ceiling. Combined with 'set to current, not padded' pins, every ceiling was a hard cap at the exact observed count, so one line landing on main in any contracts crate redded every open branch at its next fold until someone re-captured. Measured on nearai#7157: four loop_contracts re-captures, roughly once per fold, every delta <= 105 lines — the gate generating its own busywork. The growth check now allows GROWTH_TOLERANCE = 150 of working slack above each pin (sized to composition-budget precedent; the reviewed raises this gate has caught were +1,069 and +1,214 lines, far above it), and all six ceilings are re-pinned to the counts the test itself reported with every ceiling at 0 — which also removes the +400 seed padding on common/loop_contracts/prompt_envelope that contradicted the capture rule and put those crates one deleted line from the banked jaw. Sabotage-verified both ways: +1 line in host_api and -1 line in common — both red before this change — now pass; a +151-line probe still fails with the effective-ceiling arithmetic in the message. Full reborn_dependency_boundaries binary green (41/41); clippy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(budget): re-equalize composition pins to observed — restore the working window Owner-directed companion to the contracts-ceiling repair (same annoying class, other mass gate): merged main-side growth since the 2026-08-05 equalization had drifted +101 LOC and +5 Arc<dyn> sites through the tolerance windows, leaving 49 LOC / 10 sites of live headroom — the next routine composition PR would have gone red on wiring alone (the gate audit measured this the same day it was pinned). Per the TOML's own maintenance instructions: loc_ceiling/loc_observed 40423 -> 40524 and arc_dyn 814 -> 819, measured with the gate's --print, set to current not padded, dated notes appended (not overwritten), and the arch-test record (COMPOSITION_ABSOLUTE_SRC_LOC) moved in the same commit as its file requires. ceiling_bp stays 658 — the WS0 floor is deliberately not re-set. Verified: check-composition-budget.sh OK; its 76-case self-test green; reborn_restructure_baselines green; probe +100 LOC now passes (was red at 49 headroom), probe +160 LOC still fails. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(internal): record the landed zero-slack repairs in the audit report The §3.2 repair moved from recommendation to landed at owner direction; the report's answer, inventory rows, and §7 ledger now say so, with the counting-rule fix promoted to the top remaining recommendation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * gates: pin the ceiling-window arithmetic; fail preflight discovery closed Two review-round hardenings (the open CodeRabbit Majors): - reborn_dependency_boundaries.rs: extract the size-ceiling comparison into contracts_ceiling_verdict() and pin its four window edges with a committed regression test (contracts_size_ceiling_window_edges_hold) — accept at ceiling+GROWTH_TOLERANCE, reject one line past, accept at ceiling-TOLERANCE, reject one banked line further, and a zero-measure scan reads Banked, never a silent pass. The pre-repair asymmetry (tolerance consulted only downward) can no longer return silently. Live-gate behavior re-probed unchanged after the rewiring: +1 line to host_api passes, +151 fails with the same effective-ceiling message. - preflight-gates.sh: setup and changed-file discovery now fail closed — a missing repo root exits 2, and a failed merge-base/diff widens the charter run to all five crates instead of silently skipping them (the same fallback the missing-base branch already used). A broken setup may cost compile time, never a silent skip. Full boundary binary 42/42 green; clippy clean; preflight-gates.sh end-to-end green on this tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
The bug: asking the agent to connect Slack from chat dead-ends. In the 2026-08-07 QA session (thread
e79a994f, run251aec0bonironclaw-qa-testing-libsql, deployed commit389aa99) the tester typed "connect account" and the model answered "I can't initiate the Slack OAuth flow from here directly — that needs to be done through the web interface" — even though the trace shows the model had already run the designed path (extension_search→extension_install(slack)), the activation credential gate had passed, and the tester's personal Slack account was therefore already connected. The truthful answer was "you're already connected"; instead the user spent ~29 minutes hunting the web UI. The repo's own QA spec requires this flow to work (qa_3a_slack_connect,scripts/reborn_webui_v2_live_qa/run_live_qa.py:144: ask IronClaw "connect to Slack", go through the flow, expected: connected).Three causes, three fixes — all in the "model is confused about extension connect" cluster:
docs/onboard.mdx,docs/channels/overview.mdx). They still said "Asking the agent to connect a channel doesn't work… It may tell you it can't help; do it yourself in the web interface" — stale since the chat-side connect gate shipped in early July, and load-bearing becauseself_knowledge.md(injected every turn) makes docs.ironclaw.com the model's authority on IronClaw's own capabilities. The pages conflated the operator half (registering app/bot credentials — genuinely web-only) with the per-user OAuth half thatextension_installdrives through the same activation credential gate as the Channels card. Both pages now distinguish the two halves.ironclaw_extension_manager). When install-driven activation passes because the caller's declared credential requirements are all satisfied, the response carried only conditional guidance ("If WebChat shows an account connection panel… If the user's account is already connected, continue") — the gate's verdict was never surfaced, so the model couldn't tell which branch was true. The install response now appends an explicit already-connected confirmation exactly when declared requirements were verified present for the calling user, and never for credential-free extensions.ironclaw_host_runtime::surface). Host-bundled capability descriptions carrieddescription_trust = Untrusted, so the loop-tier prompt-text denylist strict-scanned compiled-in text and droppedbuiltin.extension_register_hosted_mcpfrom the prompt's capability surface on every turn (WARN each turn:"browser authorization-code flow"matched the"authorization"credential pattern). This is the same confusion surface — the disclosure protocol tells the model to use the lifecycle tools for connect/install requests, while the scanner was quietly deleting one of them from its instructions, and any future connect-flow capability whose description mentions OAuth would meet the same fate.HostBundled— the only source eligible for effective FirstParty/System trust — now maps toVerifiedCataloglike signature/digest-verified registry installs;InstalledLocal/UserRegistered/unknown keep the strict scan.Change Type
Linked Issue
None — diagnosed from the QA "Slack issues" report (thread
e79a994f-55bd-56a8-87c3-9c9032405be2, run artifact251aec0b, Railwayironclaw-rebornneari/qa-testing-libsqltrace logs).Validation
cargo fmt --all -- --checkcargo clippy -p ironclaw_loop_contracts -p ironclaw_host_runtime -p ironclaw_extension_manager --all-targets --all-features(clean; workspace-D warningslane in CI)cargo build(via test builds)ironclaw_host_runtime(446+37),ironclaw_extension_manager(141+2),ironclaw_loop_contracts(new pins),ironclaw_architecture_tests(55 suites, 0 failures),reborn_integration_extension_runtime(22 pass; the 2case_2Postgres legs require a local Docker daemon — not running on this machine — and fail identically without this diff; CI covers them)review-pr/pr-shepherdTest Strategy
User behavior: asking the agent to connect an already-connected extension in chat gets a truthful "already connected — continue" instead of a deflection to the web interface;
extension_register_hosted_mcpappears in the prompt's capability surface; the docs the model self-grounds on describe what chat can actually do.Risk areas:
Tests added or updated:
instruction_bundleretain/omit pins for auth-vocabulary descriptions by trust level (loop_contracts);install_response_confirms_connection_only_when_caller_credentials_were_verified(manager)reborn_integration_extension_runtimeexercised (no assertions changed);standalone_agent_surface_exposes_extension_lifecycle_toolsextended to pinVerifiedCatalogtrust for all model-visible lifecycle capabilities through the real host runtime;standalone_extension_activate_hosted_mcp_stages_discovery_and_publishes_toolsextended to pin the already-connected confirmation on the seeded-credential path; credential-free install pinned to NOT claim a connectiontests/fixtures/for the changed strings; none pinnedchannel_connection_required) is untouchedWhat the tests prove: each fix fails without its change — the trust assertions were RED (
Untrusted) before the surface mapping change, and the connect confirmation was RED (message absent) before the install-arm change; the untrusted-provenance strict scan is pinned so the safety posture cannot silently widen.Commands run:
cargo test -p ironclaw_host_runtime,cargo test -p ironclaw_extension_manager,cargo test -p ironclaw_loop_contracts --lib,cargo test -p ironclaw_architecture_tests,cargo test -p ironclaw_integration_tests --test reborn_integration_extension_runtime,python3 scripts/ci/docs_publication_boundary.py,cargo fmt.Compatibility, rollback, follow-ups
Compatibility. The trust widening is scoped to
HostBundled(compiled/shipped with the binary — descriptions are repo-authored constants; structural prompt limits still apply on every variant, and untrusted provenances keep the full denylist). The install message change is additive prose on the model-visible result; no wire schema change. Docs regeneratellms.txton deploy, which also updates what the self-knowledge protocol serves the model.Rollback. Each commit reverts independently; no persistence, schema, or config changes.
Follow-ups (out of scope, from the same diagnosis).
stream_options.include_usagenever sent) — fix is written and verified against the live endpoint, staged separately from this PR.logs: {available: true, entries: []}— the operator in-memory tracing ring loses run lines within seconds of completion.provider_id=unconfiguredon the NEAR AI backend.github,mcp-*) hittingOAuth gate vendor unavailable; falling through—qa_4b_github_connectwill reproduce the same complaint class.extension_activatecontinuation arm.qa_3achat-connect journey on the QA deployment once this deploys.🤖 Generated with Claude Code