Guidance layer: a family AGENTS.md for every family, a README for every crate, and a repo-wide stale sweep - #7264
Merged
Merged
Conversation
…ent (D-S) and re-walk the WS9 verify row Appends §12.13 D-S under delegated authority at owner direction, flagged for post-hoc review by Illia Polosukhin (#6696's author): the await-edge store is measured to be a pure projection over ProcessDependencyPort (that half of the shed happened inside #6696 itself), and the resolver is a genuine loop-tier responsibility journal edges cannot express (owner recovery, sanitized transcript result materialization, batch-gate resume-once drain, BlockedDependentRunGate resume policy). §6.7.3 is amended (scheduler DONE / store DONE / resolver KEEP) instead of the shed being executed; the 2.9k figure is corrected to 1,459 production + 1,448 cfg(test) lines. The §12.10 bullet, §2 divergence flag, §9 row 49, §13 validation row, CHECKLIST header/WS4 pointer, README and PLAN all carry the dated resolution. WS9 verify row ticked with evidence: one lifecycle authority (the process journal; TurnRunState/TurnRunRecord are projections via AgentTurnProcessRuntime, ProcessRecord is a capability-invocation view, no bare RunRecord exists) and §7 T4 re-walked clause-by-clause against merged code — matches, including the checkpoint-gated no-auto-retry mechanism (BeforeModel precedes ModelStage; requeue only when checkpoint-free under the 3-claim cap). Docs-only; no code, no tests, no gates touched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…dependent rederivation) and the 74-row §9 mapping audit (45 L / 15 L-A / 14 OBD / 0 NOT-LANDED; 3 findings recorded) Row 1: check-target-tree.py reports 64 workspace members == 64 documented packages, 1 documented exclusion (tools/ironclaw_silk_decoder), 0 owned exceptions (EXCEPTIONS table empty — §5 steady state); self-test 17/17; cargo-metadata name set diffed empty against an independent §5 parse. Row 2: docs/reborn/target-architecture/ws12-mapping-audit.md is the audit record — per-row executed-evidence, delete-clauses read against WS8's execution notes, all 14 open rows cite their owning CHECKLIST/PROPOSAL row or issue. Findings (recorded, not fixed): F1 prompt_envelope manifest-description fix has no owner row; F2 WS6:429's '#5618 residue deleted' overstates vs the live adopt_migrated_identity + open WS8:523; F3 stale-docs cluster where the tree is ahead of the prose (trace re-export drop, TurnRunTransitionPort, processes->resources). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… resolver retained loop-tier; WS9 rows resolved
… wiring choice, and regression pins (PROPOSAL §12.13, 2026-08-05) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…scan traversal (+node_modules), bounded sidecar output draining (capped capture + discard drain), deadlock regression asserts successful redaction (no seq), XLSX/DOCX empty-classification via extract_document, raise_for_status annotations Threads already addressed by the fold: latency.rs caller-contract wording (merged doc scopes the requirement to latency-trace callers), BodyJsonPointer coverage (the plaintext-refusal test drives all four injection shapes). Deliberately not taken: un-xfailing the four Slack-catalog projections — the xfail is a documented tripwire (unexpected-pass goes red) and clearing them is the #6520 projection-modeling follow-on its comment specs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…gate caught — main's relocated triggered_run_delivery_services carried the drift the PR fixed at its old channel_host address Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… ruling (D-R), typed quarantine flag, planner entrypoint classification, review-thread triage # Conflicts: # docs/reborn/target-architecture/PROPOSAL.md
…manifest description (F1), dated ✎ corrections for the #5618 overstatement (F2) and the stale-prose cluster (F3) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e-verification (rows 5-6) Adversarial second-reviewer pass over PROPOSAL §12.1a/b/c and the batch's own §12.13 D-R loopback carve-out, plus a re-run of the five extension journeys. Attacks were executed rather than argued: two sabotage files and a 38-shape hostile-URL probe were planted, run, and reverted. Verdicts — mint consolidation HOLDS-WITH-RESIDUAL, secrets tightening HOLDS-WITH-RESIDUAL, host/verifier colocation HOLDS, D-R HOLDS. No HOLE. Four findings recorded rather than fixed (report-not-repair): - F1 test_verified/_for_tenant are ungranted mint constructors gated only by the `test-support` feature, in no mint-name table, with nothing pinning the feature to [dev-dependencies]; the shipped binary is measured feature-free. - F2 §12.1b's products-layer residue undercounts by one (ironclaw_assistant). - F3 journey coverage hole: gsuite-with-credential-injection is proven in two halves that no committed test joins. - F4 both recorded census evasions and both fail-open reads are CLOSED on this tree, so §11.2.5/§12.1a/CHECKLIST:552/:597 now understate the seal. Rows 5-6 ticked; only lines 631-632 of CHECKLIST touched so the concurrent rows 3-4 edit folds cleanly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rity audit (0 holes, 2 residuals, 1 journey gap recorded)
Dispatch ceiling 1122 -> 814 (today's observed, nudge taken; WS0 record 827 stays within effective 829). Mass-share ceiling 2398 -> 658 bp (the WS0 baseline floor — the arch-test assert refuses lower, and observed 578 bp sits inside the nudge window). Absolute LOC re-equalized at 40423: #6831 added 4 governed LOC through the queue's tolerance window; ceiling, observed, and COMPOSITION_ABSOLUTE_SRC_LOC move together here. Both tightenings sabotage-verified red (dispatch 9-over at 790; abs 73-over at 40200). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…AL in scope), row 4 verified-but-open on two pre-existing Postgres-leg test-isolation defects WS12 rows 3-4 verification on the assembled batch tip 0c6c0cf: Row 3 (ticked): fmt, clippy default/all-features/--lib --bins, workspace tests (495 targets, 15,203 passed, 0 failed; the smoke.rs:3132 CPU-saturation flake passed first try), arch suite 285/0, the integration-feature lane 1,665/0, recorded-fixture QA (61 fixtures clean, 41/0), frontend (typecheck 1,588 files; vitest 1,088/0; build + bundle budgets), e2e smoke = the CI browser lane under the hermetic wrapper (50 + 21 + 5 passed), and all 41 scripts/ci self-tests (two mapfile/bash-3.2 casualties green under bash 5, the CI shape). Row 4 (stays open, dated note added): both-backend parity proven with legs demonstrably executed for the fabric (57 pg + 81 libsql), triggers (ADR 0003, REQUIRE_POSTGRES), hooks (ADR 0004, all three backends), composition, processes journal, extension-registry, host-runtime libSQL restart, and the backend matrix; fabric-delegated domains enumerated. Two REAL blockers (one class): the Postgres legs of the event-store and assistant-ledger contract suites assert against shared-database state and cannot pass as-written (each failing test passes alone on a virgin database; files byte-identical to origin/main; no CI lane sets their env vars). Full evidence: docs/reborn/target-architecture/ws12-gauntlet-report.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ty 8/10 with the two pre-existing Postgres isolation defects recorded # Conflicts: # docs/reborn/target-architecture/CHECKLIST.md
The base commit for the family-guidance program: one canonical home per fact, measured-not-aspirational claims, boundaries stated as exclusions, and the note that guidance files can be gate-pinned. Every family/crate document written on top of this branch follows this shape. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ity-blocking contract suites The WS12 gauntlet (ws12-gauntlet-report.md §P6/§P8) measured the Postgres legs of ironclaw_event_store's durable_event_store_contract and ironclaw_assistant's durable_ledger_contract as test-isolation-defective: absolute database-global asserts (event cursors; settled-entry prune bookkeeping) run against the single external database named by their IRONCLAW_*_POSTGRES_URL env vars. Every failing test passes alone on a virgin database - store semantics correct, suites not self-isolating (PROPOSAL §12.13 D-T). Fix: each affected test provisions a private database on the configured server - the fabric contract's IsolatedDatabase pattern (db_root_filesystem_contract.rs) ported locally into each suite: CREATE DATABASE per test, store/pool + migrations against it, courtesy DROP ... WITH (FORCE), and a once-per-binary stale-name sweep. Every assertion preserved byte-identical; libsql/jsonl twins untouched. In the ledger suite only the two retention tests move - the other six Postgres tests keep their proven fingerprint-suffix isolation. Regression pins are the fixed tests themselves: - postgres_replay_advances_next_cursor_past_trailing_filtered_records - postgres_runtime_and_audit_logs_survive_rebuild_with_filtered_cursor_semantics - postgres_settled_entry_limit_prunes_oldest_when_configured - postgres_settled_prune_interval_defers_until_interval_when_configured Green proven on a shared dirty database twice in a row (parallel default threading) and serially on a virgin database; red-first reproduction captured before the fix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…lose CHECKLIST WS12 row 4 D-T (after D-S): the WS12 gauntlet's two REAL findings were one defect class — absolute database-global asserts against the single shared env-var Postgres database — in two suites (event store cursor contract, assistant settled-ledger retention). Ruling executed in commit 864d93e: per-test isolated databases via the fabric contract's IsolatedDatabase pattern, assertions preserved; alternatives (baseline-relative asserts, serial-only, leave-open) recorded with why they lost; regression pin = the four fixed tests themselves. CHECKLIST WS12 backend-parity row ticks [x] with a dated addendum: red-first reproduction, the three green isolation runs (dirty shared DB twice in parallel; failing pairs serial on virgin), parity now green 10/10. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…idance-conventions shape, READMEs for all 4 family crates and 14 packages, duplicate-guidance consolidation The family AGENTS.md now teaches the unified extension model (extension = the only product object; channel/tool/auth are manifest surfaces; runtime is loading, never taxonomy; ExtensionId vs VendorId; retired vocabulary pinned by reborn_retired_taxonomy.rs), carries the self-containment and package-to-crate rules from families/extensions.md, the four-responsibility lookup, the measured package catalog, the exclusion list, and the armed gates by test name. Every crate and package gains a README.md (ironclaw_extension_host had no guidance of any kind). ironclaw_extension_registry and memory-native each had both an AGENTS.md and a CLAUDE.md saying overlapping things: AGENTS.md is now canonical, CLAUDE.md a pointer, and memory-native's stale v1 references (src/workspace, src/db/libsql) are dropped in the merge. The slack/telegram agent maps get package framing and a contracts-tier pointer in place of the stale ironclaw_assistant one. Every path literal verified to resolve on disk; all figures (tool counts, dep sets, consumers, layer declarations) measured from the tree at 8d13454. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ntions.md
Family AGENTS.md rewritten to the spec shape for crates/substrates/ and
crates/lanes/: boundary, crate table, exclusion lists (mechanism-not-authority
for substrates; kernel-decides-lane-executes for lanes), armed gates by test
name, and measured deviations stated as deviations (sandbox's three substrate
deps, script.rs direct spawn). The lanes wit/-is-load-bearing note is kept.
A README.md for every crate in both families (10 new), measured against
cargo metadata 2026-08-05: public surface, workspace edges, consumer counts,
and enforced invariants each citing their gate. ironclaw_libsql_runtime and
ironclaw_wasm_limiter previously had no guidance of any kind; their READMEs
carry the sole-pool-home rule (ADDITIONAL_DRIVER_ALLOWLISTS: deadpool =
{filesystem, libsql_runtime}) and the outbound-only limiter gate
(wasm_sandbox_core_module_stays_domain_free_v1_parity_kernel; no BoundaryRule
names the limiter).
Duplicate guidance consolidated per rule 1: for the six crates holding both
AGENTS.md and CLAUDE.md (filesystem, network, secrets, mcp, sandbox, wasm),
CLAUDE.md stays canonical (module spec for filesystem; gate-pinned wording for
mcp and wasm) and AGENTS.md becomes a short pointer. No gate-pinned file was
edited. Stale reference removed: safety AGENTS.md pointed at
src/NETWORK_SECURITY.md, which exists nowhere in the tree.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… the restructure crates/AGENTS.md (264 -> 175 lines): routing map only — the ten families, the read order (family AGENTS.md -> crate README.md -> working rules/module spec -> docs/reborn/contracts/), the enforced seven-layer matrix with the family/layer divergences, measured workspace facts (64 packages, 1 documented exclusion, 0 owned exceptions per scripts/ci/check-target-tree.py), and a verified command block. The 40-row per-crate map is gone: family AGENTS.md files own crate routing per docs/reborn/guidance-conventions.md. crates/README.md (141 -> 119 lines): human map — mental model in family vocabulary, the ten families with measured crate counts, the 14 extension packages (4 crates + 10 data-only), and the two workspace members outside crates/. crates/Architecture.md (1019 -> 1059 lines): audited against the live tree; every named symbol/path re-verified 2026-08-05. Corrected: retired ProductAdapter vocabulary (zero residue in code), the stale pre-rename dependency ladder that still cited the deleted gateway/TUI crates, run-state store mentions, lane-table crate anchors (sandbox/extension_support), declared-in vs minted-by owners in the core data model, and the subagent deny-filter status note (re-verified). Marked the pre-restructure 'partial or evolving' list as unmeasured rather than asserting it. Also documents that scripts/check-boundaries.sh fails on a clean tree (check-5 grep false positives) and greps the deleted v1 src/ in 4 of 6 checks — boundary enforcement for crates/ is the architecture suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…dinator sync with extensions-family agent) Every extensions package dir — the 10 data-only ones included — now ships a README.md, so both maps extend the read order to package level. The sibling branch also confirmed what this map already derived per-crate: packages/ is not uniformly products-layer (memory-native and mem0 declare substrates). The other two coordinator corrections targeted rows of the old per-crate map, which this rewrite deleted wholesale; nothing here cites memory-native's CLAUDE.md or claims ironclaw_extension_host lacks guidance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…orn/guidance-conventions.md
- crates/contracts/AGENTS.md and crates/events/AGENTS.md rewritten to the
family shape: exclusion lists with destinations, armed gates by test name,
layer-matrix rows, crossing guide, measured header counts.
- README.md added for all 10 crates (ironclaw_prompt_envelope previously had
no guidance of any kind — the CHECKLIST WS11 gap).
- One canonical guidance file per crate, other file a pointer:
A+C merges for ironclaw_host_api, ironclaw_event_log,
ironclaw_event_projections, ironclaw_event_streams; CLAUDE-only content
moved to AGENTS.md for ironclaw_loop_contracts,
ironclaw_extension_contracts, ironclaw_product_contracts (none of these are
root module-spec crates, so AGENTS.md is the working-rules home).
- Stale guidance fixed against the live tree:
* loop_contracts dep list contradicted the enforced allowlist (manifest is
host_api + extension_contracts; common/prompt_envelope are permitted,
unused).
* event_log still documented the deleted jsonl parse/replay helpers.
* event_projections still claimed EventStreamManager,
DurableMemoryAuditSink, MemoryAuditProjectionMetadata, and
PendingGateProjection — all deleted per PROPOSAL 6.3.3.
* product_contracts still carried the pre-D-E open vendor decision under
the nonexistent module name llm_config, and a Deferred section
contradicting its own operator_llm/operator_service rows.
* extension_contracts module table was missing the WS3 runtime module
while counting 18.
* common's llm_costs note carried the ModelCostTable seam claim refuted by
PROPOSAL 12.11 D-F; now cites the pricer-port ruling and the vendor
census residue.
- Deleted crates/events/ironclaw_event_projections/PENDING_GATE_PROJECTION.md:
every claim in it referenced deleted symbols or the removed v1 src/ tree,
and its only inbound reference was the crate's own CLAUDE.md.
Verified: all consumer counts reproduce via the printed grep commands; 147
path literals across the 28 touched files resolve on disk; no architecture
test reads any of these files by name; conflict-marker scan clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e program memory packages are substrates-layer, not products (families/extensions.md); memory_native declares no extension_contracts dep (PROPOSAL §6.8.4); wasm's extension_contracts edge is dev-only and the wasm 'never depends on' bullet is lane-scoped, not family-wide (families/lanes.md). Three further reported defects were checked and NOT corrected — they were misreads: the sandbox 'never above the runtime tier' rule holds (substrates sit below it), and PROPOSAL's safety consumer count already reads 17, matching the tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… READMEs, AGENTS/CLAUDE consolidation Family-guidance program, kernel family (guidance-conventions.md shape): - crates/kernel/AGENTS.md rewritten to the family shape: the nine-stage effect pipeline with stage ownership, the sealed-mint table (witness / trust ceiling / approval lease / verified-inbound evidence, each with its mint site and its seal mechanism), the per-stage fail-closed table with file:line or test citations, the sharp exclusion list, and the armed gates by test name (authorized-seal ratchet, sealed-evidence mint ratchet, BoundaryRules, same-layer edge inventory at 21 kernel edges, empty LAYER_MATRIX_EXCEPTIONS register, driver boundary, process storage scan, origin-gate matrix ratchet). - A README.md for each of the nine crates, per the crate shape: measured workspace deps and consumer counts (cargo metadata), public surface with verified citations, enforced invariants naming their gates. ironclaw_processes states the single-lifecycle-authority direction of truth (journal = store; TurnRunState/ProcessRecord/await-edge = projections; PROPOSAL §12.13 D-S); ironclaw_host_runtime documents the D-R literal-loopback carve-out and names its two regression tests. - Duplicate guidance reconciled in all nine crates: AGENTS.md is canonical (guardrails absorbed), CLAUDE.md reduced to a pointer; ironclaw_trust's CONTRACT.md untouched as the co-located cross-crate contract. - Stale references fixed inside owned paths: the deleted capability-profile conformance module (evaluate_profile_conformance — zero hits workspace-wide) removed from ironclaw_capabilities guidance; trust's 'staging branch' / 'PR3' phrasing updated; capabilities' 'later obligation slices' updated to the landed host_runtime obligations split; cross-crate path mentions fully qualified. Every path literal in all 29 kernel .md files verified to resolve on disk; every named symbol swept against crate sources. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…te READMEs, duplicate-guidance consolidation, stale-path fixes Family guidance for crates/domains/ per docs/reborn/guidance-conventions.md: - crates/domains/AGENTS.md rewritten to the family shape: charter table with go-here-when routing, the exclusion list, every armed gate named by test (BoundaryRules + identity/memory allowlists, the 5-entry in-family edge inventory, the naming gates, trusted-trigger ownership, the memory-provider residue ledger, persistence-driver boundary, the two module-charter gates). - A measured README.md for each of the 12 crates: charter, use-when / don't-use-when routing, public surface, measured normal deps + named consumers, enforced invariants with their gates, exact test commands. ironclaw_attachments and ironclaw_identity had no guidance of any kind; identity's README points at CONTRACT.md (the module spec), llm's at its CLAUDE.md module spec. - Duplicate guidance consolidated to one canonical file + pointer per crate: threads/conversations/memory/outbound rules now live in AGENTS.md (CLAUDE.md is a pointer); auth/llm keep CLAUDE.md canonical because their tests/module_charter.rs gates read it (AGENTS.md is the pointer). One misstatement fixed in the conversations merge: transcript content belongs to ironclaw_threads' SessionThreadService, not InboundConversationService. - Staleness fixed inside the family: identity CONTRACT.md two-edge allowlist claim reconciled with D-Q's three entries; trace_commons CLAUDE.md gains the capture module row and strikes its two discharged Known Gaps (recording/paths shims deleted, rename done); llm CLAUDE.md reasoning.rs caller corrected to crates/loop/ironclaw_loop_host; triggers lib.rs 'feature-gated' repo doc comments corrected; pre-family path literals in comments repointed (kernel/approvals+processes, loop/hooks, app/architecture_tests, domains/auth) and the deleted-v1-engine references in skills marked historical. Verified: cargo test -p ironclaw_llm --no-fail-fast (922 passed, exit 0 — CLAUDE.md is gate-pinned); cargo check --all-targets on all six crates with source edits; every cited path literal resolves on disk; conflict-marker scan clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… family laws kernel.md: ironclaw_authorization's 'Security & authority role' bullet has been textually corrupted since #6918 — an approvals sentence was spliced into it mid-clause, orphaning its continuation line. Reconstructed, with the spliced sentence restored to the approvals entry where it is true. lanes.md: 'a lane never depends on a substrate' is false as a family-wide law (ironclaw_sandbox holds network/safety/secrets normal deps, which its own entry licenses); the accurate law is the layer ladder, and the narrow claim holds for ironclaw_wasm alone. lanes.md + events.md: the 'every crate ships both an AGENTS.md and a CLAUDE.md' requirement is superseded by docs/reborn/guidance-conventions.md — two files restating one rule is the drift the guidance program removes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two new gates in reborn_sealed_evidence_mint_ratchet (closed paths #12/#13), per the audit's remedy spec: (a) TEST_SEAM_MINT_FNS governs test_verified/test_verified_for_tenant — any production-text call site outside ironclaw_host_api is an offender (comments/strings stripped, #[cfg(test)] blocks stripped, tests.rs / *_tests.rs and cfg-test-only files excluded via the shared census); (b) test-support may appear in no normal dependency table workspace-wide (dependencies / build-dependencies / target.* variants / workspace.dependencies), and no [features] key other than test-support may forward to it — the laundering shape that would evade (b) by one rename. [dev-dependencies] enablement stays legal (cargo-features.md bar 4, the sanctioned dev seam). Measured zero offenders on this tree in both directions before pinning; sabotage-proven red->green both ways (planted production call named with file:line-text; [dependencies] enablement named with its table path). Self-tests drive the same pipelines the gates run (zero-match principle); the definition-location and partition tests now cover the new table. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
serrrfirat
approved these changes
Aug 6, 2026
BenKurrek
added a commit
that referenced
this pull request
Aug 6, 2026
…ls, delivery heuristics deleted Re-landed PR #7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as #6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from #7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
7 tasks done
thisisjoshford
added a commit
that referenced
this pull request
Aug 6, 2026
Main's #7264 (guidance-layer sweep) rewrote three files this branch also edits. Resolutions, all take-theirs-then-reapply-fixes: * .claude/skills/ironclaw-reborn-orientation/SKILL.md: kept main's updated prompt-templates bullet (new family-depth glob and crate list) and this branch's internal-docs placement bullet beside it. * crates/loop/ironclaw_hooks/CLAUDE.md and docs/reborn/target-architecture/CHECKLIST.md: took main's newer content wholesale, then re-ran the path sweep (docs/plans, docs/adr text refs) and the relative ADR-link fixes over it. Post-merge verification: relative-link scanner reports no new broken links, boundary checker green, 21 + 52 self-tests green, cargo fmt clean, Reborn PR test planner green.
BenKurrek
added a commit
that referenced
this pull request
Aug 6, 2026
…ls, delivery heuristics deleted Re-landed PR #7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as #6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from #7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; #7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BenKurrek
added a commit
that referenced
this pull request
Aug 7, 2026
…ls, delivery heuristics deleted Re-landed PR #7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as #6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from #7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; #7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
10 of 11 tasks
personal-upstream-sync Bot
pushed a commit
to theredspoon/ironclaw
that referenced
this pull request
Aug 7, 2026
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
pull Bot
pushed a commit
to Stars1233/ironclaw
that referenced
this pull request
Aug 8, 2026
nearai#7157 follow-ups) (nearai#7377) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key observer gate notices by their gate ref One run can park on several approval/auth gates in sequence (the observer's blocked-state loop re-announces whenever the (status, gate) marker changes), but the live observer derived every gate notice's projection id with no discriminator, so all ApprovalNeeded notices in one run collapsed to a single durable delivery identity. The second gate's prompt came back AlreadyDelivered from the coordinator, was treated as success, and was never sent — the user was never told about the gate their run was parked on, and no reply route was recorded for it, so a bare `approve` could not resolve it either. Key the projection id by the notification's gate ref (the mechanism nearai#7157 added for the triggered notifier's RunBlocked notices). A repeat announcement of the SAME gate still dedupes; kinds that carry no gate ref (FinalReplyReady) keep the historical undiscriminated id shape so existing delivery identities are not re-keyed. Regression: observer_delivers_a_prompt_for_each_distinct_approval_gate drives the real DeliveryCoordinator over the real outbound store through two distinct scripted gates and asserts two delivered prompts plus a recorded reply route for each. Sabotage-verified: reverting the discriminator to None fails exactly this test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(composition): pin the notification-channels gate dance when owner ≠ actor The full builtin.notification_channels_set approval dance — raise, replay payload, user approve (store + lease mint from the stored row), approved resume, lease claim, dispatch, lease consume — driven through the real capability port on a run whose thread owner differs from its acting user. Pins two properties ahead of unifying the scope derivation onto the actor: raise and resume must derive the same scope (every store in the dance is scope-keyed, so a half-unified derivation strands the approved capability), and whose identity that scope carries (the thread owner, under the interim split nearai#7157 shipped). The approve step mints the lease from the stored request's own scope, grantee, and fingerprint — the same material the production click-approval resolution uses — never a re-derivation. Capability-host tier rather than tests/integration because the product rule "a run acts as its invoker" makes owner ≠ actor unconstructible through every product front door; the run-context shape remains legal kernel state (runs parked across the deploy boundary carry it). The owner == actor dance stays covered end-to-end at the integration tier (outbound_target.rs). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(outbound): scope the whole notification-channels gate dance as the acting user Unify the interim nearai#7157 split: resource_scope_for_run and settings_scope_for_run now derive from the acting user like caller_for_run already did, and effective_user_id (the owner-first ladder) is deleted. The approval-gate raise, the replay payload, the durable gate record, the lease, and the approval-settings read all follow the user who invoked the run — so the invoker sees and approves the gate, and their settings govern it. The raise/resume coverage added one commit earlier ran before and after this change and caught a real half-unification in between: the resume-side replay load lives in ironclaw_loop_host's synthetic-capability wrap (a different crate from the raise-side save in notification_channels_set) and still derived owner-first, stranding an approved resume with "replay payload is unavailable". The acting-identity ladder now has exactly one definition — LoopRunContext::acting_user_id in ironclaw_loop_contracts — and both sides delegate to it, so the hand-synced-copy class is closed rather than re-synced. notification_channels_set's replay/gate-record writes move from the capability_host-wide owner-first helper onto the outbound module's base_resource_scope_for_run so every store in one dance derives one user; the capability_host-wide helper itself is unchanged (thread/durable-result scoping legitimately follows thread ownership, and owner == actor on every binding created under the run-acts-as-invoker rule). Loop-contracts size ceiling: +16 lines for the shared ladder, paid for by splitting host/run_context.rs's 104-line inline #[cfg(test)] module into its run_context/tests.rs sibling; ceiling re-captured DOWN 13_115 -> 13_028 from the gate's own failure message. Runs raised before this change with owner != actor and resumed after it will miss their replay payload once and fail closed; re-requesting approval recovers. Documented in the PR body. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(conversations): key shared-route bindings per (conversation, actor) A run acts as the user who invoked it, so a shared conversation binds one thread per paired actor — each owned by that actor — instead of one conversation-wide thread owned by a configured subject. BindingKey gains a serde-defaulted shared_actor_user_id component (None for Direct routes, whose identity stays the conversation alone); the trusted-owner parameter is deliberately ignored on Shared creates (it remains the trigger lane's way to bind Direct conversations for their creator), and the legacy shared-owner backfill is removed with it. Migration is ignore-but-retain, pinned with a restart-path test: legacy Direct keys deserialize byte-identically (continuity), while legacy conversation-keyed shared rows deserialize to a key no per-actor lookup builds — retained in durable state untouched, and every participant (including the old subject) starts a fresh thread they own. Morphed legacy pins record what became structural: a shared probe/lookup can no longer address (or widen) a Direct binding at all; stored reply targets are isolated per actor; an actor's unpair cannot take the conversation away from other participants' own threads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(product)!: remove shared-route subject binding; scope = invoker Owner ruling: a run acts as the user who invoked it, in a DM and in a shared channel alike, with one thread per (conversation, user). This removes the subject half of shared-route configuration end to end and keeps the admission half, fail-closed: - ironclaw_product_contracts: subject_route becomes shared_admission — the SharedConversationAdmission port answers only "is this shared conversation connected"; ProductConversationRouteKey survives as the admission key. ResolvedBinding loses subject_user_id (retired-field JSON still deserializes; persisted-shape test updated); the actor is the one identity. - ironclaw_assistant: ProductInstallationScope drops the default-subject, static-route, and subject-resolver knobs for one shared_conversation_admission port; resolve/lookup/reset check admission fail-closed (no port wired, or an unlisted conversation, rejects with a not-connected BindingRequired); resolve passes no trusted owner — the conversations domain keys and owns shared bindings by the paired actor. Thread and turn scopes derive their owner from the binding's actor on every route kind. - ironclaw_extension_host: channel_subject_routes.rs becomes channel_shared_admission.rs; ChannelConfigSharedAdmission admits by membership in the operator-saved *_allowed_channels JSON array; the managed derived subject (user:{ext}-channel:{sha16}) is deleted; legacy *_subject_routes values are inert (pinned by test). Shared conversations are no longer offered as per-user notification delivery targets — their ownership came from the retired subject map — and stored channel-target preferences fail closed at resolution; DM targets are unchanged. - slack manifest: slack_shared_subject_user_id and slack_subject_routes are retired with a gravestone comment; slack_allowed_channels is the admission surface (saves to the retired handles already fail closed as unknown fields — the extension-config analog of the config.toml retired-section gravestone). - architecture tests: the INVERTED_PORTS row moves with the port rename. User-visible consequences (also in the PR body): each shared-channel participant now gets their own persistent thread and must be paired; no cross-user shared context; the operator's identity is never a fallback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(reborn): align guidance, specs, and live-QA scripts with invoker scope Guidance rows and the moved-ports list rename the inverted port (SharedConversationAdmission, ex ProductConversationSubjectRouteResolver); the assistant boundary prose states the new rule (one thread per (conversation, actor), admission is the only shared-conversation configuration, fail-closed on resolve/lookup/reset). The composition CONTRACT.md's never-shipped per-channel subject admin API section is excised with a dated correction; CHECKLIST/PROPOSAL get dated amendments beside the historical text. Operator docs teach slack_allowed_channels + per-user pairing. CHANGELOG records the behavior change and the retired config fields. The live-QA scripts drop subject handling for allowed-channels admission (200 script tests green), and the orphaned canary env var is removed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(telegram): connect group chats via telegram_allowed_channels Fail-closed shared-conversation admission left Telegram groups with no operator affordance to connect one — the manifest declared no *_allowed_channels handle, so every group/supergroup @-mention was unadmittable. Declare the handle (the same generic [channel.config] convention Slack uses): listed chats are served with each participant running as themselves once paired; unlisted groups stay fail-closed. Previously any group the bot was added to ran as the deployment operator, which is the exposure this branch removes. Surfaced by the integration scenario telegram_update_becomes_a_turn_and_a_coordinated_reply failing closed after the admission change — kept red until this ruling rather than narrowed to a private chat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(reborn): morph the test tier to invoker scope Every fixture and pin that carried the retired subject model moves to the per-actor rule, with a recorded rationale at each semantic morph: - extension-host channel e2e: an admitted (allowed-channels) shared channel runs as the paired actor; unlisted conversations stay rejected; stored shared-channel outbound targets and binding refs fail closed; the telegram supergroup journey admits its chat via the new telegram_allowed_channels handle and proves the reply as the invoker. - assistant contract suites: admission replaces subject-route coverage (recording/failing/admit-all doubles; not-connected rejections on resolve/lookup/reset including existing bindings — a deliberate flip from the old existing-binding exemption; admission precedes actor-pairing side effects; direct routes never consult admission; per-actor threads for two participants; lookups never surface another actor's thread). - root integration harness + journeys: the binding fake, thread/turn scopes, and the group canonical user derive from the actor; multi-actor isolation pins unchanged and strictly stronger. - parity QA binary harness: subject resolution returns the actor. - webui product API redaction pin: the new telegram admission handle joins the admin-metadata forbidden list. Suites: extension_host 390/0; assistant 1084/0; conversations 105/0; architecture suite full pass; integration bins: extension_delivery 21/0 (Postgres legs under colima), delivery_user_journeys 22/0, mcp 22/0, trace_capture 14/0, generated_gate_sequences 29/0, group_journeys 16/0, group_multiuser 14/0. Workspace cargo fmt applied. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(changelog): record the telegram_allowed_channels admission field Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(merge): reconcile composition ceilings and capability_wiring test arity Post-merge fixups after folding main (nearai#7157 squashed + nearai#7214 + the inspector prompt-diagnostic work) into run-acts-as-invoker: - Re-capture the composition absolute-mass ceiling 40_747 -> 40_811 in both the budget manifest and reborn_restructure_baselines.rs: the acting-user scope helper and shared-admission wiring add +64 production LOC on the merged tree. Recorded rather than parked in the 150-line tolerance. - Add the 10th `tool_diagnostic_sink` argument (None) to the invoker's capability_wiring test call — main grew the signature after this branch wrote that call site. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(conversations): refuse legacy-row route-kind mismatches; drop unreachable widen A Direct request is the one key shape a retained legacy conversation-scoped shared row can collide with. Resolve, lookup, reset, and link now refuse the mismatch outright (BindingRequired) instead of trusting adapters never to re-classify a conversation's route kind — pinned by a Direct-probe leg on the legacy restart-path test. The forward half of the migration contract is pinned too: per_actor_shared_bindings_keep_their_threads_across_reopen proves a new per-actor shared binding survives a restart (a deserialize-side regression would previously have orphaned every group thread silently). widen_binding_route_access and ReplyRouteAccess::allow_shared are deleted: every Shared-keyed row is born shared under per-actor keying, so both widen call sites were unreachable. The persisted flag stays for legacy reads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key gate notices by gate ref on the triggered lane too The gate-collapse fix shipped on the observer lane only; the background lane still minted undiscriminated projection ids for ApprovalNeeded/AuthRequired, so an automation run parking on a SECOND gate deduped to AlreadyDelivered, recorded the whole delivery Failed, and the gate was never announced or reply-routable (AGENTS.md: fix the sibling when a pattern bug is fixed). TriggeredNotification's discriminator now carries the gate ref for gate prompts (RunBlocked stand-ins compose their label with it), matching the observer keying, with a triggered two-gate regression pinning outcome, prompts, and both reply routes. Also pinned: same-gate re-announcement dedupe (g1->g2->g1), two distinct AUTH gates, and the refless id shapes incl. FinalReplyReady. Over-long discriminators are bounded with a stable FNV-1a suffix so a maximal legal TurnGateRef can never overflow ProjectionUpdateRef and silently lose a notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(identity): one contract derivation for every acting-identity scope LoopRunContext::acting_resource_scope joins acting_user_id on the contract type: the raise/resume scope recipe both gate-dance crates hand-synced is now declared once, and every surviving ladder delegates — composition's owner-first resource_scope_for_run (workspace/skill mounts) and the inline grant-minting copy, loop_host's synthetic resume load, and project_create_capability's effective_user_id (deleted; its doc claimed a mirror that no longer existed). On the only run shape where owner and actor differ — legacy runs parked across the deploy — mounts and grants now follow the ACTOR like the rest of the dance; the pin flip is recorded in visible_capability_request_uses_acting_user_for_runtime_scope. The ladder is unit-pinned in its owning crate (all three rungs) and the accepted deploy-boundary resume-miss is pinned on the synthetic port with an acting-scope positive control. loop_contracts ceiling re-captured 13094 -> 13107 with provenance (the +13-line contract method). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extension-host): collapse admission handles; operator-identity channels never admit ChannelConfigSharedAdmission now holds the one declared *_allowed_channels handle as a plain String (the scan returns Option<String>): 'installed but handle-less' is no longer representable and the per-request Option branch is gone. The root pub use of the admission items is removed — consumers are crate-local and use the module path. Structural closure of the no-auth-vendor residual: a channel whose actor identity is not per-user (no OAuth vendor, no pairing strategy) never receives an admission resolver at all — an operator-identity channel that admitted a group would run every participant as the operator, the exact exposure run-acts-as-invoker removed. Previously this was unreachable only by manifest inventory. The extension_manager wire-shape pin gains the telegram_allowed_channels row (production projection was already correct), and extension_delivery gains the caller-path rejection leg: a correctly-signed webhook for an UNLISTED supergroup is acknowledged but produces no turn and no reply. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test+docs(reborn): re-pin parity to per-actor scopes; align contract, docs, vocabulary The identity-parity bin now pins the run-acts-as-invoker property its fixture can actually express: the shared-room support binding keeps ONE thread (the per-actor thread model is locked at the integration tier by scenario_two_actors_own_threads), and inside that shared thread each RUN's scope is owned by its own invoking actor with identity context never crossing. The shared-admission suite gains the reset checkpoint leg (deny before rotation, thread survives), and the connect-nudge suite's shared leg is documented as the deliberate unpaired-participant silence contract. docs/reborn/contracts/conversation-binding.md (the owning contract) is amended: per-actor key in rule 8, participant widening retired in rule 14, subject ownership struck in rule 24, and the admission/retention semantics recorded. Operator docs and CHANGELOG state the real unpaired-shared behavior (silence; pairing via Extensions; DMs still nudge), the CHANGELOG gains the both-lanes gate-announcement entry and Added-first ordering, CHECKLIST's contradictory open-status is reconciled with a dated note, the new REBORN_WEBUI_V2_LIVE_QA_SLACK_ALLOWED_CHANNELS is threaded through live-canary.yml, observer.rs carries its arch-exempt annotation, and retired 'subject' vocabulary is renamed out of live test support and doc comments. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This was referenced Aug 10, 2026
Kampouse
pushed a commit
to Kampouse/ironclaw
that referenced
this pull request
Aug 13, 2026
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse
pushed a commit
to Kampouse/ironclaw
that referenced
this pull request
Aug 13, 2026
nearai#7157 follow-ups) (nearai#7377) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key observer gate notices by their gate ref One run can park on several approval/auth gates in sequence (the observer's blocked-state loop re-announces whenever the (status, gate) marker changes), but the live observer derived every gate notice's projection id with no discriminator, so all ApprovalNeeded notices in one run collapsed to a single durable delivery identity. The second gate's prompt came back AlreadyDelivered from the coordinator, was treated as success, and was never sent — the user was never told about the gate their run was parked on, and no reply route was recorded for it, so a bare `approve` could not resolve it either. Key the projection id by the notification's gate ref (the mechanism nearai#7157 added for the triggered notifier's RunBlocked notices). A repeat announcement of the SAME gate still dedupes; kinds that carry no gate ref (FinalReplyReady) keep the historical undiscriminated id shape so existing delivery identities are not re-keyed. Regression: observer_delivers_a_prompt_for_each_distinct_approval_gate drives the real DeliveryCoordinator over the real outbound store through two distinct scripted gates and asserts two delivered prompts plus a recorded reply route for each. Sabotage-verified: reverting the discriminator to None fails exactly this test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(composition): pin the notification-channels gate dance when owner ≠ actor The full builtin.notification_channels_set approval dance — raise, replay payload, user approve (store + lease mint from the stored row), approved resume, lease claim, dispatch, lease consume — driven through the real capability port on a run whose thread owner differs from its acting user. Pins two properties ahead of unifying the scope derivation onto the actor: raise and resume must derive the same scope (every store in the dance is scope-keyed, so a half-unified derivation strands the approved capability), and whose identity that scope carries (the thread owner, under the interim split nearai#7157 shipped). The approve step mints the lease from the stored request's own scope, grantee, and fingerprint — the same material the production click-approval resolution uses — never a re-derivation. Capability-host tier rather than tests/integration because the product rule "a run acts as its invoker" makes owner ≠ actor unconstructible through every product front door; the run-context shape remains legal kernel state (runs parked across the deploy boundary carry it). The owner == actor dance stays covered end-to-end at the integration tier (outbound_target.rs). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(outbound): scope the whole notification-channels gate dance as the acting user Unify the interim nearai#7157 split: resource_scope_for_run and settings_scope_for_run now derive from the acting user like caller_for_run already did, and effective_user_id (the owner-first ladder) is deleted. The approval-gate raise, the replay payload, the durable gate record, the lease, and the approval-settings read all follow the user who invoked the run — so the invoker sees and approves the gate, and their settings govern it. The raise/resume coverage added one commit earlier ran before and after this change and caught a real half-unification in between: the resume-side replay load lives in ironclaw_loop_host's synthetic-capability wrap (a different crate from the raise-side save in notification_channels_set) and still derived owner-first, stranding an approved resume with "replay payload is unavailable". The acting-identity ladder now has exactly one definition — LoopRunContext::acting_user_id in ironclaw_loop_contracts — and both sides delegate to it, so the hand-synced-copy class is closed rather than re-synced. notification_channels_set's replay/gate-record writes move from the capability_host-wide owner-first helper onto the outbound module's base_resource_scope_for_run so every store in one dance derives one user; the capability_host-wide helper itself is unchanged (thread/durable-result scoping legitimately follows thread ownership, and owner == actor on every binding created under the run-acts-as-invoker rule). Loop-contracts size ceiling: +16 lines for the shared ladder, paid for by splitting host/run_context.rs's 104-line inline #[cfg(test)] module into its run_context/tests.rs sibling; ceiling re-captured DOWN 13_115 -> 13_028 from the gate's own failure message. Runs raised before this change with owner != actor and resumed after it will miss their replay payload once and fail closed; re-requesting approval recovers. Documented in the PR body. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(conversations): key shared-route bindings per (conversation, actor) A run acts as the user who invoked it, so a shared conversation binds one thread per paired actor — each owned by that actor — instead of one conversation-wide thread owned by a configured subject. BindingKey gains a serde-defaulted shared_actor_user_id component (None for Direct routes, whose identity stays the conversation alone); the trusted-owner parameter is deliberately ignored on Shared creates (it remains the trigger lane's way to bind Direct conversations for their creator), and the legacy shared-owner backfill is removed with it. Migration is ignore-but-retain, pinned with a restart-path test: legacy Direct keys deserialize byte-identically (continuity), while legacy conversation-keyed shared rows deserialize to a key no per-actor lookup builds — retained in durable state untouched, and every participant (including the old subject) starts a fresh thread they own. Morphed legacy pins record what became structural: a shared probe/lookup can no longer address (or widen) a Direct binding at all; stored reply targets are isolated per actor; an actor's unpair cannot take the conversation away from other participants' own threads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(product)!: remove shared-route subject binding; scope = invoker Owner ruling: a run acts as the user who invoked it, in a DM and in a shared channel alike, with one thread per (conversation, user). This removes the subject half of shared-route configuration end to end and keeps the admission half, fail-closed: - ironclaw_product_contracts: subject_route becomes shared_admission — the SharedConversationAdmission port answers only "is this shared conversation connected"; ProductConversationRouteKey survives as the admission key. ResolvedBinding loses subject_user_id (retired-field JSON still deserializes; persisted-shape test updated); the actor is the one identity. - ironclaw_assistant: ProductInstallationScope drops the default-subject, static-route, and subject-resolver knobs for one shared_conversation_admission port; resolve/lookup/reset check admission fail-closed (no port wired, or an unlisted conversation, rejects with a not-connected BindingRequired); resolve passes no trusted owner — the conversations domain keys and owns shared bindings by the paired actor. Thread and turn scopes derive their owner from the binding's actor on every route kind. - ironclaw_extension_host: channel_subject_routes.rs becomes channel_shared_admission.rs; ChannelConfigSharedAdmission admits by membership in the operator-saved *_allowed_channels JSON array; the managed derived subject (user:{ext}-channel:{sha16}) is deleted; legacy *_subject_routes values are inert (pinned by test). Shared conversations are no longer offered as per-user notification delivery targets — their ownership came from the retired subject map — and stored channel-target preferences fail closed at resolution; DM targets are unchanged. - slack manifest: slack_shared_subject_user_id and slack_subject_routes are retired with a gravestone comment; slack_allowed_channels is the admission surface (saves to the retired handles already fail closed as unknown fields — the extension-config analog of the config.toml retired-section gravestone). - architecture tests: the INVERTED_PORTS row moves with the port rename. User-visible consequences (also in the PR body): each shared-channel participant now gets their own persistent thread and must be paired; no cross-user shared context; the operator's identity is never a fallback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(reborn): align guidance, specs, and live-QA scripts with invoker scope Guidance rows and the moved-ports list rename the inverted port (SharedConversationAdmission, ex ProductConversationSubjectRouteResolver); the assistant boundary prose states the new rule (one thread per (conversation, actor), admission is the only shared-conversation configuration, fail-closed on resolve/lookup/reset). The composition CONTRACT.md's never-shipped per-channel subject admin API section is excised with a dated correction; CHECKLIST/PROPOSAL get dated amendments beside the historical text. Operator docs teach slack_allowed_channels + per-user pairing. CHANGELOG records the behavior change and the retired config fields. The live-QA scripts drop subject handling for allowed-channels admission (200 script tests green), and the orphaned canary env var is removed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(telegram): connect group chats via telegram_allowed_channels Fail-closed shared-conversation admission left Telegram groups with no operator affordance to connect one — the manifest declared no *_allowed_channels handle, so every group/supergroup @-mention was unadmittable. Declare the handle (the same generic [channel.config] convention Slack uses): listed chats are served with each participant running as themselves once paired; unlisted groups stay fail-closed. Previously any group the bot was added to ran as the deployment operator, which is the exposure this branch removes. Surfaced by the integration scenario telegram_update_becomes_a_turn_and_a_coordinated_reply failing closed after the admission change — kept red until this ruling rather than narrowed to a private chat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(reborn): morph the test tier to invoker scope Every fixture and pin that carried the retired subject model moves to the per-actor rule, with a recorded rationale at each semantic morph: - extension-host channel e2e: an admitted (allowed-channels) shared channel runs as the paired actor; unlisted conversations stay rejected; stored shared-channel outbound targets and binding refs fail closed; the telegram supergroup journey admits its chat via the new telegram_allowed_channels handle and proves the reply as the invoker. - assistant contract suites: admission replaces subject-route coverage (recording/failing/admit-all doubles; not-connected rejections on resolve/lookup/reset including existing bindings — a deliberate flip from the old existing-binding exemption; admission precedes actor-pairing side effects; direct routes never consult admission; per-actor threads for two participants; lookups never surface another actor's thread). - root integration harness + journeys: the binding fake, thread/turn scopes, and the group canonical user derive from the actor; multi-actor isolation pins unchanged and strictly stronger. - parity QA binary harness: subject resolution returns the actor. - webui product API redaction pin: the new telegram admission handle joins the admin-metadata forbidden list. Suites: extension_host 390/0; assistant 1084/0; conversations 105/0; architecture suite full pass; integration bins: extension_delivery 21/0 (Postgres legs under colima), delivery_user_journeys 22/0, mcp 22/0, trace_capture 14/0, generated_gate_sequences 29/0, group_journeys 16/0, group_multiuser 14/0. Workspace cargo fmt applied. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(changelog): record the telegram_allowed_channels admission field Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(merge): reconcile composition ceilings and capability_wiring test arity Post-merge fixups after folding main (nearai#7157 squashed + nearai#7214 + the inspector prompt-diagnostic work) into run-acts-as-invoker: - Re-capture the composition absolute-mass ceiling 40_747 -> 40_811 in both the budget manifest and reborn_restructure_baselines.rs: the acting-user scope helper and shared-admission wiring add +64 production LOC on the merged tree. Recorded rather than parked in the 150-line tolerance. - Add the 10th `tool_diagnostic_sink` argument (None) to the invoker's capability_wiring test call — main grew the signature after this branch wrote that call site. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(conversations): refuse legacy-row route-kind mismatches; drop unreachable widen A Direct request is the one key shape a retained legacy conversation-scoped shared row can collide with. Resolve, lookup, reset, and link now refuse the mismatch outright (BindingRequired) instead of trusting adapters never to re-classify a conversation's route kind — pinned by a Direct-probe leg on the legacy restart-path test. The forward half of the migration contract is pinned too: per_actor_shared_bindings_keep_their_threads_across_reopen proves a new per-actor shared binding survives a restart (a deserialize-side regression would previously have orphaned every group thread silently). widen_binding_route_access and ReplyRouteAccess::allow_shared are deleted: every Shared-keyed row is born shared under per-actor keying, so both widen call sites were unreachable. The persisted flag stays for legacy reads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key gate notices by gate ref on the triggered lane too The gate-collapse fix shipped on the observer lane only; the background lane still minted undiscriminated projection ids for ApprovalNeeded/AuthRequired, so an automation run parking on a SECOND gate deduped to AlreadyDelivered, recorded the whole delivery Failed, and the gate was never announced or reply-routable (AGENTS.md: fix the sibling when a pattern bug is fixed). TriggeredNotification's discriminator now carries the gate ref for gate prompts (RunBlocked stand-ins compose their label with it), matching the observer keying, with a triggered two-gate regression pinning outcome, prompts, and both reply routes. Also pinned: same-gate re-announcement dedupe (g1->g2->g1), two distinct AUTH gates, and the refless id shapes incl. FinalReplyReady. Over-long discriminators are bounded with a stable FNV-1a suffix so a maximal legal TurnGateRef can never overflow ProjectionUpdateRef and silently lose a notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(identity): one contract derivation for every acting-identity scope LoopRunContext::acting_resource_scope joins acting_user_id on the contract type: the raise/resume scope recipe both gate-dance crates hand-synced is now declared once, and every surviving ladder delegates — composition's owner-first resource_scope_for_run (workspace/skill mounts) and the inline grant-minting copy, loop_host's synthetic resume load, and project_create_capability's effective_user_id (deleted; its doc claimed a mirror that no longer existed). On the only run shape where owner and actor differ — legacy runs parked across the deploy — mounts and grants now follow the ACTOR like the rest of the dance; the pin flip is recorded in visible_capability_request_uses_acting_user_for_runtime_scope. The ladder is unit-pinned in its owning crate (all three rungs) and the accepted deploy-boundary resume-miss is pinned on the synthetic port with an acting-scope positive control. loop_contracts ceiling re-captured 13094 -> 13107 with provenance (the +13-line contract method). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extension-host): collapse admission handles; operator-identity channels never admit ChannelConfigSharedAdmission now holds the one declared *_allowed_channels handle as a plain String (the scan returns Option<String>): 'installed but handle-less' is no longer representable and the per-request Option branch is gone. The root pub use of the admission items is removed — consumers are crate-local and use the module path. Structural closure of the no-auth-vendor residual: a channel whose actor identity is not per-user (no OAuth vendor, no pairing strategy) never receives an admission resolver at all — an operator-identity channel that admitted a group would run every participant as the operator, the exact exposure run-acts-as-invoker removed. Previously this was unreachable only by manifest inventory. The extension_manager wire-shape pin gains the telegram_allowed_channels row (production projection was already correct), and extension_delivery gains the caller-path rejection leg: a correctly-signed webhook for an UNLISTED supergroup is acknowledged but produces no turn and no reply. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test+docs(reborn): re-pin parity to per-actor scopes; align contract, docs, vocabulary The identity-parity bin now pins the run-acts-as-invoker property its fixture can actually express: the shared-room support binding keeps ONE thread (the per-actor thread model is locked at the integration tier by scenario_two_actors_own_threads), and inside that shared thread each RUN's scope is owned by its own invoking actor with identity context never crossing. The shared-admission suite gains the reset checkpoint leg (deny before rotation, thread survives), and the connect-nudge suite's shared leg is documented as the deliberate unpaired-participant silence contract. docs/reborn/contracts/conversation-binding.md (the owning contract) is amended: per-actor key in rule 8, participant widening retired in rule 14, subject ownership struck in rule 24, and the admission/retention semantics recorded. Operator docs and CHANGELOG state the real unpaired-shared behavior (silence; pairing via Extensions; DMs still nudge), the CHANGELOG gains the both-lanes gate-announcement entry and Added-first ordering, CHECKLIST's contradictory open-status is reconciled with a dated note, the new REBORN_WEBUI_V2_LIVE_QA_SLACK_ALLOWED_CHANNELS is threaded through live-canary.yml, observer.rs carries its arch-exempt annotation, and retired 'subject' vocabulary is renamed out of live test support and doc comments. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse
pushed a commit
to Kampouse/ironclaw
that referenced
this pull request
Aug 13, 2026
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse
pushed a commit
to Kampouse/ironclaw
that referenced
this pull request
Aug 13, 2026
nearai#7157 follow-ups) (nearai#7377) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key observer gate notices by their gate ref One run can park on several approval/auth gates in sequence (the observer's blocked-state loop re-announces whenever the (status, gate) marker changes), but the live observer derived every gate notice's projection id with no discriminator, so all ApprovalNeeded notices in one run collapsed to a single durable delivery identity. The second gate's prompt came back AlreadyDelivered from the coordinator, was treated as success, and was never sent — the user was never told about the gate their run was parked on, and no reply route was recorded for it, so a bare `approve` could not resolve it either. Key the projection id by the notification's gate ref (the mechanism nearai#7157 added for the triggered notifier's RunBlocked notices). A repeat announcement of the SAME gate still dedupes; kinds that carry no gate ref (FinalReplyReady) keep the historical undiscriminated id shape so existing delivery identities are not re-keyed. Regression: observer_delivers_a_prompt_for_each_distinct_approval_gate drives the real DeliveryCoordinator over the real outbound store through two distinct scripted gates and asserts two delivered prompts plus a recorded reply route for each. Sabotage-verified: reverting the discriminator to None fails exactly this test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(composition): pin the notification-channels gate dance when owner ≠ actor The full builtin.notification_channels_set approval dance — raise, replay payload, user approve (store + lease mint from the stored row), approved resume, lease claim, dispatch, lease consume — driven through the real capability port on a run whose thread owner differs from its acting user. Pins two properties ahead of unifying the scope derivation onto the actor: raise and resume must derive the same scope (every store in the dance is scope-keyed, so a half-unified derivation strands the approved capability), and whose identity that scope carries (the thread owner, under the interim split nearai#7157 shipped). The approve step mints the lease from the stored request's own scope, grantee, and fingerprint — the same material the production click-approval resolution uses — never a re-derivation. Capability-host tier rather than tests/integration because the product rule "a run acts as its invoker" makes owner ≠ actor unconstructible through every product front door; the run-context shape remains legal kernel state (runs parked across the deploy boundary carry it). The owner == actor dance stays covered end-to-end at the integration tier (outbound_target.rs). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(outbound): scope the whole notification-channels gate dance as the acting user Unify the interim nearai#7157 split: resource_scope_for_run and settings_scope_for_run now derive from the acting user like caller_for_run already did, and effective_user_id (the owner-first ladder) is deleted. The approval-gate raise, the replay payload, the durable gate record, the lease, and the approval-settings read all follow the user who invoked the run — so the invoker sees and approves the gate, and their settings govern it. The raise/resume coverage added one commit earlier ran before and after this change and caught a real half-unification in between: the resume-side replay load lives in ironclaw_loop_host's synthetic-capability wrap (a different crate from the raise-side save in notification_channels_set) and still derived owner-first, stranding an approved resume with "replay payload is unavailable". The acting-identity ladder now has exactly one definition — LoopRunContext::acting_user_id in ironclaw_loop_contracts — and both sides delegate to it, so the hand-synced-copy class is closed rather than re-synced. notification_channels_set's replay/gate-record writes move from the capability_host-wide owner-first helper onto the outbound module's base_resource_scope_for_run so every store in one dance derives one user; the capability_host-wide helper itself is unchanged (thread/durable-result scoping legitimately follows thread ownership, and owner == actor on every binding created under the run-acts-as-invoker rule). Loop-contracts size ceiling: +16 lines for the shared ladder, paid for by splitting host/run_context.rs's 104-line inline #[cfg(test)] module into its run_context/tests.rs sibling; ceiling re-captured DOWN 13_115 -> 13_028 from the gate's own failure message. Runs raised before this change with owner != actor and resumed after it will miss their replay payload once and fail closed; re-requesting approval recovers. Documented in the PR body. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(conversations): key shared-route bindings per (conversation, actor) A run acts as the user who invoked it, so a shared conversation binds one thread per paired actor — each owned by that actor — instead of one conversation-wide thread owned by a configured subject. BindingKey gains a serde-defaulted shared_actor_user_id component (None for Direct routes, whose identity stays the conversation alone); the trusted-owner parameter is deliberately ignored on Shared creates (it remains the trigger lane's way to bind Direct conversations for their creator), and the legacy shared-owner backfill is removed with it. Migration is ignore-but-retain, pinned with a restart-path test: legacy Direct keys deserialize byte-identically (continuity), while legacy conversation-keyed shared rows deserialize to a key no per-actor lookup builds — retained in durable state untouched, and every participant (including the old subject) starts a fresh thread they own. Morphed legacy pins record what became structural: a shared probe/lookup can no longer address (or widen) a Direct binding at all; stored reply targets are isolated per actor; an actor's unpair cannot take the conversation away from other participants' own threads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(product)!: remove shared-route subject binding; scope = invoker Owner ruling: a run acts as the user who invoked it, in a DM and in a shared channel alike, with one thread per (conversation, user). This removes the subject half of shared-route configuration end to end and keeps the admission half, fail-closed: - ironclaw_product_contracts: subject_route becomes shared_admission — the SharedConversationAdmission port answers only "is this shared conversation connected"; ProductConversationRouteKey survives as the admission key. ResolvedBinding loses subject_user_id (retired-field JSON still deserializes; persisted-shape test updated); the actor is the one identity. - ironclaw_assistant: ProductInstallationScope drops the default-subject, static-route, and subject-resolver knobs for one shared_conversation_admission port; resolve/lookup/reset check admission fail-closed (no port wired, or an unlisted conversation, rejects with a not-connected BindingRequired); resolve passes no trusted owner — the conversations domain keys and owns shared bindings by the paired actor. Thread and turn scopes derive their owner from the binding's actor on every route kind. - ironclaw_extension_host: channel_subject_routes.rs becomes channel_shared_admission.rs; ChannelConfigSharedAdmission admits by membership in the operator-saved *_allowed_channels JSON array; the managed derived subject (user:{ext}-channel:{sha16}) is deleted; legacy *_subject_routes values are inert (pinned by test). Shared conversations are no longer offered as per-user notification delivery targets — their ownership came from the retired subject map — and stored channel-target preferences fail closed at resolution; DM targets are unchanged. - slack manifest: slack_shared_subject_user_id and slack_subject_routes are retired with a gravestone comment; slack_allowed_channels is the admission surface (saves to the retired handles already fail closed as unknown fields — the extension-config analog of the config.toml retired-section gravestone). - architecture tests: the INVERTED_PORTS row moves with the port rename. User-visible consequences (also in the PR body): each shared-channel participant now gets their own persistent thread and must be paired; no cross-user shared context; the operator's identity is never a fallback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(reborn): align guidance, specs, and live-QA scripts with invoker scope Guidance rows and the moved-ports list rename the inverted port (SharedConversationAdmission, ex ProductConversationSubjectRouteResolver); the assistant boundary prose states the new rule (one thread per (conversation, actor), admission is the only shared-conversation configuration, fail-closed on resolve/lookup/reset). The composition CONTRACT.md's never-shipped per-channel subject admin API section is excised with a dated correction; CHECKLIST/PROPOSAL get dated amendments beside the historical text. Operator docs teach slack_allowed_channels + per-user pairing. CHANGELOG records the behavior change and the retired config fields. The live-QA scripts drop subject handling for allowed-channels admission (200 script tests green), and the orphaned canary env var is removed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(telegram): connect group chats via telegram_allowed_channels Fail-closed shared-conversation admission left Telegram groups with no operator affordance to connect one — the manifest declared no *_allowed_channels handle, so every group/supergroup @-mention was unadmittable. Declare the handle (the same generic [channel.config] convention Slack uses): listed chats are served with each participant running as themselves once paired; unlisted groups stay fail-closed. Previously any group the bot was added to ran as the deployment operator, which is the exposure this branch removes. Surfaced by the integration scenario telegram_update_becomes_a_turn_and_a_coordinated_reply failing closed after the admission change — kept red until this ruling rather than narrowed to a private chat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(reborn): morph the test tier to invoker scope Every fixture and pin that carried the retired subject model moves to the per-actor rule, with a recorded rationale at each semantic morph: - extension-host channel e2e: an admitted (allowed-channels) shared channel runs as the paired actor; unlisted conversations stay rejected; stored shared-channel outbound targets and binding refs fail closed; the telegram supergroup journey admits its chat via the new telegram_allowed_channels handle and proves the reply as the invoker. - assistant contract suites: admission replaces subject-route coverage (recording/failing/admit-all doubles; not-connected rejections on resolve/lookup/reset including existing bindings — a deliberate flip from the old existing-binding exemption; admission precedes actor-pairing side effects; direct routes never consult admission; per-actor threads for two participants; lookups never surface another actor's thread). - root integration harness + journeys: the binding fake, thread/turn scopes, and the group canonical user derive from the actor; multi-actor isolation pins unchanged and strictly stronger. - parity QA binary harness: subject resolution returns the actor. - webui product API redaction pin: the new telegram admission handle joins the admin-metadata forbidden list. Suites: extension_host 390/0; assistant 1084/0; conversations 105/0; architecture suite full pass; integration bins: extension_delivery 21/0 (Postgres legs under colima), delivery_user_journeys 22/0, mcp 22/0, trace_capture 14/0, generated_gate_sequences 29/0, group_journeys 16/0, group_multiuser 14/0. Workspace cargo fmt applied. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(changelog): record the telegram_allowed_channels admission field Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(merge): reconcile composition ceilings and capability_wiring test arity Post-merge fixups after folding main (nearai#7157 squashed + nearai#7214 + the inspector prompt-diagnostic work) into run-acts-as-invoker: - Re-capture the composition absolute-mass ceiling 40_747 -> 40_811 in both the budget manifest and reborn_restructure_baselines.rs: the acting-user scope helper and shared-admission wiring add +64 production LOC on the merged tree. Recorded rather than parked in the 150-line tolerance. - Add the 10th `tool_diagnostic_sink` argument (None) to the invoker's capability_wiring test call — main grew the signature after this branch wrote that call site. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(conversations): refuse legacy-row route-kind mismatches; drop unreachable widen A Direct request is the one key shape a retained legacy conversation-scoped shared row can collide with. Resolve, lookup, reset, and link now refuse the mismatch outright (BindingRequired) instead of trusting adapters never to re-classify a conversation's route kind — pinned by a Direct-probe leg on the legacy restart-path test. The forward half of the migration contract is pinned too: per_actor_shared_bindings_keep_their_threads_across_reopen proves a new per-actor shared binding survives a restart (a deserialize-side regression would previously have orphaned every group thread silently). widen_binding_route_access and ReplyRouteAccess::allow_shared are deleted: every Shared-keyed row is born shared under per-actor keying, so both widen call sites were unreachable. The persisted flag stays for legacy reads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key gate notices by gate ref on the triggered lane too The gate-collapse fix shipped on the observer lane only; the background lane still minted undiscriminated projection ids for ApprovalNeeded/AuthRequired, so an automation run parking on a SECOND gate deduped to AlreadyDelivered, recorded the whole delivery Failed, and the gate was never announced or reply-routable (AGENTS.md: fix the sibling when a pattern bug is fixed). TriggeredNotification's discriminator now carries the gate ref for gate prompts (RunBlocked stand-ins compose their label with it), matching the observer keying, with a triggered two-gate regression pinning outcome, prompts, and both reply routes. Also pinned: same-gate re-announcement dedupe (g1->g2->g1), two distinct AUTH gates, and the refless id shapes incl. FinalReplyReady. Over-long discriminators are bounded with a stable FNV-1a suffix so a maximal legal TurnGateRef can never overflow ProjectionUpdateRef and silently lose a notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(identity): one contract derivation for every acting-identity scope LoopRunContext::acting_resource_scope joins acting_user_id on the contract type: the raise/resume scope recipe both gate-dance crates hand-synced is now declared once, and every surviving ladder delegates — composition's owner-first resource_scope_for_run (workspace/skill mounts) and the inline grant-minting copy, loop_host's synthetic resume load, and project_create_capability's effective_user_id (deleted; its doc claimed a mirror that no longer existed). On the only run shape where owner and actor differ — legacy runs parked across the deploy — mounts and grants now follow the ACTOR like the rest of the dance; the pin flip is recorded in visible_capability_request_uses_acting_user_for_runtime_scope. The ladder is unit-pinned in its owning crate (all three rungs) and the accepted deploy-boundary resume-miss is pinned on the synthetic port with an acting-scope positive control. loop_contracts ceiling re-captured 13094 -> 13107 with provenance (the +13-line contract method). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extension-host): collapse admission handles; operator-identity channels never admit ChannelConfigSharedAdmission now holds the one declared *_allowed_channels handle as a plain String (the scan returns Option<String>): 'installed but handle-less' is no longer representable and the per-request Option branch is gone. The root pub use of the admission items is removed — consumers are crate-local and use the module path. Structural closure of the no-auth-vendor residual: a channel whose actor identity is not per-user (no OAuth vendor, no pairing strategy) never receives an admission resolver at all — an operator-identity channel that admitted a group would run every participant as the operator, the exact exposure run-acts-as-invoker removed. Previously this was unreachable only by manifest inventory. The extension_manager wire-shape pin gains the telegram_allowed_channels row (production projection was already correct), and extension_delivery gains the caller-path rejection leg: a correctly-signed webhook for an UNLISTED supergroup is acknowledged but produces no turn and no reply. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test+docs(reborn): re-pin parity to per-actor scopes; align contract, docs, vocabulary The identity-parity bin now pins the run-acts-as-invoker property its fixture can actually express: the shared-room support binding keeps ONE thread (the per-actor thread model is locked at the integration tier by scenario_two_actors_own_threads), and inside that shared thread each RUN's scope is owned by its own invoking actor with identity context never crossing. The shared-admission suite gains the reset checkpoint leg (deny before rotation, thread survives), and the connect-nudge suite's shared leg is documented as the deliberate unpaired-participant silence contract. docs/reborn/contracts/conversation-binding.md (the owning contract) is amended: per-actor key in rule 8, participant widening retired in rule 14, subject ownership struck in rule 24, and the admission/retention semantics recorded. Operator docs and CHANGELOG state the real unpaired-shared behavior (silence; pairing via Extensions; DMs still nudge), the CHANGELOG gains the both-lanes gate-announcement entry and Added-first ordering, CHECKLIST's contradictory open-status is reconciled with a dated note, the new REBORN_WEBUI_V2_LIVE_QA_SLACK_ALLOWED_CHANNELS is threaded through live-canary.yml, observer.rs carries its arch-exempt annotation, and retired 'subject' vocabulary is renamed out of live test support and doc comments. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
…ry crate, and a repo-wide stale sweep (nearai#7264) * docs(target-arch): resolve the await-edge design question by measurement (D-S) and re-walk the WS9 verify row Appends §12.13 D-S under delegated authority at owner direction, flagged for post-hoc review by Illia Polosukhin (nearai#6696's author): the await-edge store is measured to be a pure projection over ProcessDependencyPort (that half of the shed happened inside nearai#6696 itself), and the resolver is a genuine loop-tier responsibility journal edges cannot express (owner recovery, sanitized transcript result materialization, batch-gate resume-once drain, BlockedDependentRunGate resume policy). §6.7.3 is amended (scheduler DONE / store DONE / resolver KEEP) instead of the shed being executed; the 2.9k figure is corrected to 1,459 production + 1,448 cfg(test) lines. The §12.10 bullet, §2 divergence flag, §9 row 49, §13 validation row, CHECKLIST header/WS4 pointer, README and PLAN all carry the dated resolution. WS9 verify row ticked with evidence: one lifecycle authority (the process journal; TurnRunState/TurnRunRecord are projections via AgentTurnProcessRuntime, ProcessRecord is a capability-invocation view, no bare RunRecord exists) and §7 T4 re-walked clause-by-clause against merged code — matches, including the checkpoint-gated no-auto-retry mechanism (BeforeModel precedes ModelStage; requeue only when checkpoint-free under the 3-claim cap). Docs-only; no code, no tests, no gates touched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): rows 1-2 — package-set tick (64==64/1/0, gate+selftest+independent rederivation) and the 74-row §9 mapping audit (45 L / 15 L-A / 14 OBD / 0 NOT-LANDED; 3 findings recorded) Row 1: check-target-tree.py reports 64 workspace members == 64 documented packages, 1 documented exclusion (tools/ironclaw_silk_decoder), 0 owned exceptions (EXCEPTIONS table empty — §5 steady state); self-test 17/17; cargo-metadata name set diffed empty against an independent §5 parse. Row 2: docs/reborn/target-architecture/ws12-mapping-audit.md is the audit record — per-row executed-evidence, delete-clauses read against WS8's execution notes, all 14 open rows cite their owning CHECKLIST/PROPOSAL row or issue. Findings (recorded, not fixed): F1 prompt_envelope manifest-description fix has no owner row; F2 WS6:429's 'nearai#5618 residue deleted' overstates vs the live adopt_migrated_identity + open WS8:523; F3 stale-docs cluster where the tree is ahead of the prose (trace re-export drop, TurnRunTransitionPort, processes->resources). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fold(7154): squash-port fix/red-main-7119 onto family-world main — defect train nearai#7146/nearai#7115/nearai#7104/nearai#7103/nearai#7144 (+nearai#7119 CI lane), 34-hunk contribution.rs port into the split modules, planner entrypoint classification, D-R loopback exception on the widened HTTPS credential guard Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(extractors): issue-number + assertion-rationale doc refinement (rescued 844964f from rescue/7154-parked-guard) Ports only the doc/assertion refinement commit; the guard-parking commit e8f5a31 on that branch is deliberately NOT taken — superseded by the D-R loopback ruling. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): record D-R — the loopback credential-guard ruling, wiring choice, and regression pins (PROPOSAL §12.13, 2026-08-05) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7154): CodeRabbit round-1 triage — fail-closed tracing-target scan traversal (+node_modules), bounded sidecar output draining (capped capture + discard drain), deadlock regression asserts successful redaction (no seq), XLSX/DOCX empty-classification via extract_document, raise_for_status annotations Threads already addressed by the fold: latency.rs caller-contract wording (merged doc scopes the requirement to latency-trace callers), BodyJsonPointer coverage (the plaintext-refusal test drives all four injection shapes). Deliberately not taken: un-xfailing the four Slack-catalog projections — the xfail is a documented tripwire (unexpected-pass goes red) and clearing them is the nearai#6520 projection-modeling follow-on its comment specs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(assistant): re-point the one field-form tracing target the nearai#7146 gate caught — main's relocated triggered_run_delivery_services carried the drift the PR fixed at its old channel_host address Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * closure fixups: execute the mapping-audit findings — prompt_envelope manifest description (F1), dated ✎ corrections for the nearai#5618 overstatement (F2) and the stale-prose cluster (F3) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): second-reviewer security spot-audit + extension-journey re-verification (rows 5-6) Adversarial second-reviewer pass over PROPOSAL §12.1a/b/c and the batch's own §12.13 D-R loopback carve-out, plus a re-run of the five extension journeys. Attacks were executed rather than argued: two sabotage files and a 38-shape hostile-URL probe were planted, run, and reverted. Verdicts — mint consolidation HOLDS-WITH-RESIDUAL, secrets tightening HOLDS-WITH-RESIDUAL, host/verifier colocation HOLDS, D-R HOLDS. No HOLE. Four findings recorded rather than fixed (report-not-repair): - F1 test_verified/_for_tenant are ungranted mint constructors gated only by the `test-support` feature, in no mint-name table, with nothing pinning the feature to [dev-dependencies]; the shipped binary is measured feature-free. - F2 §12.1b's products-layer residue undercounts by one (ironclaw_assistant). - F3 journey coverage hole: gsuite-with-credential-injection is proven in two halves that no committed test joins. - F4 both recorded census evasions and both fail-open reads are CLOSED on this tree, so §11.2.5/§12.1a/CHECKLIST:552/:597 now understate the seal. Rows 5-6 ticked; only lines 631-632 of CHECKLIST touched so the concurrent rows 3-4 edit folds cleanly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ratchet(closure): lock the budget gate at the program's end state Dispatch ceiling 1122 -> 814 (today's observed, nudge taken; WS0 record 827 stays within effective 829). Mass-share ceiling 2398 -> 658 bp (the WS0 baseline floor — the arch-test assert refuses lower, and observed 578 bp sits inside the nudge window). Absolute LOC re-equalized at 40423: nearai#6831 added 4 governed LOC through the queue's tolerance window; ceiling, observed, and COMPOSITION_ABSOLUTE_SRC_LOC move together here. Both tightenings sabotage-verified red (dispatch 9-over at 790; abs 73-over at 40200). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): gauntlet report — row 3 ticked (full gauntlet green, 0 REAL in scope), row 4 verified-but-open on two pre-existing Postgres-leg test-isolation defects WS12 rows 3-4 verification on the assembled batch tip 0c6c0cf: Row 3 (ticked): fmt, clippy default/all-features/--lib --bins, workspace tests (495 targets, 15,203 passed, 0 failed; the smoke.rs:3132 CPU-saturation flake passed first try), arch suite 285/0, the integration-feature lane 1,665/0, recorded-fixture QA (61 fixtures clean, 41/0), frontend (typecheck 1,588 files; vitest 1,088/0; build + bundle budgets), e2e smoke = the CI browser lane under the hermetic wrapper (50 + 21 + 5 passed), and all 41 scripts/ci self-tests (two mapfile/bash-3.2 casualties green under bash 5, the CI shape). Row 4 (stays open, dated note added): both-backend parity proven with legs demonstrably executed for the fabric (57 pg + 81 libsql), triggers (ADR 0003, REQUIRE_POSTGRES), hooks (ADR 0004, all three backends), composition, processes journal, extension-registry, host-runtime libSQL restart, and the backend matrix; fabric-delegated domains enumerated. Two REAL blockers (one class): the Postgres legs of the event-store and assistant-ledger contract suites assert against shared-database state and cannot pass as-written (each failing test passes alone on a virgin database; files byte-identical to origin/main; no CI lane sets their env vars). Full evidence: docs/reborn/target-architecture/ws12-gauntlet-report.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): set the crate/family guidance convention The base commit for the family-guidance program: one canonical home per fact, measured-not-aspirational claims, boundaries stated as exclusions, and the note that guidance files can be gate-pinned. Every family/crate document written on top of this branch follows this shape. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): per-test isolated Postgres databases for the two WS12 parity-blocking contract suites The WS12 gauntlet (ws12-gauntlet-report.md §P6/§P8) measured the Postgres legs of ironclaw_event_store's durable_event_store_contract and ironclaw_assistant's durable_ledger_contract as test-isolation-defective: absolute database-global asserts (event cursors; settled-entry prune bookkeeping) run against the single external database named by their IRONCLAW_*_POSTGRES_URL env vars. Every failing test passes alone on a virgin database - store semantics correct, suites not self-isolating (PROPOSAL §12.13 D-T). Fix: each affected test provisions a private database on the configured server - the fabric contract's IsolatedDatabase pattern (db_root_filesystem_contract.rs) ported locally into each suite: CREATE DATABASE per test, store/pool + migrations against it, courtesy DROP ... WITH (FORCE), and a once-per-binary stale-name sweep. Every assertion preserved byte-identical; libsql/jsonl twins untouched. In the ledger suite only the two retention tests move - the other six Postgres tests keep their proven fingerprint-suffix isolation. Regression pins are the fixed tests themselves: - postgres_replay_advances_next_cursor_past_trailing_filtered_records - postgres_runtime_and_audit_logs_survive_rebuild_with_filtered_cursor_semantics - postgres_settled_entry_limit_prunes_oldest_when_configured - postgres_settled_prune_interval_defers_until_interval_when_configured Green proven on a shared dirty database twice in a row (parallel default threading) and serially on a virgin database; red-first reproduction captured before the fix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(reborn): record §12.13 D-T (parity-suite isolation ruling) and close CHECKLIST WS12 row 4 D-T (after D-S): the WS12 gauntlet's two REAL findings were one defect class — absolute database-global asserts against the single shared env-var Postgres database — in two suites (event store cursor contract, assistant settled-ledger retention). Ruling executed in commit 864d93e: per-test isolated databases via the fabric contract's IsolatedDatabase pattern, assertions preserved; alternatives (baseline-relative asserts, serial-only, leave-open) recorded with why they lost; regression pin = the four fixed tests themselves. CHECKLIST WS12 backend-parity row ticks [x] with a dated addendum: red-first reproduction, the three green isolation runs (dirty shared DB twice in parallel; failing pairs serial on virgin), parity now green 10/10. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(extensions): family guidance layer — AGENTS.md rewrite to the guidance-conventions shape, READMEs for all 4 family crates and 14 packages, duplicate-guidance consolidation The family AGENTS.md now teaches the unified extension model (extension = the only product object; channel/tool/auth are manifest surfaces; runtime is loading, never taxonomy; ExtensionId vs VendorId; retired vocabulary pinned by reborn_retired_taxonomy.rs), carries the self-containment and package-to-crate rules from families/extensions.md, the four-responsibility lookup, the measured package catalog, the exclusion list, and the armed gates by test name. Every crate and package gains a README.md (ironclaw_extension_host had no guidance of any kind). ironclaw_extension_registry and memory-native each had both an AGENTS.md and a CLAUDE.md saying overlapping things: AGENTS.md is now canonical, CLAUDE.md a pointer, and memory-native's stale v1 references (src/workspace, src/db/libsql) are dropped in the merge. The slack/telegram agent maps get package framing and a contracts-tier pointer in place of the stale ironclaw_assistant one. Every path literal verified to resolve on disk; all figures (tool counts, dep sets, consumers, layer declarations) measured from the tree at 8d13454. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): substrates + lanes family guidance per guidance-conventions.md Family AGENTS.md rewritten to the spec shape for crates/substrates/ and crates/lanes/: boundary, crate table, exclusion lists (mechanism-not-authority for substrates; kernel-decides-lane-executes for lanes), armed gates by test name, and measured deviations stated as deviations (sandbox's three substrate deps, script.rs direct spawn). The lanes wit/-is-load-bearing note is kept. A README.md for every crate in both families (10 new), measured against cargo metadata 2026-08-05: public surface, workspace edges, consumer counts, and enforced invariants each citing their gate. ironclaw_libsql_runtime and ironclaw_wasm_limiter previously had no guidance of any kind; their READMEs carry the sole-pool-home rule (ADDITIONAL_DRIVER_ALLOWLISTS: deadpool = {filesystem, libsql_runtime}) and the outbound-only limiter gate (wasm_sandbox_core_module_stays_domain_free_v1_parity_kernel; no BoundaryRule names the limiter). Duplicate guidance consolidated per rule 1: for the six crates holding both AGENTS.md and CLAUDE.md (filesystem, network, secrets, mcp, sandbox, wasm), CLAUDE.md stays canonical (module spec for filesystem; gate-pinned wording for mcp and wasm) and AGENTS.md becomes a short pointer. No gate-pinned file was edited. Stale reference removed: safety AGENTS.md pointed at src/NETWORK_SECURITY.md, which exists nowhere in the tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(crates-map): rewrite the three top-level maps family-first after the restructure crates/AGENTS.md (264 -> 175 lines): routing map only — the ten families, the read order (family AGENTS.md -> crate README.md -> working rules/module spec -> docs/reborn/contracts/), the enforced seven-layer matrix with the family/layer divergences, measured workspace facts (64 packages, 1 documented exclusion, 0 owned exceptions per scripts/ci/check-target-tree.py), and a verified command block. The 40-row per-crate map is gone: family AGENTS.md files own crate routing per docs/reborn/guidance-conventions.md. crates/README.md (141 -> 119 lines): human map — mental model in family vocabulary, the ten families with measured crate counts, the 14 extension packages (4 crates + 10 data-only), and the two workspace members outside crates/. crates/Architecture.md (1019 -> 1059 lines): audited against the live tree; every named symbol/path re-verified 2026-08-05. Corrected: retired ProductAdapter vocabulary (zero residue in code), the stale pre-rename dependency ladder that still cited the deleted gateway/TUI crates, run-state store mentions, lane-table crate anchors (sandbox/extension_support), declared-in vs minted-by owners in the core data model, and the subagent deny-filter status note (re-verified). Marked the pre-restructure 'partial or evolving' list as unmeasured rather than asserting it. Also documents that scripts/check-boundaries.sh fails on a clean tree (check-5 grep false positives) and greps the deleted v1 src/ in 4 of 6 checks — boundary enforcement for crates/ is the architecture suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(crates-map): package directories carry their own README.md (coordinator sync with extensions-family agent) Every extensions package dir — the 10 data-only ones included — now ships a README.md, so both maps extend the read order to package level. The sibling branch also confirmed what this map already derived per-crate: packages/ is not uniformly products-layer (memory-native and mem0 declare substrates). The other two coordinator corrections targeted rows of the old per-crate map, which this rewrite deleted wholesale; nothing here cites memory-native's CLAUDE.md or claims ironclaw_extension_host lacks guidance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): contracts + events family guidance layer per docs/reborn/guidance-conventions.md - crates/contracts/AGENTS.md and crates/events/AGENTS.md rewritten to the family shape: exclusion lists with destinations, armed gates by test name, layer-matrix rows, crossing guide, measured header counts. - README.md added for all 10 crates (ironclaw_prompt_envelope previously had no guidance of any kind — the CHECKLIST WS11 gap). - One canonical guidance file per crate, other file a pointer: A+C merges for ironclaw_host_api, ironclaw_event_log, ironclaw_event_projections, ironclaw_event_streams; CLAUDE-only content moved to AGENTS.md for ironclaw_loop_contracts, ironclaw_extension_contracts, ironclaw_product_contracts (none of these are root module-spec crates, so AGENTS.md is the working-rules home). - Stale guidance fixed against the live tree: * loop_contracts dep list contradicted the enforced allowlist (manifest is host_api + extension_contracts; common/prompt_envelope are permitted, unused). * event_log still documented the deleted jsonl parse/replay helpers. * event_projections still claimed EventStreamManager, DurableMemoryAuditSink, MemoryAuditProjectionMetadata, and PendingGateProjection — all deleted per PROPOSAL 6.3.3. * product_contracts still carried the pre-D-E open vendor decision under the nonexistent module name llm_config, and a Deferred section contradicting its own operator_llm/operator_service rows. * extension_contracts module table was missing the WS3 runtime module while counting 18. * common's llm_costs note carried the ModelCostTable seam claim refuted by PROPOSAL 12.11 D-F; now cites the pricer-port ruling and the vendor census residue. - Deleted crates/events/ironclaw_event_projections/PENDING_GATE_PROJECTION.md: every claim in it referenced deleted symbols or the removed v1 src/ tree, and its only inbound reference was the crate's own CLAUDE.md. Verified: all consumer counts reproduce via the printed grep commands; 147 path literals across the 28 touched files resolve on disk; no architecture test reads any of these files by name; conflict-marker scan clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): three measured corrections surfaced by the guidance program memory packages are substrates-layer, not products (families/extensions.md); memory_native declares no extension_contracts dep (PROPOSAL §6.8.4); wasm's extension_contracts edge is dev-only and the wasm 'never depends on' bullet is lane-scoped, not family-wide (families/lanes.md). Three further reported defects were checked and NOT corrected — they were misreads: the sandbox 'never above the runtime tier' rule holds (substrates sit below it), and PROPOSAL's safety consumer count already reads 17, matching the tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(kernel): family guidance layer — perimeter AGENTS.md, nine crate READMEs, AGENTS/CLAUDE consolidation Family-guidance program, kernel family (guidance-conventions.md shape): - crates/kernel/AGENTS.md rewritten to the family shape: the nine-stage effect pipeline with stage ownership, the sealed-mint table (witness / trust ceiling / approval lease / verified-inbound evidence, each with its mint site and its seal mechanism), the per-stage fail-closed table with file:line or test citations, the sharp exclusion list, and the armed gates by test name (authorized-seal ratchet, sealed-evidence mint ratchet, BoundaryRules, same-layer edge inventory at 21 kernel edges, empty LAYER_MATRIX_EXCEPTIONS register, driver boundary, process storage scan, origin-gate matrix ratchet). - A README.md for each of the nine crates, per the crate shape: measured workspace deps and consumer counts (cargo metadata), public surface with verified citations, enforced invariants naming their gates. ironclaw_processes states the single-lifecycle-authority direction of truth (journal = store; TurnRunState/ProcessRecord/await-edge = projections; PROPOSAL §12.13 D-S); ironclaw_host_runtime documents the D-R literal-loopback carve-out and names its two regression tests. - Duplicate guidance reconciled in all nine crates: AGENTS.md is canonical (guardrails absorbed), CLAUDE.md reduced to a pointer; ironclaw_trust's CONTRACT.md untouched as the co-located cross-crate contract. - Stale references fixed inside owned paths: the deleted capability-profile conformance module (evaluate_profile_conformance — zero hits workspace-wide) removed from ironclaw_capabilities guidance; trust's 'staging branch' / 'PR3' phrasing updated; capabilities' 'later obligation slices' updated to the landed host_runtime obligations split; cross-crate path mentions fully qualified. Every path literal in all 29 kernel .md files verified to resolve on disk; every named symbol swept against crate sources. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(domains): family guidance layer — AGENTS.md boundary doc, 12 crate READMEs, duplicate-guidance consolidation, stale-path fixes Family guidance for crates/domains/ per docs/reborn/guidance-conventions.md: - crates/domains/AGENTS.md rewritten to the family shape: charter table with go-here-when routing, the exclusion list, every armed gate named by test (BoundaryRules + identity/memory allowlists, the 5-entry in-family edge inventory, the naming gates, trusted-trigger ownership, the memory-provider residue ledger, persistence-driver boundary, the two module-charter gates). - A measured README.md for each of the 12 crates: charter, use-when / don't-use-when routing, public surface, measured normal deps + named consumers, enforced invariants with their gates, exact test commands. ironclaw_attachments and ironclaw_identity had no guidance of any kind; identity's README points at CONTRACT.md (the module spec), llm's at its CLAUDE.md module spec. - Duplicate guidance consolidated to one canonical file + pointer per crate: threads/conversations/memory/outbound rules now live in AGENTS.md (CLAUDE.md is a pointer); auth/llm keep CLAUDE.md canonical because their tests/module_charter.rs gates read it (AGENTS.md is the pointer). One misstatement fixed in the conversations merge: transcript content belongs to ironclaw_threads' SessionThreadService, not InboundConversationService. - Staleness fixed inside the family: identity CONTRACT.md two-edge allowlist claim reconciled with D-Q's three entries; trace_commons CLAUDE.md gains the capture module row and strikes its two discharged Known Gaps (recording/paths shims deleted, rename done); llm CLAUDE.md reasoning.rs caller corrected to crates/loop/ironclaw_loop_host; triggers lib.rs 'feature-gated' repo doc comments corrected; pre-family path literals in comments repointed (kernel/approvals+processes, loop/hooks, app/architecture_tests, domains/auth) and the deleted-v1-engine references in skills marked historical. Verified: cargo test -p ironclaw_llm --no-fail-fast (922 passed, exit 0 — CLAUDE.md is gate-pinned); cargo check --all-targets on all six crates with source edits; every cited path literal resolves on disk; conflict-marker scan clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): repair the corrupted kernel bullet and correct two family laws kernel.md: ironclaw_authorization's 'Security & authority role' bullet has been textually corrupted since nearai#6918 — an approvals sentence was spliced into it mid-clause, orphaning its continuation line. Reconstructed, with the spliced sentence restored to the approvals entry where it is true. lanes.md: 'a lane never depends on a substrate' is false as a family-wide law (ironclaw_sandbox holds network/safety/secrets normal deps, which its own entry licenses); the accurate law is the layer ladder, and the narrow claim holds for ironclaw_wasm alone. lanes.md + events.md: the 'every crate ships both an AGENTS.md and a CLAUDE.md' requirement is superseded by docs/reborn/guidance-conventions.md — two files restating one rule is the drift the guidance program removes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(arch): govern the ProtocolAuthEvidence test seam — WS12 audit F1 Two new gates in reborn_sealed_evidence_mint_ratchet (closed paths #12/#13), per the audit's remedy spec: (a) TEST_SEAM_MINT_FNS governs test_verified/test_verified_for_tenant — any production-text call site outside ironclaw_host_api is an offender (comments/strings stripped, #[cfg(test)] blocks stripped, tests.rs / *_tests.rs and cfg-test-only files excluded via the shared census); (b) test-support may appear in no normal dependency table workspace-wide (dependencies / build-dependencies / target.* variants / workspace.dependencies), and no [features] key other than test-support may forward to it — the laundering shape that would evade (b) by one rename. [dev-dependencies] enablement stays legal (cargo-features.md bar 4, the sanctioned dev seam). Measured zero offenders on this tree in both directions before pinning; sabotage-proven red->green both ways (planted production call named with file:line-text; [dependencies] enablement named with its table path). Self-tests drive the same pipelines the gates run (zero-match principle); the definition-location and partition tests now cover the new table. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(integration): join the gsuite credential-injection journey — WS12 audit F3 WS12 row 5 leg 3 was verified in two halves no committed test joined: gsuite handler -> staged credential (crate tier) and staged obligation -> wire (GitHub/Slack only). Scenario 5 already drives gmail.list_messages through production dispatch on a Google-OAuth-configured group; it now also asserts the JOIN: the seeded google account's token (itest-google-token) lands on the recorded outbound gmail.googleapis.com request as 'authorization: Bearer ...', injected at the host egress chokepoint (apply_credential_injection) per the gmail manifest's declared recipe — store -> dispatch-time staging -> chokepoint -> wire, through the caller. Sabotage-proven: disabling the Header injection arm reds exactly this scenario with 'no network egress request matching url gmail.googleapis.com has header authorization' while the request itself still reaches the wire (headers seen: content-type only) — the injection reason, not a setup error; restore -> green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): correct five measured dependency claims in families/domains.md conversations does not depend on safety (its BoundaryRule now forbids it); triggers depends on libsql_runtime + safety and NOT filesystem, so its 'filesystem-routed persistence path alongside SQL' is one path, not two; memory's live set is host_api alone (prompt_envelope is allowlisted, unused); auth was short by extension_contracts + product_contracts. Each verified against the manifest before editing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): family AGENTS.md + crate READMEs + guidance consolidation for loop/product/app Family-guidance program, families 8-10 (the top of the stack), per docs/reborn/guidance-conventions.md: - Rewrite crates/{loop,product,app}/AGENTS.md from routing stubs to the spec's family shape: exclusion lists, armed gates by test name, layer rows, crossing guides. Loop carries the trust story + the declared Loop*Port decorator chain; product carries the frozen-surface rule, the transports-consume-contracts rule (with the D-B frozen-constant qualification), the evidence-mint prohibition, and the two vendor exceptions; app carries the wires-owners-never-becomes-one charter, the binary-names-packages rule, config's zero-dep guarantee, and the composition mass ratchet (loc 40423 / Arc<dyn> 814). - Add a README.md to all 13 crates (12 new; webui's rewritten to the spec shape) with measured public surface, deps, and consumer counts. - Consolidate duplicate AGENTS.md/CLAUDE.md per spec rule 1: AGENTS.md is canonical and CLAUDE.md a pointer for agent_loop, loop_host, turn_runner, hooks, host_ingress, openai_compat, operator, and architecture_tests; CLAUDE.md stays canonical (module spec / gate-pinned) for webui, composition, and assistant, with composition's AGENTS.md reduced to the pointer. - Fix stale references in owned paths: hooks' dependency diagram and AgentLoopDriver home (ironclaw_loop_contracts, not ironclaw_turns), loop_host/agent_loop port-home claims, turn_runner's pre-nearai#6696 scheduler description, webui's ProductSurface path (product_contracts, not host_api), route count (93, measured), and webui's allowed-dependency list (7 of 10 were listed), the D-S await-edge ruling reflected in turn_runner guidance, composition's llm_admin residue (nearai_login_serve left for operator). Verified: cargo test -p ironclaw_architecture_tests --no-fail-fast (39 binaries, 0 failures — covers the CLI AGENTS.md phrase pin and the composition guidance-markdown scan), scripts/ci/check-target-tree.py, path-literal resolution over all 37 changed files, conflict-marker scan. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): record the closed scan evasions (F4) and the secrets-consumer correction (F2) The sealed-mint census weaknesses PROPOSAL §11.2.5/§12.1a and CHECKLIST recorded as live and owed to WS10 are all closed on this tree, verified by re-attacking the seam with both evasions at once; the docs understated the seal. Ratchet is 23 tests. One residual replaces them: the test_verified test-seam constructors, now pinned by two gates. §12.1b's 'only products-layer crate with the edge' is false by one — ironclaw_assistant carries ironclaw_secrets as port-declaration vocabulary with no expose_secret call. Not a value-reach bypass; joins nearai#7095's inventory. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): correct the app-family layer, config's consumer set, and the webui route count ironclaw_config declares layer=substrates while living in crates/app/; its consumers include operator, extension_manager and extension_host, not just the assembly crate and the binary; webui is 93 contract-locked routes, not 92 (nearai#6780 landed after the last recount). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(stale-sweep): fix agent guidance outside crates/ for the family restructure Audit-and-fix pass over every stale document outside crates/ (PR 2 of the family-guidance program). Live guidance verified against the tree; records kept with dated notes instead of rewrites. Guidance fixes (verified against HEAD before writing): - .claude/commands/trace.md: MCP tool prefix codebase-memory -> codebase-memory-mcp (allowed-tools never matched the real server), ProductSurface home -> ironclaw_product_contracts, capabilities host.rs -> host/ module split, scripts lane -> script-sandbox; deleted the redundant v1-anchors section. - .claude/commands/add-sse-event.md: deleted the banner-quarantined v1 scaffold steps (every path deleted with the monolith); now an honest redirect to the Reborn projection/SSE path. Frontmatter no longer advertises a working scaffold. - .claude/commands/deslop-reborn.md: three dead crates/*/Cargo.toml globs (family layout added a level), ls crates/ -> family-aware listing, v1-only consumer logic retired, per-crate --features integration phrasing. - .claude/rules/type-placement.md: crates/*/src globs matched nothing; recipes re-pointed and numbers re-measured 2026-08 (3,495 structs/enums, 385 traits, fan-in host_api 53 / common 20 / turns 12). - .claude/rules/skills.md: paths trigger pointed at a nonexistent bundled_skills.rs (rule never fired); SKILLS_REGEX_ACTIVATION_ENABLED / SKILLS_MAX_TOKENS env vars are read by nothing -> documented the real config-file setting and DEFAULT_MAX_SKILL_CONTEXT_TOKENS. - .claude/rules/testing.md, ironclaw-reborn-testing skill, CONTRIBUTING.md, .github/pull_request_template.md, testing-playbook, deslop: the workspace-root `integration` feature is empty with zero consumers - all "cargo test --features integration" guidance re-pointed to crate-level suites. - .claude/skills/reborn-extension-surfaces: four pre-colocation assets/ paths, CapabilitySurfaceKind home, conformance-suite move to ironclaw_extension_contracts, ingestion test move to the registry crate, gate-banned migration exemplar replaced with the live behavioral pin, [mcp] instead-of claim softened (nearai-mcp pins a static [[tools]]). - .claude/skills/ironclaw-reborn-orientation: turn_runner labels, prompt-crate list re-derived (turns/first_party_extension_ports out; host_api, loop_contracts, assistant in), consumer-grep glob fixed. - .claude/skills/reborn-feature + docs/reborn/how-to-port-channel-to-reborn.md: ProductSurface/ProductView/descriptors/caller types live in ironclaw_product_contracts; recipes re-pointed. - CLAUDE.md: dead root --features integration line replaced; project tree redrawn with the ten families; trait homes corrected; ProviderId -> VendorId; CapabilitySurfaceKind + ChannelAdapter homes; [channel.config] -> [channel.connection]/[admin_configuration]; v1 Job State Machine section deleted (no such machine in Reborn); prompt-crates recipe fixed; MCP server name; LLM backend list re-derived from LlmBackendKind. - docs/extensions/building-a-tool.md: product-adapter crates row -> channel surface model; package registration -> PACKAGES collector in ironclaw_extension_support (available_extensions.rs is being dissolved); hosted-MCP policy home -> ironclaw_extension_host/src/mcp.rs; dead v1 bullets dropped. - docs/internal/mutation-audit.md: runnable command blocks re-pointed (family paths; ironclaw_dispatcher example replaced - crate deleted in WS0). - docs/reborn/harness/e2e.md: dispatcher row -> the capabilities dispatch contract suites. docs/reborn/contracts/host-api.md: three ironclaw_dispatcher mentions -> capabilities dispatch module. standard-operations.md: renamed crate + arch-test package name. - scripts: mutation-audit.sh usage header, check-hermetic-env.sh env_helpers pointer, check-generic-without-concrete.sh mirror pointer, telegram_smoke/README regression step (target deleted with v1 in nearai#6375). - .env.example: dead SKILLS_REGEX_ACTIVATION_ENABLED entry -> config-file doc. - docs/qa/telegram-coverage-map.md: nine not-automated reasons re-worded to the crate-level integration tier. Records (dated notes, no rewrites): ADR 0003/0004 path notes (evidence pinned to their measured SHA), FEATURE_PARITY state-migration paragraph marked historical with a git-show recovery pointer, engine-v2 parity record's "coexist on main" claim corrected with a historical note, subagent-spawn legacy scope re-tensed. Pre-family path reproduction count: 73 -> 70 files; every remaining file is a dated record (docs/plans, docs/superpowers, ADRs, audits, CHANGELOG history, historical-marked train docs) or a deliberate past-tense mention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): tick row 7 — the fresh-agent placement probe passed on the final tree All three placements correct with high confidence, each naming the trait, the tests, and the tempting wrong place it rejected. The probe doubled as a docs audit and independently hit four defects, three of which the stacked guidance PR fixes — it succeeded despite them. WS12 is now 7/7. The restructure is complete. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): the product→loop_host recount was wrong on the day it was written Eight importing files across four seams, not seven across three — the fourth being a skill-activation-observer seam (projection.rs, projection/live_progress.rs) this bullet never named, which §6.4.7's own same-day note already implied. Surfaced by the plan-conformance audit. The recount history is 3→5→6→7→8, wrong at four of five attempts. That retires the prose count as a method: the sever slice should land an inventory ratchet before or with the move, not another number. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7263): CodeRabbit round-1 triage — 4 code fixes (2 sabotage-proven, 2 red-first) + 6 doc-truth corrections Code, each verified red-first or by sabotage matrix: - sealed-mint ratchet: per-name sighting floor for TEST_SEAM_MINT_FNS (closed path #12). Proven: renaming test_verified_for_tenant away plus one extra legitimate sibling mention passed the old aggregate floor (silent disarm) and fails the new per-name floor naming the constructor; suite 23/23 after revert. (CodeRabbit's claimed baseline ">2 mentions today" is wrong — each name has exactly one kept sighting — but the doc/enforcement mismatch and at-threshold fragility were real.) - trace credit: non-finite novelty_score/duplicate_score are treated as absent before clamping (clamp preserves NaN, which poisoned online_score and credit_points_estimate); NaN cases added to the nearai#7144 regression test, red first. - trace submission: a 2xx whose body stream dies mid-read now maps through request_failed (network telemetry kind, true I/O cause) instead of collapsing to an empty body that the nearai#7144 strict parse misreported as response_invalid/Submission; truncated-body regression test, red first. - Postgres contract suites (event store + assistant ledger): isolated-DB names now carry a creation epoch and the once-per-binary sweep is age-gated (1h), closing the cross-process window where a sibling's fresh zero-backend database (between CREATE DATABASE and first connection) was sweepable; legacy pid-scheme leftovers still collect immediately. Proven on live Postgres 16: planted stale name swept, planted fresh name survives, 13/13 x2 and 20/20 x2 with zero leftovers. Docs (target-architecture truth pass): - PROPOSAL section 9: the WS6 rename sweep (nearai#7152) had rewritten the source column of the 12 renamed rows to their post-rename names, turning their rename dispositions into no-ops (rows 13/14/28/30/49/51/59/61/64/66/67/70); pre-restructure names restored with a dated footnote. - PROPOSAL:69: removed the superseded 3->5->6->7 recount sentence (the corrected 3->5->6->7->8 passage subsumes it). - PROPOSAL row 34: ToolPermissionOverrideStorePort deletion marked landed (2026-08-05 WS8, matching section 6.5.3; zero workspace hits). - CHECKLIST:631: dated note recording that the WS12 F3 gsuite join landed in this batch (scenario_uninstalled_tool_call_denied_until_active.rs asserts the seeded google token on the gmail.googleapis.com wire; suite run green). - CHECKLIST:632: dated note spending F4 (the audit's 19 was correct at its SHA; the ratchet file now holds 23 tests, re-counted at lines 552/597). - ws12-gauntlet-report P6 heading: first of TWO real failures (one class), matching P8 and the report's own summary. - ws12-mapping-audit rows 49/137: dated D-S closure notes (await-edge store half = journal projection already; resolver retained loop-tier; no shed owed) so the backlog register no longer lists it as in-flight. Not fixed, with evidence: the span-helper macros gate suggestion (info_span!(target = ...) is a hard compile error, E0425 — no silent trap), the webui tracing-subscriber workspace-dep suggestion (no [workspace.dependencies] entry exists; suggestion would not build; 8 siblings use the identical direct shape), and the mapping-audit regeneration (the audit is accurate at its pinned SHA; the in-batch F1 fix is recorded in its dated coordinator note). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7263): CodeRabbit round-2 — rejection-body read keeps its cause; 200 {} is not a submission acknowledgement; lanes.md family dep rule matches measured Cargo.tomls - submission.rs non-2xx path: a failed rejection-body read no longer collapses to an empty detail via .unwrap_or_default() (banned by .claude/rules/error-handling.md); the read error folds into the http_rejection detail so the received status keeps driving the 401/403 auth-retry and the Credential/HttpRejection telemetry split. Regression: submit_preserves_rejection_body_read_failure_cause_with_status. - TraceSubmissionReceipt.status: serde default removed — it fabricated status "submitted" from a proxy's 200 {} (the nearai#7144 synthesis, resurfacing through the wire type's defaults), after which the flush caller recorded Submitted and deleted the only retryable queued copy. The acknowledgement is the server naming what happened to the submission — every workspace fixture sends status and callers persist it unconditionally as server_status — so a status-less 2xx body now fails the strict receipt parse as response_invalid. Regression: submit_rejects_success_response_without_explicit_server_status (covers 200 {} and a status-less non-empty object). - docs(lanes.md): the family Dependency-direction rule no longer claims every lane takes the extension-surface vocabulary crate — measured across crates/lanes/*/Cargo.toml: mcp + sandbox hold ironclaw_extension_contracts under [dependencies], wasm only under [dev-dependencies]; dated ✎ cross-references the ironclaw_wasm entry's 2026-08-05 correction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7263): CodeRabbit round-3 — shared Postgres test provisioner (the "new dep edge" premise measured false), entrypoint self-test armed (sabotage-proven), six doc self-contradictions reconciled Code: - ironclaw_filesystem gains a `postgres_isolation` test-support module — the single home of the per-test isolated-database scaffolding (once-per-binary age-gated stale sweep, epoch-in-name convention, DROP WITH (FORCE) cleanup), parameterised by suite/env-var/prefix/unreachable-policy. Zero new production edges: event_store already normal-deps filesystem, filesystem already owns tokio-postgres, and the dev-dep+feature pattern is the one 17 crates already use. The event-store and product-workflow-ledger suites migrate onto it; both Postgres legs proven live against postgres:16 (12 tests, zero leftover databases). The fabric original keeps its older variant with the differences documented at its IsolatedDatabase. - ironclaw_event_store drops the duplicate tokio-postgres dev-dep (the normal dep already reaches tests). - test-reborn-docker-entrypoint.sh: the missing-argv check now exits the command-substitution subshell instead of incrementing a counter the parent never sees — red-proven (a migrate-but-never-exec entrypoint passed with 7 FAIL lines printed), green after the fix both sabotaged and restored. - trace_commons submission test additionally pins !auth_rejection() for the 503 rejection (the structural assert the API affords; the prescribed payload asserts are refuted — status is private and source is None by design, with the message derived from the structured status in the same constructor). Docs (each reconciled to one canonical statement, measured): - kernel.md: lease ownership decided from code — authorization stores, matches, and expires leases (CapabilityLeaseStore + port + expiry all live there); approvals constructs and issues into that store. The round-1 re-homing of the spliced sentence into approvals was wrong and is corrected in the dated repair note. - app.md: "nothing depends on app" scoped to the three app-layer crates; ironclaw_config's consumers restated by dependency kind (normal: composition, cli, operator, extension_host; dev-only: extension_manager, root integration-tests package). - lanes.md: the mediated-services sentence now states the family law as layer-ladder + injected authority; the no-secrets/network/filesystem-dep claim is scoped to ironclaw_wasm, matching the file's own corrections. - CHECKLIST 429/430: the one open traces clause is named (ScopedFilesystem adoption); the stale "other two" count corrected against the F3a strike. - PROPOSAL:69 + CHECKLIST:72: the project-create route repointed — first_party_extension_ports dissolved into loop_host::skill_activation (WS8, §9 row 55) — still unattempted. - PROPOSAL §9 rows 57/62 synced to §6.8.4 (telegram: dependency-set equality with Slack's four contract-tier crates) and §6.9.4 (webui -> assistant is a charter-permanent edge, §12.11 D-B). - PLAN top summary records Wave 6's design question as resolved (D-S, 2026-08-05). - deploy-reborn-cli-docker.md: the two migration paragraphs unified on the entrypoint's actual behavior — only enabled = false beside signing_secret_env/bot_token_env is migrated; every other retired-key shape fails startup with the migration pointer. - composition-budget.toml: the stale "2398 bp, a true ratchet" header replaced with the WS0-floor truth the baselines test asserts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: move the guidance convention into this PR so its citations resolve families/lanes.md and families/events.md cite docs/reborn/guidance-conventions.md when superseding their 'every crate ships both an AGENTS.md and a CLAUDE.md' requirement, but the file was only on the stacked guidance branch — a forward reference that dangles if this PR merges alone. The convention is the rule those notes invoke, so it belongs with them. Caught by the CodeRabbit round-3 pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): give the hoisted postgres provisioner its safety rationales The round-3 hoist moved test provisioning into a production src/ path, so check_no_panics flagged its four panic/expect sites and reddened Code Style via fast-checks. The gate is right to flag them: it deliberately does NOT exempt #[cfg(feature = "test-support")] modules, because a cargo feature is not a privilege boundary in this workspace (PROPOSAL 12.1a proved exactly that) — so a test-support module still compiles into a build where any sibling enables the feature. Suppressed with the gate's documented inline rationale, which must trail the statement rather than precede it. The panics themselves stay: a configured but unusable Postgres must fail the suite loudly rather than skip it, which is the inert-guard rule the isolation fix exists to serve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): classify the three planner-unknown paths this PR touches The Reborn PR test planner fails closed on any unclassified path and raises on the FIRST failure in sorted order, so CI only ever showed .github/pull_request_template.md. Classifying that unmasked two more paths in this PR's own diff: scripts/mutation-audit.sh and scripts/telegram_smoke/README.md. All three are classified; the fail-closed arm is untouched: * .github/pull_request_template.md -> IGNORED_PREFIXES, beside its exact sibling .github/ISSUE_TEMPLATE/ (both GitHub UI templates; classify-test-scope.sh already pairs them in its docs-only arm). * scripts/mutation-audit.sh -> PR_STATIC_CONTROL_PATHS, beside its self-test scripts/test-mutation-audit.sh; both run only in nightly-deep-ci.yml's mutation-frontier job. * scripts/telegram_smoke/ -> QA_HARNESS_PREFIXES; a live, by-hand release smoke harness referenced by no workflow, same class as scripts/reborn_qa_matrix/. Each entry is pinned red-first in test_reborn_pr_test_plan.py (entry commented out, new assertion fails with the exact production error, entry restored, green): a new PR-template test with paired accept-AND-select-nothing assertions plus unknown-.github/-sibling refusal probes, and the two existing class tests extended. Planner self-test: 65 tests OK. The planner CLI over this PR's full 209-path diff now exits 0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
…ls, delivery heuristics deleted (nearai#7157) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): chart the notification-channel handlers in the WebUI charter map `handlers_module_charter` failed: `get_notification_channels` and `set_notification_channels` — handlers this PR adds — had no sub-owner row in `CONTRACT.md`'s enforced charter map, and the row they belong to still named `get_outbound_preferences`, `set_outbound_preferences` and `outbound_preferences_activity_id`, all deleted by this PR. Both halves are fixed together because the gate checks both in one test: unclaimed items first, then entries naming items that no longer exist. Only the first had fired, so the stale half was still latent behind it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): re-capture the loop_contracts ceiling after main's merge Main's nearai#7361/nearai#7363 landed 66 lines in `ironclaw_loop_contracts/src/instruction_bundle.rs`, which this branch picked up when folding onto main. That put the crate at 13,181 against the 13,115 ceiling re-captured earlier in this PR. Not growth from this PR. The gate's upward check is a hard `lines > ceiling` with no headroom — `TOLERANCE` (400) governs only the downward ratchet-nudge — so a ceiling captured at the exact observed value reddens every open branch the moment anyone adds a line to that crate, including from main. Re-captured at the measured value; count read from the gate's own failure message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer
pushed a commit
to l3ocifer/frick-ironclaw
that referenced
this pull request
Sep 3, 2026
nearai#7157 follow-ups) (nearai#7377) * feat: explicit channel delivery tool — two lanes, notification channels, delivery heuristics deleted Re-landed PR nearai#7157 on current main (aa8b748 → 967f33a, across the WS2 composition inversion, the WS6 crate renames, and the WS7 family moves), with all 44 review comments dispositioned. Two-lane delivery model: a run's final reply always lands in its own conversation (lane 1); reaching any other surface is the model's explicit `builtin.outbound_deliver` call (lane 2 — bot identity, one catalog target per call, synchronous through the DeliveryCoordinator, provider-issued message refs as evidence). Background-run notices fan out to a user-configured notification-channel set (new record field + read-side legacy migration, `builtin.notification_channels_set`, first-approve-wins, WebUI multi-select). The stored delivery heuristics are deleted (route_current, builtin:web_app, outbound_delivery_target_set, per-trigger delivery_target_id + precedence chains + four-slot preference fallback), with an idempotent boot migration of stored trigger targets into explicit prompt steps and the retired vocabulary pinned in reborn_retired_taxonomy.rs. Port notes (old → new homes): ironclaw_reborn_composition→app/ironclaw_composition, ironclaw_product→product/ironclaw_assistant, first_party_extensions→extensions/packages, run-profile vocabulary→ironclaw_loop_contracts, PreferenceTargetCodec→ ironclaw_extension_contracts, wire DTOs→ironclaw_product_contracts::product_wire. The model-delivery implementation moved extension_host→assistant (CoordinatedModelChannelDelivery) — the WS2 port inversion forbids extension_host naming product types; the deferred-slot registration and post-coordinator bind now live in composition's production assembly, mirroring TriggeredRunDeliveryDriver. Re-folded 2026-08-05 onto b2023bc (nearai#7258): channel-adapter vocabulary re-imported from ironclaw_extension_contracts, product-adapter/inbound vocabulary from host_api/product_contracts, module-charter map's outbound row renamed to the two-lane vocabulary. Post-branch CI gates adapted in the same change: skills/ classified in the PR test planner (test-first, sabotage-verified), panic baseline ratcheted down, nested test fixtures renamed to the scanner-sanctioned support_tests.rs shape, composition's inline trigger-migration tests split to tests.rs (mass budget green with no ceiling raise), extension_contracts size ceiling 7727 -> 7748 (+21: the ActivePreferenceTargetCodecs port), loop_contracts ceiling re-captured down 14479 -> 13850 after the delivery-vocabulary deletion. Third fold 2026-08-05 onto b72d7da (nearai#6831, standardized messaging framework): the two-lane guidance moved into the canonical messaging core prompt (host_api prompts/messaging/send_message.core.md), now naming builtin__outbound_deliver with the arrive-twice and trigger caveats for every messaging extension; slack vendor addendum/manifest taken as nearai#6831 shipped them; ceiling-table union (host_api 18570 beside this PR's two re-captures); retired slack schema embed and deleted preferences capability stay deleted; golden context-surfacing snapshot regenerated (one surface-hash line). Fourth fold 2026-08-06 onto c69ed2d (nearai#7263 program-closure batch + sibling fixes): ceiling-table union (product_contracts 15685 from nearai#7230 beside this PR's re-captures) and main's tracing-target syntax sweep (target = -> target:, gate-enforced) applied over this PR's kept lines; deleted delivery-heuristic code stays deleted. Fifth fold 2026-08-06 onto 0c297cb (nearai#7264 guidance-layer sweep): zero conflicts; guidance/doc-pointer changes auto-merged over this delta. Routing-UX slice 2026-08-06 (product thread + follow-ups): result routing is prompt-owned with a pinned source-surface default (bare "send me" = the surface you asked from; web app = no delivery step; explicit destinations override, one delivery step each) — iterated against live recordings until a real model followed it, with two live-recorded QA fixtures (bare-webui, multi-channel) plus contracts and replays. The automations-page panel is retained as the notification-channel selector (notices only); the conversational notification_channels_set tool writes the same validated set. Delivery-evidence fix (theredspoon's flag; nearai#7029 fixes the same swallow on main): mark_terminal reports whether the durable write committed and a confirmed send whose Delivered row failed to commit returns DeliveredUnconfirmed (refs retained, durably_recorded: false), never a fabricated Delivered — regression-tested and sabotage-verified. Plus a CodeRabbit triage batch: correctable coordinator errors stay model-visible, omitted target_ids no longer clears the set, the success schema requires evidence, the composition outbound facade is dissolved, and guidance/contract docs are aligned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: harden channel delivery routines and replay * fix: preserve automation loading and identity freshness * fix: close channel delivery review defects * fix: skip paused routine catch-up slots * ci: record channel delivery composition budget * test: align composition baseline with channel delivery * fix(auth): survive interrupted OAuth callbacks * fix(auth): keep callback coordination panic-free * fix(ci): reconcile channel delivery merge seams * fix(ci): recapture merged contracts ceiling * fix(delivery): close review findings across delivery, migration, and guidance Fixes the findings from the multi-agent review of this PR. Every behavioral fix ships with a regression test that fails before it. CI (red on this head) - `standalone_yolo_notification_channels_set_bypasses_approval_gate` expected the shared "invalid outbound delivery request" summary for a `builtin__notification_channels_set` call. Production deliberately specializes that message per operation and pins it with `notification_channel_failure_names_the_operation_the_model_can_correct`; the assertion was the stale side. Delivery evidence (kernel + assistant + outbound) - `AlreadyDelivered` replays reported `delivered: false` "unverified" because the ledger row retains no provider refs, inviting the duplicate resend the at-most-once claim exists to prevent. Evidence gained `already_delivered`; a replay now reads as delivered with an honest "not resent" summary. The classification suite had no `AlreadyDelivered` case at all, which is why this shipped. - `DeliveredUnconfirmed` is the one non-`Delivered` outcome that actually sent something, but `delivered_messages_from_outcome` dropped its refs, so gate reply-routes went unrecorded and a live OAuth prompt could never be retracted on that path. - `content` is now rejected when empty: the input schema advertised minLength 1 and nothing enforced it, so empty content reached the channel as an empty part and returned an opaque provider error. Background-run notifier (assistant) - One run legitimately emits several `RunBlocked` notices (re-auth stand-in, unserviceable-auth cancellation, run failure), but all three derived the same projection ref, and the delivery id hashes it. The second notice to a target came back `AlreadyDelivered`, was treated as success, and was never sent — a user could be told a routine needed re-authorization and never told it then failed. Notices carry a discriminator; once-per-run kinds keep their historical id shape, so existing delivery identities are unchanged. - When every catalog lookup failed, the empty result was recorded as `NoDefaultConfigured`, reporting a backend outage as the benign "user configured nothing" state. It now records `Failed`. Boot migration (composition) - The retired `builtin:web_app` target meant "no external delivery". It was being rewritten into a delivery step to an id nothing can resolve, inverting the stored intent on every later fire. It now clears without adding a step. - One unmigratable row aborted the entire composition boot, with the error telling the operator to shorten a prompt through the UI that no longer starts. It now pauses its own routine — a paused trigger cannot fire, so "never fire unrouted" still holds per record — and boot continues. Only a systemic store failure stays boot-fatal. A row deleted during the CAS retry ends that record instead of failing boot. - The CAS retry loop, its bounded exhaustion, and the vanished-row arm had no caller-level coverage; adds a delegating repository double that forces CAS misses. The prior fail-closed test is rewritten to pin the invariant it documented (route survives, record not half-migrated) under the new per-record mechanism. Model-visible messages (composition) - The targets-list denial said "not permitted to change the outbound delivery target" for a read-only call, and the lease denial named the retired delivery-target concept on the notification-channel path that is its only production caller. Both are now operation-specific and pinned. WebUI (frontend) - `setNotificationChannels()` with no argument posted `target_ids: []`, turning an omitted argument into a destructive clear-all and defeating the backend contract that deliberately rejects an omitted field. - The notification-channels panel stayed editable after a failed read, so toggling one row full-replaced the stored set from an empty baseline and silently dropped every channel the user never saw. Editing is now locked on a failed read, with a rendered explanation. - Adds the missing `tools.description.builtin.notification_channels_set` key to all 11 locales, plus save-failure coverage for the hook (which was correct, but untested) and locale-parity tests. Guidance - The new `.claude/rules/tools.md` was ported from a pre-restructure branch: it named `ironclaw_dispatcher` (deleted) and `ironclaw_extensions` (never existed), and its review command grepped three paths removed by WS6/WS7. Its `paths:` frontmatter also never matched the product/composition callers its rules govern, so the rule never loaded for them. - `ironclaw_loop_contracts` now records both embedded prompt assets; this PR added a second one while the crate's Known-debt entry still said one. - Bumps `skills/delegation` (rewritten guidance, unlike its two siblings in this PR which both bumped) and fixes a pre-rename path in the extension-runtime checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): unify the DM-rule enforcement point and re-ratchet composition CI (composition mass budget, red on the previous head): the per-record migration quarantine pushed composition 6 LOC over its absolute ceiling. Resolved by the reduction the budget file itself blesses rather than a raise — `runtime/approval.rs`'s 417-line inline `#[cfg(test)]` module split verbatim into `runtime/approval/tests.rs` (the gate excludes test-only files but counts inline test modules). Composition is now 40,432 LOC, smaller than before these fixes, and `loc_ceiling`/`loc_observed` plus the arch-test record are re-captured together at the measured value per the gate's one-directional ratchet rule. The codec-scan that decodes a binding and enforces "an OAuth authorization URL only ever lands in a personal DM" existed twice — once in `TriggeredReplyTargetAuthority`, once as `CodecChannelTargetResolver` — with both copies commented as "the single enforcement point". They are now one implementation, shared by the notifier and `builtin.outbound_deliver`, with a context label so each path keeps its own diagnostic. That rule turned out to be UNGUARDED: sabotaging it (`if false && ...`) failed no test in the crate. The vendor codecs pin the predicate in isolation and the coordinator test pins rejection handling with a double that decides the verdict itself, so nothing covered the wiring that joins them. Adds a contract test driving the real resolver through `DeliveryCoordinator::deliver` for both verdicts, asserting a non-DM target never reaches the vendor adapter. Sabotage-verified: the test fails with the rule disabled and passes with it restored. Smaller findings: the notification-channel schema cap now derives from `ironclaw_outbound::NOTIFICATION_TARGETS_CAP` instead of hand-mirroring `8`; `triggered_run_delivery`'s module and trait docs described the retired result-push model this PR deletes; the two new notification strings used a different brand spelling and dash style from the nine siblings in their own module; and several new comments navigated by pre-rename paths (`ironclaw_product::`, `local_dev::`, `crates/ironclaw_webui/`) plus a citation of a test symbol that does not exist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): scope the delivery catalog to the authenticated actor `builtin.outbound_deliver` resolved its destination catalog under `ResourceScope.user_id` while performing the send as `authenticated_actor_user_id`. Those are the same user on a personal thread and on an automation fire, but they diverge on a shared-route channel conversation: the scope user is the route's SUBJECT (`TurnScope::explicit_owner_user_id`) and the actor is whoever sent the message. Any participant of such a channel could therefore name the subject's target ids and push bot-identity content into the subject's own destinations — their personal DM included — from a conversation the subject may never read. The catalog now follows the actor, so a caller stays inside their own connected surfaces on every path and an unfamiliar target simply does not resolve. Behavior is unchanged wherever owner and actor already agree, which is every non-shared-route path. Regression test drives the divergent case through the port (participant denied with `TargetUnavailable`, nothing reaching a vendor adapter) plus a control proving the owner's own delivery still works. Sabotage-verified: restoring owner-scoping fails it. NOT changed here, and flagged for a product decision: the sibling `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derive their caller from the same owner-preferring `effective_user_id`, so on a shared route a participant can still enumerate — and, with the approval gate auto-approved, rewrite — the subject's notification channels. That helper also scopes approval gates and capability leases, so flipping its precedence risks breaking approval raise/resume matching in a path no test covers; it needs its own change with that coverage. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): stop rewriting the DM-target row on every message The post-admission backfill calls `FilesystemChannelDmTargetStore::upsert` for every admitted inbound direct message, and the store unconditionally wrote a fresh row. After the first message the stored record is already correct, so the steady state was one durable backend write per DM message, forever, whose only effect was a new `updated_at` — and each message's reply-delivery observation was serialized behind it. An unchanged record now short-circuits; the existing row is loaded here anyway to preserve `created_at`, so the comparison costs nothing. Also adds the regression test the `NoDefaultConfigured` -> `Failed` classification fix landed without: the notifier's `SkipEntry` lookup lane had no coverage at all (no test ever made a catalog lookup error), so neither the skip nor the all-failed arm was exercised. The triggered harness gains an injectable catalog provider for it. Sabotage-verified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(loop): bound the connected-channels line so the runtime slice fits Confirmed live, not theoretical: a worst-case runtime context renders 4,391 bytes against the 4,096-byte `PromptTextSurface::SafeSummary` cap that `instruction_bundle::push_runtime_context` validates the whole slice on — and exceeding it is a run-ending error on EVERY prompt build for that user, not a one-off. This PR's fixed ~1.1 KiB delivery-guidance block is what pushes a previously-fitting context over. The individual parts are each bounded (location 200 chars at its producer, locale 35, per-label safe-text validation), but nothing bounded their SUM, and the connected-channels line is the one part that grows without limit: up to 20 entries whose names and presentation hints are only individually capped. That line now renders as many channels as fit a 1 KiB budget and folds the rest into the "+N more" counter it already carried, so the fixed guidance can never be squeezed out by variable content. The worst-case test that found this stays as the pin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(outbound): resolve outbound capabilities as the acting user `builtin.outbound_delivery_targets_list` and `builtin.notification_channels_set` derived their caller from `effective_user_id`, which prefers the thread owner over the actor. Those agree on a direct message and on an automation fire, but diverge on a shared-route channel conversation, where the owner is the route's configured subject — the deployment operator by default (`channel_workflow.rs`) — and the actor is whoever posted. So any participant of a shared channel could enumerate the operator's connected destinations and rewrite the operator's notification-channel set, which is where approval prompts, re-auth prompts and failure notices are delivered. The caller now follows the acting user, matching the fix already applied to `builtin.outbound_deliver`. This deliberately REVERSES a previously pinned preference. Two tests asserted the owner won when the two differ; that pin predates shared-route subjects defaulting to the operator, and it contradicts the rule that a run acts as whoever invoked it. Both are updated to pin the actor, with the reversal recorded at each site rather than silently relaxed, and the notification-channel write is now asserted to land under the acting user with the thread owner's own set left untouched. INTERIM, by design: `resource_scope_for_run` and `settings_scope_for_run` still follow the owner, because they scope the approval-gate raise and the capability lease and those must stay matched between raise and resume. Unifying them belongs with the follow-up that removes shared-route subject binding entirely so a shared channel runs wholly as its invoker; that needs approval raise/resume coverage which does not exist yet. A new test pins the split so the interim state is explicit rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): keep loop_contracts under its size ceiling and ratchet down The runtime-context byte-budget fix and its worst-case pin pushed `ironclaw_loop_contracts` to 14,032 production lines against a 13,949 ceiling. Resolved by the reduction the gate prefers over a raise: `runtime_context.rs`'s 919-line inline `#[cfg(test)]` module split verbatim into a `runtime_context/tests.rs` sibling, which `production_rust_files` excludes (an inline test module inside a production file is counted; a test-only file is not). The crate now measures 13,115 — 834 lines below the previous ceiling and smaller than before this review round — so the ceiling is re-captured downward at the measured value rather than raised, per the gate's one-directional ratchet. Count read from the gate's own failure message, not by eye. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key observer gate notices by their gate ref One run can park on several approval/auth gates in sequence (the observer's blocked-state loop re-announces whenever the (status, gate) marker changes), but the live observer derived every gate notice's projection id with no discriminator, so all ApprovalNeeded notices in one run collapsed to a single durable delivery identity. The second gate's prompt came back AlreadyDelivered from the coordinator, was treated as success, and was never sent — the user was never told about the gate their run was parked on, and no reply route was recorded for it, so a bare `approve` could not resolve it either. Key the projection id by the notification's gate ref (the mechanism nearai#7157 added for the triggered notifier's RunBlocked notices). A repeat announcement of the SAME gate still dedupes; kinds that carry no gate ref (FinalReplyReady) keep the historical undiscriminated id shape so existing delivery identities are not re-keyed. Regression: observer_delivers_a_prompt_for_each_distinct_approval_gate drives the real DeliveryCoordinator over the real outbound store through two distinct scripted gates and asserts two delivered prompts plus a recorded reply route for each. Sabotage-verified: reverting the discriminator to None fails exactly this test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(composition): pin the notification-channels gate dance when owner ≠ actor The full builtin.notification_channels_set approval dance — raise, replay payload, user approve (store + lease mint from the stored row), approved resume, lease claim, dispatch, lease consume — driven through the real capability port on a run whose thread owner differs from its acting user. Pins two properties ahead of unifying the scope derivation onto the actor: raise and resume must derive the same scope (every store in the dance is scope-keyed, so a half-unified derivation strands the approved capability), and whose identity that scope carries (the thread owner, under the interim split nearai#7157 shipped). The approve step mints the lease from the stored request's own scope, grantee, and fingerprint — the same material the production click-approval resolution uses — never a re-derivation. Capability-host tier rather than tests/integration because the product rule "a run acts as its invoker" makes owner ≠ actor unconstructible through every product front door; the run-context shape remains legal kernel state (runs parked across the deploy boundary carry it). The owner == actor dance stays covered end-to-end at the integration tier (outbound_target.rs). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(outbound): scope the whole notification-channels gate dance as the acting user Unify the interim nearai#7157 split: resource_scope_for_run and settings_scope_for_run now derive from the acting user like caller_for_run already did, and effective_user_id (the owner-first ladder) is deleted. The approval-gate raise, the replay payload, the durable gate record, the lease, and the approval-settings read all follow the user who invoked the run — so the invoker sees and approves the gate, and their settings govern it. The raise/resume coverage added one commit earlier ran before and after this change and caught a real half-unification in between: the resume-side replay load lives in ironclaw_loop_host's synthetic-capability wrap (a different crate from the raise-side save in notification_channels_set) and still derived owner-first, stranding an approved resume with "replay payload is unavailable". The acting-identity ladder now has exactly one definition — LoopRunContext::acting_user_id in ironclaw_loop_contracts — and both sides delegate to it, so the hand-synced-copy class is closed rather than re-synced. notification_channels_set's replay/gate-record writes move from the capability_host-wide owner-first helper onto the outbound module's base_resource_scope_for_run so every store in one dance derives one user; the capability_host-wide helper itself is unchanged (thread/durable-result scoping legitimately follows thread ownership, and owner == actor on every binding created under the run-acts-as-invoker rule). Loop-contracts size ceiling: +16 lines for the shared ladder, paid for by splitting host/run_context.rs's 104-line inline #[cfg(test)] module into its run_context/tests.rs sibling; ceiling re-captured DOWN 13_115 -> 13_028 from the gate's own failure message. Runs raised before this change with owner != actor and resumed after it will miss their replay payload once and fail closed; re-requesting approval recovers. Documented in the PR body. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(conversations): key shared-route bindings per (conversation, actor) A run acts as the user who invoked it, so a shared conversation binds one thread per paired actor — each owned by that actor — instead of one conversation-wide thread owned by a configured subject. BindingKey gains a serde-defaulted shared_actor_user_id component (None for Direct routes, whose identity stays the conversation alone); the trusted-owner parameter is deliberately ignored on Shared creates (it remains the trigger lane's way to bind Direct conversations for their creator), and the legacy shared-owner backfill is removed with it. Migration is ignore-but-retain, pinned with a restart-path test: legacy Direct keys deserialize byte-identically (continuity), while legacy conversation-keyed shared rows deserialize to a key no per-actor lookup builds — retained in durable state untouched, and every participant (including the old subject) starts a fresh thread they own. Morphed legacy pins record what became structural: a shared probe/lookup can no longer address (or widen) a Direct binding at all; stored reply targets are isolated per actor; an actor's unpair cannot take the conversation away from other participants' own threads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(product)!: remove shared-route subject binding; scope = invoker Owner ruling: a run acts as the user who invoked it, in a DM and in a shared channel alike, with one thread per (conversation, user). This removes the subject half of shared-route configuration end to end and keeps the admission half, fail-closed: - ironclaw_product_contracts: subject_route becomes shared_admission — the SharedConversationAdmission port answers only "is this shared conversation connected"; ProductConversationRouteKey survives as the admission key. ResolvedBinding loses subject_user_id (retired-field JSON still deserializes; persisted-shape test updated); the actor is the one identity. - ironclaw_assistant: ProductInstallationScope drops the default-subject, static-route, and subject-resolver knobs for one shared_conversation_admission port; resolve/lookup/reset check admission fail-closed (no port wired, or an unlisted conversation, rejects with a not-connected BindingRequired); resolve passes no trusted owner — the conversations domain keys and owns shared bindings by the paired actor. Thread and turn scopes derive their owner from the binding's actor on every route kind. - ironclaw_extension_host: channel_subject_routes.rs becomes channel_shared_admission.rs; ChannelConfigSharedAdmission admits by membership in the operator-saved *_allowed_channels JSON array; the managed derived subject (user:{ext}-channel:{sha16}) is deleted; legacy *_subject_routes values are inert (pinned by test). Shared conversations are no longer offered as per-user notification delivery targets — their ownership came from the retired subject map — and stored channel-target preferences fail closed at resolution; DM targets are unchanged. - slack manifest: slack_shared_subject_user_id and slack_subject_routes are retired with a gravestone comment; slack_allowed_channels is the admission surface (saves to the retired handles already fail closed as unknown fields — the extension-config analog of the config.toml retired-section gravestone). - architecture tests: the INVERTED_PORTS row moves with the port rename. User-visible consequences (also in the PR body): each shared-channel participant now gets their own persistent thread and must be paired; no cross-user shared context; the operator's identity is never a fallback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(reborn): align guidance, specs, and live-QA scripts with invoker scope Guidance rows and the moved-ports list rename the inverted port (SharedConversationAdmission, ex ProductConversationSubjectRouteResolver); the assistant boundary prose states the new rule (one thread per (conversation, actor), admission is the only shared-conversation configuration, fail-closed on resolve/lookup/reset). The composition CONTRACT.md's never-shipped per-channel subject admin API section is excised with a dated correction; CHECKLIST/PROPOSAL get dated amendments beside the historical text. Operator docs teach slack_allowed_channels + per-user pairing. CHANGELOG records the behavior change and the retired config fields. The live-QA scripts drop subject handling for allowed-channels admission (200 script tests green), and the orphaned canary env var is removed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * feat(telegram): connect group chats via telegram_allowed_channels Fail-closed shared-conversation admission left Telegram groups with no operator affordance to connect one — the manifest declared no *_allowed_channels handle, so every group/supergroup @-mention was unadmittable. Declare the handle (the same generic [channel.config] convention Slack uses): listed chats are served with each participant running as themselves once paired; unlisted groups stay fail-closed. Previously any group the bot was added to ran as the deployment operator, which is the exposure this branch removes. Surfaced by the integration scenario telegram_update_becomes_a_turn_and_a_coordinated_reply failing closed after the admission change — kept red until this ruling rather than narrowed to a private chat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * test(reborn): morph the test tier to invoker scope Every fixture and pin that carried the retired subject model moves to the per-actor rule, with a recorded rationale at each semantic morph: - extension-host channel e2e: an admitted (allowed-channels) shared channel runs as the paired actor; unlisted conversations stay rejected; stored shared-channel outbound targets and binding refs fail closed; the telegram supergroup journey admits its chat via the new telegram_allowed_channels handle and proves the reply as the invoker. - assistant contract suites: admission replaces subject-route coverage (recording/failing/admit-all doubles; not-connected rejections on resolve/lookup/reset including existing bindings — a deliberate flip from the old existing-binding exemption; admission precedes actor-pairing side effects; direct routes never consult admission; per-actor threads for two participants; lookups never surface another actor's thread). - root integration harness + journeys: the binding fake, thread/turn scopes, and the group canonical user derive from the actor; multi-actor isolation pins unchanged and strictly stronger. - parity QA binary harness: subject resolution returns the actor. - webui product API redaction pin: the new telegram admission handle joins the admin-metadata forbidden list. Suites: extension_host 390/0; assistant 1084/0; conversations 105/0; architecture suite full pass; integration bins: extension_delivery 21/0 (Postgres legs under colima), delivery_user_journeys 22/0, mcp 22/0, trace_capture 14/0, generated_gate_sequences 29/0, group_journeys 16/0, group_multiuser 14/0. Workspace cargo fmt applied. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * docs(changelog): record the telegram_allowed_channels admission field Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eyxusG6UGBXBfJ8BnQDy * fix(merge): reconcile composition ceilings and capability_wiring test arity Post-merge fixups after folding main (nearai#7157 squashed + nearai#7214 + the inspector prompt-diagnostic work) into run-acts-as-invoker: - Re-capture the composition absolute-mass ceiling 40_747 -> 40_811 in both the budget manifest and reborn_restructure_baselines.rs: the acting-user scope helper and shared-admission wiring add +64 production LOC on the merged tree. Recorded rather than parked in the 150-line tolerance. - Add the 10th `tool_diagnostic_sink` argument (None) to the invoker's capability_wiring test call — main grew the signature after this branch wrote that call site. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(conversations): refuse legacy-row route-kind mismatches; drop unreachable widen A Direct request is the one key shape a retained legacy conversation-scoped shared row can collide with. Resolve, lookup, reset, and link now refuse the mismatch outright (BindingRequired) instead of trusting adapters never to re-classify a conversation's route kind — pinned by a Direct-probe leg on the legacy restart-path test. The forward half of the migration contract is pinned too: per_actor_shared_bindings_keep_their_threads_across_reopen proves a new per-actor shared binding survives a restart (a deserialize-side regression would previously have orphaned every group thread silently). widen_binding_route_access and ReplyRouteAccess::allow_shared are deleted: every Shared-keyed row is born shared under per-actor keying, so both widen call sites were unreachable. The persisted flag stays for legacy reads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(delivery): key gate notices by gate ref on the triggered lane too The gate-collapse fix shipped on the observer lane only; the background lane still minted undiscriminated projection ids for ApprovalNeeded/AuthRequired, so an automation run parking on a SECOND gate deduped to AlreadyDelivered, recorded the whole delivery Failed, and the gate was never announced or reply-routable (AGENTS.md: fix the sibling when a pattern bug is fixed). TriggeredNotification's discriminator now carries the gate ref for gate prompts (RunBlocked stand-ins compose their label with it), matching the observer keying, with a triggered two-gate regression pinning outcome, prompts, and both reply routes. Also pinned: same-gate re-announcement dedupe (g1->g2->g1), two distinct AUTH gates, and the refless id shapes incl. FinalReplyReady. Over-long discriminators are bounded with a stable FNV-1a suffix so a maximal legal TurnGateRef can never overflow ProjectionUpdateRef and silently lose a notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(identity): one contract derivation for every acting-identity scope LoopRunContext::acting_resource_scope joins acting_user_id on the contract type: the raise/resume scope recipe both gate-dance crates hand-synced is now declared once, and every surviving ladder delegates — composition's owner-first resource_scope_for_run (workspace/skill mounts) and the inline grant-minting copy, loop_host's synthetic resume load, and project_create_capability's effective_user_id (deleted; its doc claimed a mirror that no longer existed). On the only run shape where owner and actor differ — legacy runs parked across the deploy — mounts and grants now follow the ACTOR like the rest of the dance; the pin flip is recorded in visible_capability_request_uses_acting_user_for_runtime_scope. The ladder is unit-pinned in its owning crate (all three rungs) and the accepted deploy-boundary resume-miss is pinned on the synthetic port with an acting-scope positive control. loop_contracts ceiling re-captured 13094 -> 13107 with provenance (the +13-line contract method). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extension-host): collapse admission handles; operator-identity channels never admit ChannelConfigSharedAdmission now holds the one declared *_allowed_channels handle as a plain String (the scan returns Option<String>): 'installed but handle-less' is no longer representable and the per-request Option branch is gone. The root pub use of the admission items is removed — consumers are crate-local and use the module path. Structural closure of the no-auth-vendor residual: a channel whose actor identity is not per-user (no OAuth vendor, no pairing strategy) never receives an admission resolver at all — an operator-identity channel that admitted a group would run every participant as the operator, the exact exposure run-acts-as-invoker removed. Previously this was unreachable only by manifest inventory. The extension_manager wire-shape pin gains the telegram_allowed_channels row (production projection was already correct), and extension_delivery gains the caller-path rejection leg: a correctly-signed webhook for an UNLISTED supergroup is acknowledged but produces no turn and no reply. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test+docs(reborn): re-pin parity to per-actor scopes; align contract, docs, vocabulary The identity-parity bin now pins the run-acts-as-invoker property its fixture can actually express: the shared-room support binding keeps ONE thread (the per-actor thread model is locked at the integration tier by scenario_two_actors_own_threads), and inside that shared thread each RUN's scope is owned by its own invoking actor with identity context never crossing. The shared-admission suite gains the reset checkpoint leg (deny before rotation, thread survives), and the connect-nudge suite's shared leg is documented as the deliberate unpaired-participant silence contract. docs/reborn/contracts/conversation-binding.md (the owning contract) is amended: per-actor key in rule 8, participant widening retired in rule 14, subject ownership struck in rule 24, and the admission/retention semantics recorded. Operator docs and CHANGELOG state the real unpaired-shared behavior (silence; pairing via Extensions; DMs still nudge), the CHANGELOG gains the both-lanes gate-announcement entry and Added-first ordering, CHECKLIST's contradictory open-status is reconciled with a dated note, the new REBORN_WEBUI_V2_LIVE_QA_SLACK_ALLOWED_CHANNELS is threaded through live-canary.yml, observer.rs carries its arch-exempt annotation, and retired 'subject' vocabulary is renamed out of live test support and doc comments. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #7263 (base
program-closure) — review that one first; this PR's diff is docs and guidance only, no production code.The restructure gave the tree ten families and a documented boundary for each. The guidance layer never caught up:
crates/AGENTS.mdwas last refreshed 2026-07-02 — before the 13 crate renames (#7152), the 56-crate family move (#7206), and the extension-package colocation — and it routed readers tocrates/<crate>/…paths that no longer exist. There was one crate README in the entire workspace. 36 crates carried both anAGENTS.mdand aCLAUDE.mdstating overlapping rules that had quietly drifted apart. And 73 markdown files still spelled pre-family crate paths.This PR makes the answer to "what family does this belong to, and what may it hold" findable in one hop, and hard to drift again.
What lands
docs/reborn/guidance-conventions.md— the normative convention, written first so eight parallel agents produced one system rather than eight dialects. Four documents with distinct jobs; one canonical home per fact (a second file becomes a pointer, never a copy); claims measured, not aspirational; boundaries stated as exclusions ("what never belongs here, and where it goes instead"); and the warning that guidance can be gate-pinned.AGENTS.mdfor all ten families, rewritten to that shape — each with its crate table, its exclusion list, its layer-matrix row, and its enforcing gates named by test name so a reader can run them. They grew from 18–44 thin lines each to 118–215 substantive ones.README.mdfor every crate (58 family crates + the 14 extension packages), each with use-when / don't-use-when-→-go-here, measured public surface, measured dependencies and consumer counts, enforced invariants citing their gate, and test commands.crates/AGENTS.mdas a family-first routing map (264 → 175 lines; the 40-row per-crate table is gone, because family docs own crate routing now),crates/README.mdas the human map, andcrates/Architecture.mdaudited claim-by-claim..claude/,docs/, rootCLAUDE.md, and script headers.The drift this surfaced — the part worth reading
The sweep was not cosmetic. Ranked by blast radius:
.claude/rules/skills.md'spaths:trigger cited a file that does not exist (ironclaw_composition/src/extension_host/bundled_skills.rs; the real home iscrates/extensions/ironclaw_extension_host/src/bundled_skills.rs). The rule never fired for the code it governs. It also documented two env vars (SKILLS_REGEX_ACTIVATION_ENABLED,SKILLS_MAX_TOKENS) that nothing incrates/reads — the live mechanisms are a config-file key and aDEFAULT_MAX_SKILL_CONTEXT_TOKENSconstant.reborn-extension-surfacespointed authors at a gate-tripping pattern — its migration exemplar was deleted by owner ruling on 2026-08-04 and its identifiers are now on the retired-taxonomy banned list, so following the skill would red the build. Replaced with the live behavioral pin.CLAUDE.mdtaught a retired type name and two wrong crate homes —ProviderIdfor what the manifest callsVendorId, andChannelAdapter/CapabilitySurfaceKindplaced inironclaw_host_apiwhen both live inironclaw_extension_contracts. This file is loaded into every agent's context in this repo. Its "Job State Machine" section described a v1 state machine with no counterpart incrates/and is deleted..claude/rules/type-placement.md's own verification recipes returned nothing —crates/*/srcstopped matching when the family level was added, making the rule unfalsifiable. Re-pointed and re-measured (3,495 pub structs/enums, 385 pub traits; fan-inhost_api53 /common20 /turns12 —turnsandcommonhad both drifted).ironclaw_event_projections' two guidance files jointly claimed ownership of a stream manager and a write path that were deleted, which the family design forbids outright.ironclaw_loop_contracts' guidance contradicted its own armed gate — it said "host_api,common,prompt_envelope. Nothing else, ever" while the manifest holdsextension_contractsand the allowlist permits four..claude/commands/trace.md's allowed-tools named the MCP server wrongly, so no graph tool was ever pre-permitted;add-sse-event.md's six steps targeted deleted monolith files and are replaced by an honest redirect rather than an invented scaffold.Crates that had no guidance of any kind now have it:
ironclaw_prompt_envelope(the open WS11 gap),ironclaw_extension_host,ironclaw_identity,ironclaw_attachments,ironclaw_libsql_runtime,ironclaw_wasm_limiter.Method and verification
Eight Fable agents on isolated worktrees with strict, disjoint path ownership, each branched from the convention commit. Every claim was measured against the tree — consumer counts re-derived with the exact command printed in the doc, every path literal checked to resolve on disk, every cited command executed before being written down. Gate-pinned guidance was identified first (
rg -l '<file>' crates/app/ironclaw_architecture_tests/tests …) and the owning suite run where one was touched:cargo test -p ironclaw_architecture_testsexit 0 (39 binaries),cargo test -p ironclaw_llmexit 0,check-target-tree.pyexit 0,check-composition-budget.shexit 0. Two files stayed deliberately untouched because tests pin their exact phrases, including retired vocabulary that must not be "cleaned up" (ironclaw_wasm/CLAUDE.md,ironclaw_mcp/CLAUDE.md).docs/reborn/target-architecture/**was off-limits to these agents and belongs to #7263; measured contradictions found there were routed to that PR and landed as dated✎corrections — including afamilies/kernel.mdbullet that has been textually corrupted since #6918 (an approvals sentence spliced into authorization's, mid-clause).Deliberately not done
openwiki/is untouched — it is generated.CHANGELOG.md, and audit reports get dated path-notes; their decisions and pinned evidence stand. 70 files still contain pre-family paths on purpose — dated execution plans, historical specs, and evidence quoted at a fixed SHA. Forcing that number to zero would falsify the record.scripts/check-boundaries.shfails on a clean tree and four of its six checks grep the deleted v1src/, passing vacuously. Reported, not fixed — it needs its owner's call on repair vs deletion.docs/capabilities/anddocs/channels/passed the stale-pattern sweep but have not had per-claim verification; stated so rather than implied clean.🤖 Generated with Claude Code