Repository navigation
fix(loop): make routine delivery steering deterministic under progressive disclosure - #7390
Conversation
A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since #6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring #7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
🚅 Deployed to the ironclaw-pr-7390 environment in ironclaw-ci-preview
|
📝 WalkthroughSummary by CodeRabbit
WalkthroughTrigger guidance now routes named external destinations through pinned ChangesTrigger delivery
Estimated code review effort: 3 (Moderate) | ~20 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
🧭 IronLoop Run · ReviewThis comment updates in place as the Run moves through its stages. 🟩 Final result · Completed
Automatic trigger · attempt 1 of 3 · completed in 11m 33s IronLoop completed the review and posted it to GitHub. 🔗 Result |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In
`@crates/kernel/ironclaw_host_runtime/src/first_party_tools/trigger_management.rs`:
- Line 43: Create one crate-owned multiline delivery-guidance prompt fragment
under prompts/ and load it with include_str!(). In
crates/kernel/ironclaw_host_runtime/src/first_party_tools/trigger_management.rs:43-43,
replace the inline delivery-policy section in TRIGGER_CREATE_DESCRIPTION with
that fragment; in
crates/kernel/ironclaw_host_runtime/src/first_party_tools/schemas.rs:910-910,
load the same fragment and retain only the schema-specific surrounding text so
both locations share identical guidance.
In `@crates/loop/ironclaw_loop_host/src/tool_disclosure.rs`:
- Around line 58-67: Update representative_tool_fixture() to include fixtures
for outbound_deliver and outbound_delivery_targets_list, ensuring the
wide-catalog benchmark measures both newly Core-disclosed tools. Run the
benchmark and replace the recorded baseline and history values with the measured
result.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 3a5540fd-fa0d-4d85-9436-c355843b445e
📒 Files selected for processing (4)
crates/kernel/ironclaw_host_runtime/src/first_party_tools/schemas.rscrates/kernel/ironclaw_host_runtime/src/first_party_tools/trigger_management.rscrates/kernel/ironclaw_host_runtime/src/first_party_tools/trigger_management/tests.rscrates/loop/ironclaw_loop_host/src/tool_disclosure.rs
There was a problem hiding this comment.
🔍 IronLoop review
🟢 No actionable findings
No actionable findings in the four-file normal PR diff.
Validation
- ✅ Formatting — cargo fmt -p ironclaw_loop_host -p ironclaw_host_runtime -- --check
- ✅ Loop-host tests — cargo test -p ironclaw_loop_host --lib — 630 passed.
- ✅ Host-runtime tests — cargo test -p ironclaw_host_runtime --lib first_party_tools — 125 passed.
- ✅ Clippy — cargo clippy -p ironclaw_loop_host -p ironclaw_host_runtime --all-targets --all-features -- -D warnings
- ✅ Diff integrity — git diff --check refs/ironloop/merge-base refs/ironloop/head passed.
Review details
- Run:
f570b939-bef6-4032-8884-468094f1eab2 - Workflow: Review
- Attempts: 1
…ure (nearai#7390) A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since nearai#6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ure (nearai#7390) A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since nearai#6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ure (nearai#7390) A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since nearai#6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
builtin__outbound_deliverstep. This PR removes both sources of that nondeterminism.builtin.outbound_deliver+builtin.outbound_delivery_targets_listwere Discoverable, so on catalogs past the defer threshold the bridged surface (default since feat(reborn): enable progressive tool disclosure by default #6958) drops them fromvisible_capabilities— and thedelivery.mdguidance block (which contains exactly the missing rule: "'Send it to me' is bot delivery viabuiltin__outbound_deliver, not an act-as-user send") renders only while both are visible (delivery_tools_visible, loop driver host).trigger_createis Core, so a creation turn could write routines while blind on the delivery lane; whether steering existed depended on whether an earliertool_searchhappened to disclose the pair. Both tools are now Core-tier, the same reasoning the trigger lifecycle already uses.TRIGGER_CREATE_DESCRIPTIONand the trigger_createprompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". Correct for the unnamed-destination default, but over-applied — the qa_8d canary creation reasoned verbatim "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly (external-surface messages to the requester go throughbuiltin__outbound_deliver, never integration messaging tools).outbound_deliver: content+target_id;targets_list: no args).Change Type
Linked Issue
Companion to #7389 (live-canary two-lane delivery contract), which documents the live incident evidence. The same root causes explain the qa_8d marker-in-final-answer-only delivery that PR fixes on the canary side.
Validation
cargo fmt -p ironclaw_loop_host -p ironclaw_host_runtime -- --checkcargo clippy -p ironclaw_loop_host -p ironclaw_host_runtime --all-targets --all-features -- -D warningscargo test -p ironclaw_loop_host --lib— 630 passed (core-name census, disclosure caps fit, token-reduction benchmark unchanged)cargo test -p ironclaw_host_runtime --lib first_party_tools— 125 passed (new description/schema pins)cargo test -p ironclaw_architecture_tests— full suite green, including the extension-specificity gate (first draft named "slack" in generic kernel text; the gate caught it and the wording is now extension-neutral)Test Strategy
User behavior: routine creation turns deterministically receive delivery steering and pin
builtin__outbound_deliversteps for named external destinations; fire-time runs can always call the delivery pair directly.Risk areas:
Tests added or updated:
core_builtin_names_are_backed_by_known_capability_idsnow pins both outbound tools with their capability ids (exhaustive census — a name missing fromCORE_TOOL_NAMESfails);trigger_create_description_teaches_prompt_owned_delivery_with_no_stored_targetgains the scoped-clause/named-destination assertions; newtrigger_create_prompt_description_scopes_web_app_no_delivery_to_unnamed_destinationspins the schema sibling.production_default_defers_wide_catalog_to_bridge_meta_tools(tests/integration/tool_disclosure.rs) pins bridge membership and vendor-tool deferral, which this change does not alter; CI runs it.reborn-webui-v2-live-qadelivery cases (qa_3d/8d/9b/9d, no-retry) exercise creation-time pinning end-to-end every 3 hours; fix(live-qa): verify triggered Slack delivery through the two-lane contract #7389 additionally hard-fails a delivered message missing its marker.What the tests prove: the delivery pair cannot silently drop out of the Core tier; the categorical never-clause cannot come back in either guidance home; the named-destination steering exists in both.
Commands run: the five above.
Security Impact
None. Tier changes affect only which already-authorized tools are advertised versus deferred — authorization, gating (
outbound_deliverremains Allow + standing-grant as reviewed in #7157), and the capability surface policy are untouched. Guidance text changes steer the model toward the bot-identity delivery lane and away from act-as-user vendor sends for self-directed messages — a strictly safer default; the vendor lane itself remains available by design (spec decision 2, #7157).Database Impact
None.
Blast Radius
ironclaw_loop_hosttool-disclosure Core tier (advertised-token cost +2 small schemas on wide catalogs);ironclaw_host_runtimetrigger_create description + input schema text. Worst case: slightly larger advertised surface; steering text regressions are pinned by tests.Rollback Plan
Single commit;
git revertrestores Discoverable tier and the previous wording. No persisted state involved.Review Follow-Through
Review track: B
🤖 Generated with Claude Code