feat(reborn): enable progressive tool disclosure by default - #6958
Conversation
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
|
🚅 Deployed to the ironclaw-pr-6958 environment in ironclaw-ci-preview
|
|
/canary cases=qa_3b_endpoint_status_live_chat,qa_9b_routine_dm_delivery_exactly_once,qa_10a_slack_self_attribution |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (12)
📝 WalkthroughSummary by CodeRabbit
Walkthrough
ChangesTool disclosure default
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant IntegrationHarness
participant Runtime
participant ToolCatalog
IntegrationHarness->>Runtime: select production-default disclosure mode
Runtime->>ToolCatalog: submit wide GitHub catalog
ToolCatalog-->>Runtime: expose tool_search, tool_describe, and tool_call
Runtime-->>IntegrationHarness: exclude github__get_repo
Possibly related issues
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Started Reborn WebUI v2 live canary for |
🔎 Review · PR #6958
The target changed before this Run could finish. Automatic · PR opened · attempt 0 of 3 · cancelled after <1s Run details
|
|
/canary |
Live canary comparisonTargeted PR canary 30631326910 completed successfully against the production change at Compared with the matching cases from the newer main canary 30631705598 at
The prior fully successful scheduled-main run 30621834378 also totaled 11 calls across these cases, so the PR is flat versus that baseline and lower than the freshest main sample. This is a single live sample and model behavior is stochastic, so it is evidence against an obvious day-to-day regression rather than a statistical performance claim. The rollback remains |
|
/canary all |
|
Started Reborn WebUI v2 live canary for |
Coverage ratchetReborn integration-tier coverageLine coverage (Reborn crates): 85.8% — 320228 / 373238 lines Per-crate breakdown (60 crates, lowest-covered first)
This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors. Exemptions (17 entry/entries excluded from the accounting above)
|
|
/canary all |
|
Started Reborn WebUI v2 live canary for |
…-default-on # Conflicts: # tests/CLAUDE.md # tests/integration/tool_disclosure.rs
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
…ure (nearai#7390) A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since nearai#6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ntract (nearai#7389) * fix(live-qa): verify triggered Slack delivery through the two-lane contract Since nearai#7157 a triggered fire's result is never pushed by the completion driver: the fire itself calls builtin.outbound_deliver, and the background-run notifier's triggered-run-delivery record describes NOTICE delivery only — a cleanly completed fire records `skipped`. The delivery cases still required that record to say `delivered`, which no longer exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC) even though all four live fires verifiably delivered (three had the marker sitting in Slack history; the fourth was provider-confirmed). The waiter now verifies what the product actually guarantees: - success = the fire's durable outbound/deliveries model-delivery record for the exact run (delivered, expected DM) PLUS the independent Slack history read-back finding the marker; - notifier records: `skipped`/`no_default_configured`/`delivered` are healthy terminals, only `failed`/`denied` fail the case, and unknown future vocabulary surfaces through timeout diagnostics; - a completed outbound_deliver whose composed content lacks the marker fails deterministically (the qa_8d mode: the stale prompt bound the marker to the final answer, which is no longer the delivered payload); - the readback-inconclusive flake classification accepts an exactly-one-verified-send through either lane. Case prompts now bind the marker to the delivered Slack message itself (and still to the final answer), via one shared prompt-requirement helper. Also fixes the QA 6D-6E strict-scrub false positive: progressive tool disclosure (nearai#6958) records tool_search output in traces, and the builtin.extension_register_hosted_mcp description's "bearer for a static API token or PAT sent as a Bearer token" prose tripped the bearer pattern, deleting the trace and failing the shard with all cases green. The bearer pattern now requires 16+ token-alphabet characters. All delivery-wait decision logic is pinned by new unit tests against the production record shapes captured from the failing canary artifacts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(live-qa): close review findings on the two-lane delivery contract - Gate marker_deliver_count on completed previews: a failed or in-flight outbound_deliver whose content carries the marker never reached Slack, and counting it could fake the exactly-one-verified-send inconclusive classification or suppress the deterministic markerless red. Fixture gains a failed marker-bearing preview, observed red before the fix. - Align emit_results_json.py's bearer pattern with the scrub script's 16-char floor so description prose in results.json is not mangled to "Bearer <REDACTED>"; prose-preservation regression added. - Namespace the readback-inconclusive evidence per lane (vendor_evidence/deliver_evidence) — both dicts carry parse_error_count and the flat merge let one overwrite the other. - Reuse the production root_filesystem schema helpers in the new test fixtures instead of a hand-written CREATE TABLE. - Document why the deterministic content check keys on skipped/no_default_configured rather than the whole healthy_terminal class: `delivered` includes a fire parked on an approval gate whose run resumes — and may deliver — after the notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(live-qa): pin the exact 16-char bearer floor on both redaction rules The prose-preservation tests prove prose survives but not the threshold itself — a {15,} regression would have passed both. Pin the 15/16 boundary explicitly in the emitter suite and the shell scrubber suite, since the two rule sets are documented as kept in sync. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ure (nearai#7390) A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since nearai#6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ntract (nearai#7389) * fix(live-qa): verify triggered Slack delivery through the two-lane contract Since nearai#7157 a triggered fire's result is never pushed by the completion driver: the fire itself calls builtin.outbound_deliver, and the background-run notifier's triggered-run-delivery record describes NOTICE delivery only — a cleanly completed fire records `skipped`. The delivery cases still required that record to say `delivered`, which no longer exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC) even though all four live fires verifiably delivered (three had the marker sitting in Slack history; the fourth was provider-confirmed). The waiter now verifies what the product actually guarantees: - success = the fire's durable outbound/deliveries model-delivery record for the exact run (delivered, expected DM) PLUS the independent Slack history read-back finding the marker; - notifier records: `skipped`/`no_default_configured`/`delivered` are healthy terminals, only `failed`/`denied` fail the case, and unknown future vocabulary surfaces through timeout diagnostics; - a completed outbound_deliver whose composed content lacks the marker fails deterministically (the qa_8d mode: the stale prompt bound the marker to the final answer, which is no longer the delivered payload); - the readback-inconclusive flake classification accepts an exactly-one-verified-send through either lane. Case prompts now bind the marker to the delivered Slack message itself (and still to the final answer), via one shared prompt-requirement helper. Also fixes the QA 6D-6E strict-scrub false positive: progressive tool disclosure (nearai#6958) records tool_search output in traces, and the builtin.extension_register_hosted_mcp description's "bearer for a static API token or PAT sent as a Bearer token" prose tripped the bearer pattern, deleting the trace and failing the shard with all cases green. The bearer pattern now requires 16+ token-alphabet characters. All delivery-wait decision logic is pinned by new unit tests against the production record shapes captured from the failing canary artifacts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(live-qa): close review findings on the two-lane delivery contract - Gate marker_deliver_count on completed previews: a failed or in-flight outbound_deliver whose content carries the marker never reached Slack, and counting it could fake the exactly-one-verified-send inconclusive classification or suppress the deterministic markerless red. Fixture gains a failed marker-bearing preview, observed red before the fix. - Align emit_results_json.py's bearer pattern with the scrub script's 16-char floor so description prose in results.json is not mangled to "Bearer <REDACTED>"; prose-preservation regression added. - Namespace the readback-inconclusive evidence per lane (vendor_evidence/deliver_evidence) — both dicts carry parse_error_count and the flat merge let one overwrite the other. - Reuse the production root_filesystem schema helpers in the new test fixtures instead of a hand-written CREATE TABLE. - Document why the deterministic content check keys on skipped/no_default_configured rather than the whole healthy_terminal class: `delivered` includes a fire parked on an approval gate whose run resumes — and may deliver — after the notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(live-qa): pin the exact 16-char bearer floor on both redaction rules The prose-preservation tests prove prose survives but not the threshold itself — a {15,} regression would have passed both. Pin the 15/16 boundary explicitly in the emitter suite and the shell scrubber suite, since the two rule sets are documented as kept in sync. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ure (nearai#7390) A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since nearai#6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ntract (nearai#7389) * fix(live-qa): verify triggered Slack delivery through the two-lane contract Since nearai#7157 a triggered fire's result is never pushed by the completion driver: the fire itself calls builtin.outbound_deliver, and the background-run notifier's triggered-run-delivery record describes NOTICE delivery only — a cleanly completed fire records `skipped`. The delivery cases still required that record to say `delivered`, which no longer exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC) even though all four live fires verifiably delivered (three had the marker sitting in Slack history; the fourth was provider-confirmed). The waiter now verifies what the product actually guarantees: - success = the fire's durable outbound/deliveries model-delivery record for the exact run (delivered, expected DM) PLUS the independent Slack history read-back finding the marker; - notifier records: `skipped`/`no_default_configured`/`delivered` are healthy terminals, only `failed`/`denied` fail the case, and unknown future vocabulary surfaces through timeout diagnostics; - a completed outbound_deliver whose composed content lacks the marker fails deterministically (the qa_8d mode: the stale prompt bound the marker to the final answer, which is no longer the delivered payload); - the readback-inconclusive flake classification accepts an exactly-one-verified-send through either lane. Case prompts now bind the marker to the delivered Slack message itself (and still to the final answer), via one shared prompt-requirement helper. Also fixes the QA 6D-6E strict-scrub false positive: progressive tool disclosure (nearai#6958) records tool_search output in traces, and the builtin.extension_register_hosted_mcp description's "bearer for a static API token or PAT sent as a Bearer token" prose tripped the bearer pattern, deleting the trace and failing the shard with all cases green. The bearer pattern now requires 16+ token-alphabet characters. All delivery-wait decision logic is pinned by new unit tests against the production record shapes captured from the failing canary artifacts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(live-qa): close review findings on the two-lane delivery contract - Gate marker_deliver_count on completed previews: a failed or in-flight outbound_deliver whose content carries the marker never reached Slack, and counting it could fake the exactly-one-verified-send inconclusive classification or suppress the deterministic markerless red. Fixture gains a failed marker-bearing preview, observed red before the fix. - Align emit_results_json.py's bearer pattern with the scrub script's 16-char floor so description prose in results.json is not mangled to "Bearer <REDACTED>"; prose-preservation regression added. - Namespace the readback-inconclusive evidence per lane (vendor_evidence/deliver_evidence) — both dicts carry parse_error_count and the flat merge let one overwrite the other. - Reuse the production root_filesystem schema helpers in the new test fixtures instead of a hand-written CREATE TABLE. - Document why the deterministic content check keys on skipped/no_default_configured rather than the whole healthy_terminal class: `delivered` includes a fire parked on an approval gate whose run resumes — and may deliver — after the notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(live-qa): pin the exact 16-char bearer floor on both redaction rules The prose-preservation tests prove prose survives but not the threshold itself — a {15,} regression would have passed both. Pin the 15/16 boundary explicitly in the emitter suite and the shell scrubber suite, since the two rule sets are documented as kept in sync. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…p, stale events floor recapture (nearai#6966) * fix(ci): use histogram diff in the changed-coverage gate `reborn_changed_coverage.py` built its denominator from `git diff --unified=0` with no `--diff-algorithm`, so it inherited git's default (myers). Myers anchors greedily: on a deletion-shaped diff it shreds one large removal into interleaved -/+ hunks and re-emits surviving, unchanged text as added lines. The gate then demands 100% coverage for code the PR never touched. Discovered on nearai#6964 (deleting the verified-dead half of `llm::reasoning`), where the gate saw 478 changed lines / 208 branch arms and failed at 83.89% / 65.87%. Measured on that same range with the gate's own parser: myers (old): 917 changed production lines histogram (new): 14 changed production lines All 14 are doc comments and imports — zero executable — so the true changed-testable denominator was zero and the 478 was entirely artifact. The regression test asserts the invocation rather than re-staging a myers pathology: the pathology depends on git's internal heuristics, so a fixture built around one can quietly stop reproducing on a future git and leave a vacuous green test. Verified red-then-green — removing the flag fails the new case. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(deps): bump wasmtime 47.0.2 -> 47.0.3 (RUSTSEC-2026-0222, RUSTSEC-2026-0223) Two Wasmtime advisories published today fail `cargo deny check advisories` on every branch, which is the "Fast deterministic checks" red cascading into the required Code Style aggregate: RUSTSEC-2026-0222 — stores can mix up type indices between engines RUSTSEC-2026-0223 — preemption/traps during bulk operations can break internal VM state Both name `>=47.0.3` as the fix for the 47.x line. `cargo update -p wasmtime`. Cargo.lock-only. 28 packages move, every one of them on Wasmtime's own lockstep release train — wasmtime* 47.0.2 -> 47.0.3, cranelift* 0.134.2 -> 0.134.3 (its codegen backend), pulley* 47.0.2 -> 47.0.3 (its interpreter). Nothing outside that family changed; no package added or removed. Verified locally with cargo-deny 0.19.9: both advisories reproduce on the old lock and `advisories ok` after. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(coverage): recapture the stale ironclaw_events floor (inherited from nearai#6943) The ironclaw_events floor was captured before nearai#6943 deleted `events::{parse_jsonl, replay_jsonl}`, so main has been sitting under its own covered-lines floor ever since: observed 1197 covered / 1486 total against a floor of 1252 covered (effective 1232). Every branch that reaches the coverage job fails on it — PR nearai#6958, which touches no events file, fails with the byte-identical block. Recaptured to the observed numbers per coverage-floor.toml's own same-PR recapture workflow for legitimate deletion-driven shrinkage. The entry is copied byte-for-byte from nearai#6964's commit 7ca468d, which carries the same fix. Identical text on both branches means git merges them cleanly in either order. nearai#6964's ironclaw_llm entry is deliberately NOT brought along — that shrinkage is caused by that PR's own deletion and belongs to it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…(WS8 closeout) (nearai#6964) * refactor(llm): delete the verified-dead half of the reasoning module (WS8 closeout) `ironclaw_llm::reasoning` was half-live. `ironclaw_runner`'s model gateway calls `clean_response`, `contains_codex_text_tool_call_syntax`, and `recover_codex_text_tool_calls_from_tool_names` on the live model-response path; everything else in the module was a v1 engine remnant with no production consumer. PR nearai#6943 excluded the module for exactly this reason and left an enumeration to re-verify. Re-verified all 20 enumerated names against this base: the module is private (`mod reasoning;`), so its entire external surface is the two `pub use` blocks in lib.rs, and no crate in the workspace imports any of the 20. All 20 deleted. Three private helpers — `truncate_at_tool_tags`, `closing_tag_for`, `TOOL_TAG_PATTERNS` — were not enumerated but are transitively dead: all 10 of their non-test call sites were inside the deleted `impl Reasoning` block. The three live helpers are untouched, byte-for-byte, and stay where they are. Placement is deferred to the Wave 3 runner shed. Un-masking: ironclaw_llm 1000 -> 884 (-116, all reasoning-module tests belonging to deleted code, each classified); ironclaw_runner 467 -> 467 with an identical roster. No surviving test was edited — the test-module diff has zero added lines. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(checklist): tick the WS8 llm::reasoning row as landed via nearai#6964 Records the outcome on that row only: 20/20 enumerated names re-verified dead and deleted with no exclusions, three transitively-dead helpers found beyond the enumeration, the un-masking counts, and the explicit "deferred to the Wave 3 runner shed" answer to the row's placement question. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(coverage): recapture ironclaw_llm and ironclaw_events ratchet floors Both floors are the documented legitimate-shrinkage case from coverage-floor.toml's own header (a code+test deletion lowering covered lines updates the entry in the same PR). Numbers copied verbatim from this PR's coverage-report run; no tests were added to chase the old floors. ironclaw_llm — caused by this PR. Deleting the verified-dead half of `llm::reasoning` removed 1,871 instrumented lines (28,235 -> 26,364, a material -6.63% denominator move). That dead half carried denser test coverage than the crate average — 116 of the crate's tests exercised it — so removing code and tests together lowered the percentage even though no live path lost coverage. Recaptured to observed: 79.22% / 20,885 covered. ironclaw_events — inherited from main, not caused by this PR. The previous floor was captured before nearai#6943 deleted `events::{parse_jsonl, replay_jsonl}`, so main has been sitting under its own floor (-59 instrumented lines and their covered code); this branch is simply the first the ratchet caught. PR nearai#6958, which touches no events file, fails identically. Recaptured to observed: 80.55% / 1,197 covered. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Instrument canary model and tool usage * ci: unblock the queue — histogram diff gate fix, wasmtime RUSTSEC bump, stale events floor recapture (nearai#6966) * fix(ci): use histogram diff in the changed-coverage gate `reborn_changed_coverage.py` built its denominator from `git diff --unified=0` with no `--diff-algorithm`, so it inherited git's default (myers). Myers anchors greedily: on a deletion-shaped diff it shreds one large removal into interleaved -/+ hunks and re-emits surviving, unchanged text as added lines. The gate then demands 100% coverage for code the PR never touched. Discovered on nearai#6964 (deleting the verified-dead half of `llm::reasoning`), where the gate saw 478 changed lines / 208 branch arms and failed at 83.89% / 65.87%. Measured on that same range with the gate's own parser: myers (old): 917 changed production lines histogram (new): 14 changed production lines All 14 are doc comments and imports — zero executable — so the true changed-testable denominator was zero and the 478 was entirely artifact. The regression test asserts the invocation rather than re-staging a myers pathology: the pathology depends on git's internal heuristics, so a fixture built around one can quietly stop reproducing on a future git and leave a vacuous green test. Verified red-then-green — removing the flag fails the new case. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(deps): bump wasmtime 47.0.2 -> 47.0.3 (RUSTSEC-2026-0222, RUSTSEC-2026-0223) Two Wasmtime advisories published today fail `cargo deny check advisories` on every branch, which is the "Fast deterministic checks" red cascading into the required Code Style aggregate: RUSTSEC-2026-0222 — stores can mix up type indices between engines RUSTSEC-2026-0223 — preemption/traps during bulk operations can break internal VM state Both name `>=47.0.3` as the fix for the 47.x line. `cargo update -p wasmtime`. Cargo.lock-only. 28 packages move, every one of them on Wasmtime's own lockstep release train — wasmtime* 47.0.2 -> 47.0.3, cranelift* 0.134.2 -> 0.134.3 (its codegen backend), pulley* 47.0.2 -> 47.0.3 (its interpreter). Nothing outside that family changed; no package added or removed. Verified locally with cargo-deny 0.19.9: both advisories reproduce on the old lock and `advisories ok` after. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(coverage): recapture the stale ironclaw_events floor (inherited from nearai#6943) The ironclaw_events floor was captured before nearai#6943 deleted `events::{parse_jsonl, replay_jsonl}`, so main has been sitting under its own covered-lines floor ever since: observed 1197 covered / 1486 total against a floor of 1252 covered (effective 1232). Every branch that reaches the coverage job fails on it — PR nearai#6958, which touches no events file, fails with the byte-identical block. Recaptured to the observed numbers per coverage-floor.toml's own same-PR recapture workflow for legitimate deletion-driven shrinkage. The entry is copied byte-for-byte from nearai#6964's commit 7ca468d, which carries the same fix. Identical text on both branches means git merges them cleanly in either order. nearai#6964's ironclaw_llm entry is deliberately NOT brought along — that shrinkage is caused by that PR's own deletion and belongs to it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Render aggregate metrics in canary PR reports * test(reborn): allow instrumented runtime paths to settle * Map live QA harness in Reborn test planner * Preserve selected Reborn coverage mode * Ignore workspace MSRV in selected Reborn lanes --------- Co-authored-by: Benjamin Kurrek <57506486+BenKurrek@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
) * feat(reborn): enable progressive tool disclosure by default * test(reborn): pin scripted QA tool disclosure * test(reborn): pin composition tool surfaces * test(reborn): pin hook runtime tool surface * test(reborn): make flat tool fixtures explicit * test(e2e): pin flat disclosure fixtures * test(e2e): pin responses fixtures to flat tools * chore(reborn): clarify tool disclosure wording
…ure (nearai#7390) A web-created routine asking for GitHub-issue summaries "in a Slack message" was created with a stored prompt instructing the fire to use the vendor send-message tool instead of a pinned builtin__outbound_deliver step; an identical retry produced the correct pinned step. Two compounding causes, both observed live: 1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list were Discoverable-tier, so on a catalog past the defer threshold the bridged disclosure surface (default since nearai#6958) drops them from visible_capabilities — and the delivery guidance block renders only while both are visible (delivery_tools_visible). trigger_create is Core, so the model could create routines while blind on the delivery lane and without the "'Send it to me' is bot delivery via builtin__outbound_deliver" steering. Whether the steering existed depended on whether an earlier tool_search happened to disclose the pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools coupling as a deterministic fact. The wide-catalog reduction benchmark is unchanged (82.9%): its synthetic fixture carries no outbound tools. 2. The trigger_create description and its prompt-field schema said "never call builtin__outbound_deliver in a web-app-created routine". The clause is correct for the no-named-destination default, but creation turns over-apply it — the qa_8d canary creation verbatim reasoned "I'm in the web app, so there's no outbound delivery target to pin — let me use the Slack extension's tools for the send step" before recovering. Both texts now scope the no-delivery default to "no external destination named" and state the named-destination rule explicitly: reaching the user or anyone else on an external surface goes through builtin__outbound_deliver with a pinned target id, never through integration messaging tools (concrete extension names kept out per the specificity gate). Regression tests: the core-name census pins both tools with their capability ids; the description tests pin the scoped clause, the absence of the categorical never-clause, and the named-destination steering on both the tool description and the prompt-field schema. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ntract (nearai#7389) * fix(live-qa): verify triggered Slack delivery through the two-lane contract Since nearai#7157 a triggered fire's result is never pushed by the completion driver: the fire itself calls builtin.outbound_deliver, and the background-run notifier's triggered-run-delivery record describes NOTICE delivery only — a cleanly completed fire records `skipped`. The delivery cases still required that record to say `delivered`, which no longer exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC) even though all four live fires verifiably delivered (three had the marker sitting in Slack history; the fourth was provider-confirmed). The waiter now verifies what the product actually guarantees: - success = the fire's durable outbound/deliveries model-delivery record for the exact run (delivered, expected DM) PLUS the independent Slack history read-back finding the marker; - notifier records: `skipped`/`no_default_configured`/`delivered` are healthy terminals, only `failed`/`denied` fail the case, and unknown future vocabulary surfaces through timeout diagnostics; - a completed outbound_deliver whose composed content lacks the marker fails deterministically (the qa_8d mode: the stale prompt bound the marker to the final answer, which is no longer the delivered payload); - the readback-inconclusive flake classification accepts an exactly-one-verified-send through either lane. Case prompts now bind the marker to the delivered Slack message itself (and still to the final answer), via one shared prompt-requirement helper. Also fixes the QA 6D-6E strict-scrub false positive: progressive tool disclosure (nearai#6958) records tool_search output in traces, and the builtin.extension_register_hosted_mcp description's "bearer for a static API token or PAT sent as a Bearer token" prose tripped the bearer pattern, deleting the trace and failing the shard with all cases green. The bearer pattern now requires 16+ token-alphabet characters. All delivery-wait decision logic is pinned by new unit tests against the production record shapes captured from the failing canary artifacts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(live-qa): close review findings on the two-lane delivery contract - Gate marker_deliver_count on completed previews: a failed or in-flight outbound_deliver whose content carries the marker never reached Slack, and counting it could fake the exactly-one-verified-send inconclusive classification or suppress the deterministic markerless red. Fixture gains a failed marker-bearing preview, observed red before the fix. - Align emit_results_json.py's bearer pattern with the scrub script's 16-char floor so description prose in results.json is not mangled to "Bearer <REDACTED>"; prose-preservation regression added. - Namespace the readback-inconclusive evidence per lane (vendor_evidence/deliver_evidence) — both dicts carry parse_error_count and the flat merge let one overwrite the other. - Reuse the production root_filesystem schema helpers in the new test fixtures instead of a hand-written CREATE TABLE. - Document why the deterministic content check keys on skipped/no_default_configured rather than the whole healthy_terminal class: `delivered` includes a fire parked on an approval gate whose run resumes — and may deliver — after the notice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(live-qa): pin the exact 16-char bearer floor on both redaction rules The prose-preservation tests prove prose survives but not the threshold itself — a {15,} regression would have passed both. Pin the 15/16 boundary explicitly in the emitter suite and the shell scrubber suite, since the two rule sets are documented as kept in sync. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
REBORN_TOOL_DISCLOSUREis unset or emptyREBORN_TOOL_DISCLOSURE=offand invalid/non-Unicode values as fail-closed rollback pathsOffmain, port the mode change to its currentironclaw_loop_hostowner, and inherit ci: unblock the queue — histogram diff gate fix, wasmtime RUSTSEC bump, stale events floor recapture #6966's legitimateironclaw_eventscoverage-floor recaptureChange Type
Linked Issue
Closes #6810
Follows #5659 and #6859.
Validation
cargo fmt --all -- --checkcargo clippy --all --benches --tests --examples --all-features -- -D warnings— not run locally; the focused test builds passed and PR CI will run the repository clippy matrix.cargo build— not run separately; the owning-package and integration tests compiled the changed paths.ironclaw_loop_hosttests (553 unit tests plus contract suites), all 27reborn_integration_tool_disclosuretests, and the fullironclaw_reborn_compositionpackage suite (one pre-existing TODO test and one doctest ignored).cargo test --features integrationif database-backed or integration behavior changed — Not applicable: no database or backend behavior changed./canaryon this PR.review-prorpr-shepherd --fixwas run before requesting review — the post-merge diff againstmainwas manually audited and the owner, whole-turn, and composition suites were run.The branch is merged with current
main. The earlier aggregate coverage failure was an upstreamironclaw_eventscode-and-test deletion already recaptured by #6966; this PR inherits that fix and does not change the coverage floor itself.Test Strategy
User behavior:
Users with wide capability catalogs get the discovery bridge by default. Small catalogs remain direct. Operators can immediately restore the flat surface with
REBORN_TOOL_DISCLOSURE=off.Risk areas:
Tests added or updated:
ironclaw_loop_hostproves unset and empty configuration resolve toBridged; invalid, non-Unicode, and explicitoffresolve toOff.tool_search,tool_describe, andtool_callwhile deferring the flat GitHub list. The explicit-Off and hermetic-default controls remain green.ironclaw_loop_hostandironclaw_reborn_compositionpackage suites plus the production capability-chain disclosure integration target.What the tests prove:
OffCommands run:
cargo fmt --all -- --checkcargo test -p ironclaw_loop_hostcargo test -p ironclaw_reborn_integration_tests --test reborn_integration_tool_disclosurecargo test -p ironclaw_reborn_compositionSecurity Impact
The model-visible default changes from a flat catalog to threshold-gated progressive disclosure. Authorization, capability allow-set filtering, approvals, target registration, and mediated dispatch remain unchanged. #5659 already landed the narrowed-surface metadata-isolation prerequisite. Invalid/unreadable configuration and explicit
offcontinue to fail closed.Reborn Trust-Boundary Checklist
ToolDisclosureMode's default only was changed.serde(default)fields fail closed or have migration tests: no serialized fields changed; invalid environment configuration remainsOff.Transient,Permanent,Misconfigured,PolicyDeniedor equivalent): no error class changed.Database Impact
None.
Blast Radius
Reborn model requests whose effective catalog exceeds 32 tools or 12k estimated schema tokens now receive the discovery bridge by default. Below-threshold surfaces are unchanged. The primary regression risk is extra model calls or incorrect discovery behavior on weaker tool-calling models.
Rollback Plan
Set
REBORN_TOOL_DISCLOSURE=offto restore the flat catalog without a deploy. Revert this commit to restore default-Off configuration semantics. No migration or persisted-state rollback is required.Review Follow-Through
/canaryresult and a per-case tool-call comparison against the latest successful scheduled canary.offuntil canary evidence is reviewed; remove it through bounded cohorts afterward.Review track: C (runtime default change)