Skip to content

feat(reborn): enable progressive tool disclosure by default - #6958

Merged
serrrfirat merged 12 commits into
mainfrom
codex/tool-disclosure-default-on
Aug 7, 2026
Merged

serrrfirat merged 12 commits into
mainfrom
codex/tool-disclosure-default-on

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 31, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Closes #6810

Follows #5659 and #6859.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — not run locally; the focused test builds passed and PR CI will run the repository clippy matrix.
  • cargo build — not run separately; the owning-package and integration tests compiled the changed paths.
  • Relevant tests pass: all ironclaw_loop_host tests (553 unit tests plus contract suites), all 27 reborn_integration_tool_disclosure tests, and the full ironclaw_reborn_composition package suite (one pre-existing TODO test and one doctest ignored).
  • cargo test --features integration if database-backed or integration behavior changed — Not applicable: no database or backend behavior changed.
  • Manual testing: Not applicable before PR canary; live verification will be posted by /canary on this PR.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review — the post-merge diff against main was manually audited and the owner, whole-turn, and composition suites were run.

The branch is merged with current main. The earlier aggregate coverage failure was an upstream ironclaw_events code-and-test deletion already recaptured by #6966; this PR inherits that fix and does not change the coverage floor itself.

Test Strategy

User behavior:

Users with wide capability catalogs get the discovery bridge by default. Small catalogs remain direct. Operators can immediately restore the flat surface with REBORN_TOOL_DISCLOSURE=off.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: ironclaw_loop_host proves unset and empty configuration resolve to Bridged; invalid, non-Unicode, and explicit off resolve to Off.
  • Reborn integration: a wide catalog under the production enum default advertises tool_search, tool_describe, and tool_call while deferring the flat GitHub list. The explicit-Off and hermetic-default controls remain green.
  • Recorded fixture: Not applicable: deterministic surface selection is covered below the vendor seam.
  • Browser E2E: Not applicable: no browser contract changes.
  • Backend or runtime: ironclaw_loop_host and ironclaw_reborn_composition package suites plus the production capability-chain disclosure integration target.
  • Live canary: completed on the pre-merge product head; the full run passed 11/12 shards, and repeated tool-call samples showed enough variance that no cost-reduction claim is made.

What the tests prove:

  • unset/empty production configuration selects bridged disclosure
  • invalid configuration and the explicit kill switch stay fail-closed
  • wide catalogs expose the complete discovery protocol by default
  • general integration harnesses remain hermetically pinned to Off

Commands run:

  • cargo fmt --all -- --check
  • cargo test -p ironclaw_loop_host
  • cargo test -p ironclaw_reborn_integration_tests --test reborn_integration_tool_disclosure
  • cargo test -p ironclaw_reborn_composition

Security Impact

The model-visible default changes from a flat catalog to threshold-gated progressive disclosure. Authorization, capability allow-set filtering, approvals, target registration, and mediated dispatch remain unchanged. #5659 already landed the narrowed-surface metadata-isolation prerequisite. Invalid/unreadable configuration and explicit off continue to fail closed.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: no constructors or trust-bearing types changed.
  • Untrusted content enters prompts only through an envelope/escaping primitive: no prompt content or untrusted input handling changed.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check: Not applicable; no hash behavior changed.
  • New/changed status, exit, policy, runtime, or error variants: downstream match sites audited. Command/output: no variants changed; ToolDisclosureMode's default only was changed.
  • Security/durability serde(default) fields fail closed or have migration tests: no serialized fields changed; invalid environment configuration remains Off.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: no queue/map/buffer/counter changed; the existing 32-tool / 12k-token disclosure budget remains enforced.
  • Driver/operator-visible errors have stable class semantics (Transient, Permanent, Misconfigured, PolicyDenied or equivalent): no error class changed.
  • Sandbox/native/host names accurately describe trust boundary: no runtime-lane or sandbox names changed.

Database Impact

None.

Blast Radius

Reborn model requests whose effective catalog exceeds 32 tools or 12k estimated schema tokens now receive the discovery bridge by default. Below-threshold surfaces are unchanged. The primary regression risk is extra model calls or incorrect discovery behavior on weaker tool-calling models.

Rollback Plan

Set REBORN_TOOL_DISCLOSURE=off to restore the flat catalog without a deploy. Revert this commit to restore default-Off configuration semantics. No migration or persisted-state rollback is required.

Review Follow-Through

  • Post the exact-head /canary result and a per-case tool-call comparison against the latest successful scheduled canary.
  • Keep the production deployment override at off until canary evidence is reviewed; remove it through bounded cohorts afterward.

Review track: C (runtime default change)

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 July 31, 2026 12:37 Destroyed
@railway-app

railway-app Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6958 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ❌ Build Failed (View Logs) Web Aug 7, 2026 at 2:42 pm

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/canary cases=qa_3b_endpoint_status_live_chat,qa_9b_routine_dm_delivery_exactly_once,qa_10a_slack_self_attribution

@coderabbitai

coderabbitai Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 935c4cfb-aa29-4e86-885a-bb85f64d577f

📥 Commits

Reviewing files that changed from the base of the PR and between 81724a6 and 34431dc.

📒 Files selected for processing (12)
  • crates/app/ironclaw_composition/src/runtime/tests/core.rs
  • crates/app/ironclaw_composition/src/runtime/tests/outbound_delivery.rs
  • crates/app/ironclaw_composition/tests/runtime.rs
  • crates/app/ironclaw_composition/tests/webui_v2_e2e.rs
  • crates/loop/ironclaw_loop_host/src/tool_disclosure_mode.rs
  • tests/CLAUDE.md
  • tests/e2e/reborn_webui_harness.py
  • tests/e2e/scenarios/test_reborn_responses_api.py
  • tests/integration/support/builder.rs
  • tests/integration/tool_disclosure.rs
  • tests/reborn_qa_routines.rs
  • tests/support/reborn_parity_qa/binary_e2e.rs

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Progressive tool disclosure is now enabled by default, making large tool catalogs easier to navigate.
    • Large GitHub tool catalogs can be accessed through tool_search, tool_describe, and tool_call.
  • Bug Fixes

    • Explicitly disabling tool disclosure remains supported.
    • Invalid configuration values continue to fail safely with disclosure disabled.
  • Tests

    • Expanded coverage validates default, bridged, and disabled tool-disclosure behavior.

Walkthrough

ToolDisclosureMode now defaults to Bridged when unset or empty. Integration coverage verifies bridge tools for wide catalogs. Existing test harnesses explicitly select Off to preserve flat tool surfaces.

Changes

Tool disclosure default

Layer / File(s) Summary
Runtime default and parsing semantics
crates/loop/ironclaw_loop_host/src/tool_disclosure_mode.rs
Unset and empty configuration now resolve to Bridged. Explicit off and invalid values resolve to Off.
Production-default integration coverage
tests/integration/support/builder.rs, tests/integration/tool_disclosure.rs, tests/CLAUDE.md
The integration builder selects the production default. Coverage verifies tool_search, tool_describe, and tool_call, while excluding github__get_repo.
Backward-compatible test pinning
tests/reborn_qa_routines.rs, tests/support/reborn_parity_qa/binary_e2e.rs, crates/app/ironclaw_composition/src/runtime/tests/*, crates/app/ironclaw_composition/tests/*, tests/e2e/*
Existing QA, replay, composition, runtime, WebUI, and Responses API tests explicitly disable tool disclosure.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant IntegrationHarness
  participant Runtime
  participant ToolCatalog
  IntegrationHarness->>Runtime: select production-default disclosure mode
  Runtime->>ToolCatalog: submit wide GitHub catalog
  ToolCatalog-->>Runtime: expose tool_search, tool_describe, and tool_call
  Runtime-->>IntegrationHarness: exclude github__get_repo
Loading

Possibly related issues

  • Issue 7166 — Covers the same progressive disclosure default and explicit off rollback behavior.

Possibly related PRs

Suggested reviewers: benkurrek

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The change implements default-on behavior, fail-closed rollback, threshold handling, and coverage, but lacks evidence for several requirements in #6810. Add or link rollout metrics, comparison tooling, deterministic recall and quality validation, and health-gate evidence, or track them explicitly as follow-up work for #6810.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title uses Conventional Commits format and accurately states the main change: enabling progressive tool disclosure by default.
Description check ✅ Passed The description follows the template and documents scope, linked issues, validation, tests, security impact, blast radius, rollback, and follow-through.
Out of Scope Changes check ✅ Passed The test harness pinning, production-default coverage, runtime changes, and documentation updates directly support the tool-disclosure objectives in #6810.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: S 10-49 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 31, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Started Reborn WebUI v2 live canary for codex/tool-disclosure-default-on at 4bc3c01ed6 with cases qa_3b_endpoint_status_live_chat,qa_9b_routine_dm_delivery_exactly_once,qa_10a_slack_self_attribution: https://github.com/nearai/ironclaw/actions/runs/30631326910

@ironloopai

ironloopai Bot commented Jul 31, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6958

⚫ Cancelled · Target changed

The target changed before this Run could finish.

Automatic · PR opened · attempt 0 of 3 · cancelled after <1s

Run details
  • Repository: nearai/ironclaw
  • Base: main at 73d45cc
  • Head: codex/tool-disclosure-default-on at 4bc3c01
  • Created: Jul 31, 2026, 12:42 PM UTC
  • Updated: Jul 31, 2026, 12:42 PM UTC
  • Run: f608dfa0-0724-4e13-a8a2-23d3bacfcd61

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/canary

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 July 31, 2026 12:56 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 July 31, 2026 13:31 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Live canary comparison

Targeted PR canary 30631326910 completed successfully against the production change at 4bc3c01ed6. The current PR head adds test-only explicit-mode pins; it does not change the shipping binary exercised by this run.

Compared with the matching cases from the newer main canary 30631705598 at a88bcb2522:

Case Latest main PR default-on Delta Result
qa_3b_endpoint_status_live_chat 1 1 0 Passed; builtin.http in both
qa_9b_routine_dm_delivery_exactly_once 8 7 -1 Passed; PR used 2 ironclaw.tool_describe calls and still delivered exactly once
qa_10a_slack_self_attribution 4 3 -1 Passed; PR used 1 ironclaw.tool_describe call
Total 13 11 -2 (-15.4%) All matched cases passed

The prior fully successful scheduled-main run 30621834378 also totaled 11 calls across these cases, so the PR is flat versus that baseline and lower than the freshest main sample.

This is a single live sample and model behavior is stochastic, so it is evidence against an obvious day-to-day regression rather than a statistical performance claim. The rollback remains REBORN_TOOL_DISCLOSURE=off.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 July 31, 2026 13:46 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 July 31, 2026 13:50 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/canary all

@github-actions

Copy link
Copy Markdown
Contributor

Started Reborn WebUI v2 live canary for codex/tool-disclosure-default-on at 5886468f06 with cases all: https://github.com/nearai/ironclaw/actions/runs/30636283879

@github-actions

github-actions Bot commented Jul 31, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.8% (320228 / 373238 lines)
  floor:    85.11% (tolerance 0.5pp -> effective floor 84.61%)
  denominator: 373238 lines now vs 375097 at floor capture (-1859 lines, -0.5%) — not a material change

RATCHET PASS: ironclaw_runner
  observed: 86% (15127 / 17590 lines)
  floor:    85.55% (tolerance 0.5pp -> effective floor 85.05%)
  floor_covered_lines: 14658 (tolerance 20 lines -> effective floor 14638)
  denominator: 17590 lines now vs 17133 at floor capture (+457 lines, +2.67%) — not a material change

RATCHET PASS: ironclaw_processes
  observed: 88.79% (5891 / 6635 lines)
  floor:    88.07% (tolerance 0.5pp -> effective floor 87.57%)
  floor_covered_lines: 5839 (tolerance 20 lines -> effective floor 5819)
  denominator: 6635 lines now vs 6630 at floor capture (+5 lines, +0.08%) — not a material change

RATCHET PASS: ironclaw_turns
  observed: 86.54% (9863 / 11397 lines)
  floor:    85.11% (tolerance 0.5pp -> effective floor 84.61%)
  floor_covered_lines: 9515 (tolerance 20 lines -> effective floor 9495)
  denominator: 11397 lines now vs 11179 at floor capture (+218 lines, +1.95%) — not a material change

RATCHET PASS: ironclaw_authorization
  observed: 86.59% (723 / 835 lines)
  floor:    62.51% (tolerance 0.5pp -> effective floor 62.01%)
  floor_covered_lines: 612 (tolerance 20 lines -> effective floor 592)
  denominator: 835 lines now vs 979 at floor capture (-144 lines, -14.71%) — material change (>5%)

RATCHET PASS: ironclaw_approvals
  observed: 91.05% (1820 / 1999 lines)
  floor:    85.86% (tolerance 0.5pp -> effective floor 85.36%)
  floor_covered_lines: 1822 (tolerance 20 lines -> effective floor 1802)
  denominator: 1999 lines now vs 2122 at floor capture (-123 lines, -5.8%) — material change (>5%)

RATCHET PASS: ironclaw_secrets
  observed: 85.77% (2887 / 3366 lines)
  floor:    84.01% (tolerance 0.5pp -> effective floor 83.51%)
  floor_covered_lines: 2795 (tolerance 20 lines -> effective floor 2775)
  denominator: 3366 lines now vs 3327 at floor capture (+39 lines, +1.17%) — not a material change

RATCHET PASS: ironclaw_filesystem
  observed: 76.23% (5849 / 7673 lines)
  floor:    75.93% (tolerance 0.5pp -> effective floor 75.43%)
  floor_covered_lines: 5826 (tolerance 20 lines -> effective floor 5806)
  denominator: 7673 lines now vs 7673 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_llm
  observed: 79.96% (22644 / 28318 lines)
  floor:    79.92% (tolerance 0.5pp -> effective floor 79.42%)
  floor_covered_lines: 22566 (tolerance 20 lines -> effective floor 22546)
  denominator: 28318 lines now vs 28235 at floor capture (+83 lines, +0.29%) — not a material change

RATCHET PASS: ironclaw_triggers
  observed: 94.88% (3092 / 3259 lines)
  floor:    86.04% (tolerance 0.5pp -> effective floor 85.54%)
  floor_covered_lines: 2804 (tolerance 20 lines -> effective floor 2784)
  denominator: 3259 lines now vs 3259 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_product
  observed: 87.49% (22827 / 26090 lines)
  floor:    86.94% (tolerance 0.5pp -> effective floor 86.44%)
  floor_covered_lines: 21367 (tolerance 20 lines -> effective floor 21347)
  denominator: 26090 lines now vs 24576 at floor capture (+1514 lines, +6.16%) — material change (>5%)

RATCHET PASS: ironclaw_outbound
  observed: 94.68% (4271 / 4511 lines)
  floor:    93.49% (tolerance 0.5pp -> effective floor 92.99%)
  floor_covered_lines: 4105 (tolerance 20 lines -> effective floor 4085)
  denominator: 4511 lines now vs 4391 at floor capture (+120 lines, +2.73%) — not a material change

RATCHET PASS: ironclaw_extension_host
  observed: 83.82% (22271 / 26569 lines)
  floor:    83.82% (tolerance 0.5pp -> effective floor 83.32%)
  floor_covered_lines: 22271 (tolerance 20 lines -> effective floor 22251)
  denominator: 26569 lines now vs 26569 at floor capture (+0 lines, +0%) — not a material change

RATCHET FAIL: ironclaw_events
  observed: 80.55% (1197 / 1486 lines)
  floor:    81.04% (tolerance 0.5pp -> effective floor 80.54%)
  floor_covered_lines: 1252 (tolerance 20 lines -> effective floor 1232)
  denominator: 1486 lines now vs 1545 at floor capture (-59 lines, -3.82%) — not a material change
  To fix:
    - If this is a real coverage regression: add tests, don't touch the floor file.
    - If this is a legitimate denominator/numerator shift EITHER direction —
      growth (new/renamed module entering instrumentation, e.g. #5656) OR
      shrinkage (a code+test deletion legitimately lowering covered lines,
      even when the denominator barely moves): update
      tests/integration/coverage-floor.toml's [[crate]] entry for
      ironclaw_events IN THIS PR — bump (or lower) captured_total_lines and
      floor_percent/floor_covered_lines to the new observed numbers, set
      captured_date, and add a one-line rationale + issue link.
      See the file's own header for the schema.

RATCHET PASS: ironclaw_safety
  observed: 92.75% (4468 / 4817 lines)
  floor:    92.44% (tolerance 0.5pp -> effective floor 91.94%)
  floor_covered_lines: 3973 (tolerance 20 lines -> effective floor 3953)
  denominator: 4817 lines now vs 4298 at floor capture (+519 lines, +12.08%) — material change (>5%)

RATCHET PASS: ironclaw_host_runtime
  observed: 88.38% (21083 / 23855 lines)
  floor:    88.23% (tolerance 0.5pp -> effective floor 87.73%)
  floor_covered_lines: 20538 (tolerance 20 lines -> effective floor 20518)
  denominator: 23855 lines now vs 23277 at floor capture (+578 lines, +2.48%) — not a material change

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.8% — 320228 / 373238 lines

Per-crate breakdown (60 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_host_ingress 42.5% 17 / 40
ironclaw_memory 53.48% 630 / 1178
ironclaw_projects 72.36% 233 / 322
ironclaw_trust 73.71% 670 / 909
ironclaw_capabilities 74.7% 2884 / 3861
ironclaw_extractors 75.88% 538 / 709
ironclaw_reborn_cli 76.1% 11084 / 14566
ironclaw_observability 76.19% 32 / 42
ironclaw_filesystem 76.23% 5849 / 7673
ironclaw_wasm 78.84% 704 / 893
ironclaw_llm 79.96% 22644 / 28318
ironclaw_events 80.55% 1197 / 1486
ironclaw_auth 82.03% 5960 / 7266
ironclaw_first_party_extensions 82.57% 6784 / 8216
ironclaw_memory_native 82.85% 2850 / 3440
ironclaw_libsql_runtime 83.3% 384 / 461
ironclaw_host_api 83.42% 9718 / 11649
ironclaw_extension_host 83.82% 22271 / 26569
ironclaw_operator 84.47% 5309 / 6285
ironclaw_hooks 84.58% 9906 / 11712
ironclaw_event_projections 84.81% 854 / 1007
ironclaw_network 84.92% 890 / 1048
ironclaw_reborn_event_store 84.93% 1206 / 1420
ironclaw_reborn_config 85.29% 2110 / 2474
ironclaw_reborn_composition 85.51% 21781 / 25473
ironclaw_secrets 85.77% 2887 / 3366
ironclaw_runner 86% 15127 / 17590
ironclaw_turns 86.54% 9863 / 11397
ironclaw_authorization 86.59% 723 / 835
ironclaw_webui 86.93% 11821 / 13598
ironclaw_wasm_limiter 87.06% 74 / 85
ironclaw_common 87.33% 1641 / 1879
ironclaw_product 87.49% 22827 / 26090
ironclaw_reborn_traces 87.61% 11720 / 13377
ironclaw_extensions 87.84% 4870 / 5544
ironclaw_scripts 87.87% 420 / 478
ironclaw_skills 88.13% 2770 / 3143
ironclaw_threads 88.14% 5189 / 5887
ironclaw_host_runtime 88.38% 21083 / 23855
ironclaw_telegram_extension 88.52% 586 / 662
ironclaw_process_sandbox 88.64% 281 / 317
ironclaw_processes 88.79% 5891 / 6635
ironclaw_reborn_openai_compat 89.4% 3644 / 4076
ironclaw_telegram_v2_adapter 89.43% 1573 / 1759
ironclaw_loop_host 90.47% 18045 / 19946
ironclaw_resources 90.76% 4084 / 4500
ironclaw_approvals 91.05% 1820 / 1999
ironclaw_reborn_identity 91.3% 451 / 494
ironclaw_mcp 91.91% 1318 / 1434
ironclaw_conversations 92.08% 2383 / 2588
ironclaw_event_streams 92.5% 1048 / 1133
ironclaw_safety 92.75% 4468 / 4817
ironclaw_agent_loop 93.52% 10428 / 11151
ironclaw_slack_extension 93.94% 3689 / 3927
ironclaw_first_party_extension_ports 94.66% 3758 / 3970
ironclaw_outbound 94.68% 4271 / 4511
ironclaw_triggers 94.88% 3092 / 3259
ironclaw_prompt_envelope 97.46% 192 / 197
ironclaw_runtime_policy 97.6% 855 / 876
ironclaw_attachments 98.23% 831 / 846

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (17 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657
crates/ironclaw_attachments/src/lib.rs Declarative crate facade: module declarations, constants, and re-exports only; executable attachment modules remain covered. #6524
crates/ironclaw_extension_host/src/ingress/mod.rs Declarative ingress module facade and documentation only; executable router modules remain covered. #6524
crates/ironclaw_host_api/src/lib.rs Declarative crate facade: module declarations and re-exports only; executable host API modules remain covered. #6524
crates/ironclaw_host_api/src/product_adapter/mod.rs Declarative product-adapter facade: module declarations and re-exports only; executable adapter modules remain covered. #6524
crates/ironclaw_llm/src/rig_adapter/tests/finish_reason_tests.rs Test-only module stored under src/ for private adapter access; cargo-llvm-cov omits test harness source from production LCOV while the exercised rig_adapter.rs production lines remain coverage-gated. #6284
crates/ironclaw_outbound/src/error.rs Declarative error vocabulary only; variants have no LLVM-instrumentable production statements. #6524
crates/ironclaw_outbound/src/lib.rs Declarative crate facade: module declarations and re-exports only; executable outbound modules remain covered. #6524
crates/ironclaw_product/src/lib.rs Declaration-only public facade with no executable Rust statements; rustc emits no LCOV source record. Executable product behavior remains covered in the owned implementation modules. #6524
crates/ironclaw_product/src/lib.rs Declarative crate facade: module declarations and re-exports only; executable product modules remain covered. #6524
crates/ironclaw_product/src/scoped_fs/mod.rs Declarative scoped-filesystem facade and documentation only; executable scoped filesystem modules remain covered. #6524
crates/ironclaw_reborn_composition/src/support/fs/mod.rs Declarative composition support facade: module declarations and re-exports only; executable filesystem adapters remain covered. #6524
crates/ironclaw_slack_extension/src/lib.rs Declarative Slack crate facade: module declarations and re-exports only; executable Slack modules remain covered. #6524
crates/ironclaw_telegram_extension/src/lib.rs Declarative Telegram crate facade: module declarations and re-exports only; executable Telegram modules remain covered. #6524
crates/ironclaw_threads/src/lib.rs Declaration-only public facade with no executable Rust statements; rustc emits no LCOV source record. Executable thread behavior remains covered in the owned implementation modules. #6524
crates/ironclaw_webui/src/webui_v2/mod.rs Declaration-only WebUI v2 facade with no executable Rust statements; rustc emits no LCOV source record. Executable route behavior remains covered in the owned implementation modules. #6524

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 July 31, 2026 14:36 Destroyed
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

/canary all

@github-actions

Copy link
Copy Markdown
Contributor

Started Reborn WebUI v2 live canary for codex/tool-disclosure-default-on at ace0caeabf with cases all: https://github.com/nearai/ironclaw/actions/runs/30639369563

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 July 31, 2026 15:05 Destroyed
…-default-on

# Conflicts:
#	tests/CLAUDE.md
#	tests/integration/tool_disclosure.rs
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6958 August 7, 2026 14:29 Destroyed
@coderabbitai

coderabbitai Bot commented Aug 7, 2026

Copy link
Copy Markdown

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@serrrfirat
serrrfirat added this pull request to the merge queue Aug 7, 2026
Merged via the queue into main with commit e6a650f Aug 7, 2026
44 of 45 checks passed
@serrrfirat
serrrfirat deleted the codex/tool-disclosure-default-on branch August 7, 2026 17:40
getong pushed a commit to getong/ironclaw that referenced this pull request Aug 8, 2026
…ure (nearai#7390)

A web-created routine asking for GitHub-issue summaries "in a Slack
message" was created with a stored prompt instructing the fire to use
the vendor send-message tool instead of a pinned
builtin__outbound_deliver step; an identical retry produced the correct
pinned step. Two compounding causes, both observed live:

1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list
   were Discoverable-tier, so on a catalog past the defer threshold the
   bridged disclosure surface (default since nearai#6958) drops them from
   visible_capabilities — and the delivery guidance block renders only
   while both are visible (delivery_tools_visible). trigger_create is
   Core, so the model could create routines while blind on the delivery
   lane and without the "'Send it to me' is bot delivery via
   builtin__outbound_deliver" steering. Whether the steering existed
   depended on whether an earlier tool_search happened to disclose the
   pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools
   coupling as a deterministic fact. The wide-catalog reduction
   benchmark is unchanged (82.9%): its synthetic fixture carries no
   outbound tools.

2. The trigger_create description and its prompt-field schema said
   "never call builtin__outbound_deliver in a web-app-created routine".
   The clause is correct for the no-named-destination default, but
   creation turns over-apply it — the qa_8d canary creation verbatim
   reasoned "I'm in the web app, so there's no outbound delivery target
   to pin — let me use the Slack extension's tools for the send step"
   before recovering. Both texts now scope the no-delivery default to
   "no external destination named" and state the named-destination rule
   explicitly: reaching the user or anyone else on an external surface
   goes through builtin__outbound_deliver with a pinned target id,
   never through integration messaging tools (concrete extension names
   kept out per the specificity gate).

Regression tests: the core-name census pins both tools with their
capability ids; the description tests pin the scoped clause, the
absence of the categorical never-clause, and the named-destination
steering on both the tool description and the prompt-field schema.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
personal-upstream-sync Bot pushed a commit to theredspoon/ironclaw that referenced this pull request Aug 8, 2026
…ntract (nearai#7389)

* fix(live-qa): verify triggered Slack delivery through the two-lane contract

Since nearai#7157 a triggered fire's result is never pushed by the completion
driver: the fire itself calls builtin.outbound_deliver, and the
background-run notifier's triggered-run-delivery record describes NOTICE
delivery only — a cleanly completed fire records `skipped`. The delivery
cases still required that record to say `delivered`, which no longer
exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every
scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC)
even though all four live fires verifiably delivered (three had the
marker sitting in Slack history; the fourth was provider-confirmed).

The waiter now verifies what the product actually guarantees:

- success = the fire's durable outbound/deliveries model-delivery record
  for the exact run (delivered, expected DM) PLUS the independent Slack
  history read-back finding the marker;
- notifier records: `skipped`/`no_default_configured`/`delivered` are
  healthy terminals, only `failed`/`denied` fail the case, and unknown
  future vocabulary surfaces through timeout diagnostics;
- a completed outbound_deliver whose composed content lacks the marker
  fails deterministically (the qa_8d mode: the stale prompt bound the
  marker to the final answer, which is no longer the delivered payload);
- the readback-inconclusive flake classification accepts an
  exactly-one-verified-send through either lane.

Case prompts now bind the marker to the delivered Slack message itself
(and still to the final answer), via one shared prompt-requirement
helper.

Also fixes the QA 6D-6E strict-scrub false positive: progressive tool
disclosure (nearai#6958) records tool_search output in traces, and the
builtin.extension_register_hosted_mcp description's "bearer for a static
API token or PAT sent as a Bearer token" prose tripped the bearer
pattern, deleting the trace and failing the shard with all cases green.
The bearer pattern now requires 16+ token-alphabet characters.

All delivery-wait decision logic is pinned by new unit tests against the
production record shapes captured from the failing canary artifacts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(live-qa): close review findings on the two-lane delivery contract

- Gate marker_deliver_count on completed previews: a failed or in-flight
  outbound_deliver whose content carries the marker never reached Slack,
  and counting it could fake the exactly-one-verified-send inconclusive
  classification or suppress the deterministic markerless red. Fixture
  gains a failed marker-bearing preview, observed red before the fix.
- Align emit_results_json.py's bearer pattern with the scrub script's
  16-char floor so description prose in results.json is not mangled to
  "Bearer <REDACTED>"; prose-preservation regression added.
- Namespace the readback-inconclusive evidence per lane
  (vendor_evidence/deliver_evidence) — both dicts carry
  parse_error_count and the flat merge let one overwrite the other.
- Reuse the production root_filesystem schema helpers in the new test
  fixtures instead of a hand-written CREATE TABLE.
- Document why the deterministic content check keys on
  skipped/no_default_configured rather than the whole healthy_terminal
  class: `delivered` includes a fire parked on an approval gate whose
  run resumes — and may deliver — after the notice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(live-qa): pin the exact 16-char bearer floor on both redaction rules

The prose-preservation tests prove prose survives but not the threshold
itself — a {15,} regression would have passed both. Pin the 15/16
boundary explicitly in the emitter suite and the shell scrubber suite,
since the two rule sets are documented as kept in sync.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…ure (nearai#7390)

A web-created routine asking for GitHub-issue summaries "in a Slack
message" was created with a stored prompt instructing the fire to use
the vendor send-message tool instead of a pinned
builtin__outbound_deliver step; an identical retry produced the correct
pinned step. Two compounding causes, both observed live:

1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list
   were Discoverable-tier, so on a catalog past the defer threshold the
   bridged disclosure surface (default since nearai#6958) drops them from
   visible_capabilities — and the delivery guidance block renders only
   while both are visible (delivery_tools_visible). trigger_create is
   Core, so the model could create routines while blind on the delivery
   lane and without the "'Send it to me' is bot delivery via
   builtin__outbound_deliver" steering. Whether the steering existed
   depended on whether an earlier tool_search happened to disclose the
   pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools
   coupling as a deterministic fact. The wide-catalog reduction
   benchmark is unchanged (82.9%): its synthetic fixture carries no
   outbound tools.

2. The trigger_create description and its prompt-field schema said
   "never call builtin__outbound_deliver in a web-app-created routine".
   The clause is correct for the no-named-destination default, but
   creation turns over-apply it — the qa_8d canary creation verbatim
   reasoned "I'm in the web app, so there's no outbound delivery target
   to pin — let me use the Slack extension's tools for the send step"
   before recovering. Both texts now scope the no-delivery default to
   "no external destination named" and state the named-destination rule
   explicitly: reaching the user or anyone else on an external surface
   goes through builtin__outbound_deliver with a pinned target id,
   never through integration messaging tools (concrete extension names
   kept out per the specificity gate).

Regression tests: the core-name census pins both tools with their
capability ids; the description tests pin the scoped clause, the
absence of the categorical never-clause, and the named-destination
steering on both the tool description and the prompt-field schema.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…ntract (nearai#7389)

* fix(live-qa): verify triggered Slack delivery through the two-lane contract

Since nearai#7157 a triggered fire's result is never pushed by the completion
driver: the fire itself calls builtin.outbound_deliver, and the
background-run notifier's triggered-run-delivery record describes NOTICE
delivery only — a cleanly completed fire records `skipped`. The delivery
cases still required that record to say `delivered`, which no longer
exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every
scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC)
even though all four live fires verifiably delivered (three had the
marker sitting in Slack history; the fourth was provider-confirmed).

The waiter now verifies what the product actually guarantees:

- success = the fire's durable outbound/deliveries model-delivery record
  for the exact run (delivered, expected DM) PLUS the independent Slack
  history read-back finding the marker;
- notifier records: `skipped`/`no_default_configured`/`delivered` are
  healthy terminals, only `failed`/`denied` fail the case, and unknown
  future vocabulary surfaces through timeout diagnostics;
- a completed outbound_deliver whose composed content lacks the marker
  fails deterministically (the qa_8d mode: the stale prompt bound the
  marker to the final answer, which is no longer the delivered payload);
- the readback-inconclusive flake classification accepts an
  exactly-one-verified-send through either lane.

Case prompts now bind the marker to the delivered Slack message itself
(and still to the final answer), via one shared prompt-requirement
helper.

Also fixes the QA 6D-6E strict-scrub false positive: progressive tool
disclosure (nearai#6958) records tool_search output in traces, and the
builtin.extension_register_hosted_mcp description's "bearer for a static
API token or PAT sent as a Bearer token" prose tripped the bearer
pattern, deleting the trace and failing the shard with all cases green.
The bearer pattern now requires 16+ token-alphabet characters.

All delivery-wait decision logic is pinned by new unit tests against the
production record shapes captured from the failing canary artifacts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(live-qa): close review findings on the two-lane delivery contract

- Gate marker_deliver_count on completed previews: a failed or in-flight
  outbound_deliver whose content carries the marker never reached Slack,
  and counting it could fake the exactly-one-verified-send inconclusive
  classification or suppress the deterministic markerless red. Fixture
  gains a failed marker-bearing preview, observed red before the fix.
- Align emit_results_json.py's bearer pattern with the scrub script's
  16-char floor so description prose in results.json is not mangled to
  "Bearer <REDACTED>"; prose-preservation regression added.
- Namespace the readback-inconclusive evidence per lane
  (vendor_evidence/deliver_evidence) — both dicts carry
  parse_error_count and the flat merge let one overwrite the other.
- Reuse the production root_filesystem schema helpers in the new test
  fixtures instead of a hand-written CREATE TABLE.
- Document why the deterministic content check keys on
  skipped/no_default_configured rather than the whole healthy_terminal
  class: `delivered` includes a fire parked on an approval gate whose
  run resumes — and may deliver — after the notice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(live-qa): pin the exact 16-char bearer floor on both redaction rules

The prose-preservation tests prove prose survives but not the threshold
itself — a {15,} regression would have passed both. Pin the 15/16
boundary explicitly in the emitter suite and the shell scrubber suite,
since the two rule sets are documented as kept in sync.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…ure (nearai#7390)

A web-created routine asking for GitHub-issue summaries "in a Slack
message" was created with a stored prompt instructing the fire to use
the vendor send-message tool instead of a pinned
builtin__outbound_deliver step; an identical retry produced the correct
pinned step. Two compounding causes, both observed live:

1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list
   were Discoverable-tier, so on a catalog past the defer threshold the
   bridged disclosure surface (default since nearai#6958) drops them from
   visible_capabilities — and the delivery guidance block renders only
   while both are visible (delivery_tools_visible). trigger_create is
   Core, so the model could create routines while blind on the delivery
   lane and without the "'Send it to me' is bot delivery via
   builtin__outbound_deliver" steering. Whether the steering existed
   depended on whether an earlier tool_search happened to disclose the
   pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools
   coupling as a deterministic fact. The wide-catalog reduction
   benchmark is unchanged (82.9%): its synthetic fixture carries no
   outbound tools.

2. The trigger_create description and its prompt-field schema said
   "never call builtin__outbound_deliver in a web-app-created routine".
   The clause is correct for the no-named-destination default, but
   creation turns over-apply it — the qa_8d canary creation verbatim
   reasoned "I'm in the web app, so there's no outbound delivery target
   to pin — let me use the Slack extension's tools for the send step"
   before recovering. Both texts now scope the no-delivery default to
   "no external destination named" and state the named-destination rule
   explicitly: reaching the user or anyone else on an external surface
   goes through builtin__outbound_deliver with a pinned target id,
   never through integration messaging tools (concrete extension names
   kept out per the specificity gate).

Regression tests: the core-name census pins both tools with their
capability ids; the description tests pin the scoped clause, the
absence of the categorical never-clause, and the named-destination
steering on both the tool description and the prompt-field schema.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Kampouse pushed a commit to Kampouse/ironclaw that referenced this pull request Aug 13, 2026
…ntract (nearai#7389)

* fix(live-qa): verify triggered Slack delivery through the two-lane contract

Since nearai#7157 a triggered fire's result is never pushed by the completion
driver: the fire itself calls builtin.outbound_deliver, and the
background-run notifier's triggered-run-delivery record describes NOTICE
delivery only — a cleanly completed fire records `skipped`. The delivery
cases still required that record to say `delivered`, which no longer
exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every
scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC)
even though all four live fires verifiably delivered (three had the
marker sitting in Slack history; the fourth was provider-confirmed).

The waiter now verifies what the product actually guarantees:

- success = the fire's durable outbound/deliveries model-delivery record
  for the exact run (delivered, expected DM) PLUS the independent Slack
  history read-back finding the marker;
- notifier records: `skipped`/`no_default_configured`/`delivered` are
  healthy terminals, only `failed`/`denied` fail the case, and unknown
  future vocabulary surfaces through timeout diagnostics;
- a completed outbound_deliver whose composed content lacks the marker
  fails deterministically (the qa_8d mode: the stale prompt bound the
  marker to the final answer, which is no longer the delivered payload);
- the readback-inconclusive flake classification accepts an
  exactly-one-verified-send through either lane.

Case prompts now bind the marker to the delivered Slack message itself
(and still to the final answer), via one shared prompt-requirement
helper.

Also fixes the QA 6D-6E strict-scrub false positive: progressive tool
disclosure (nearai#6958) records tool_search output in traces, and the
builtin.extension_register_hosted_mcp description's "bearer for a static
API token or PAT sent as a Bearer token" prose tripped the bearer
pattern, deleting the trace and failing the shard with all cases green.
The bearer pattern now requires 16+ token-alphabet characters.

All delivery-wait decision logic is pinned by new unit tests against the
production record shapes captured from the failing canary artifacts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(live-qa): close review findings on the two-lane delivery contract

- Gate marker_deliver_count on completed previews: a failed or in-flight
  outbound_deliver whose content carries the marker never reached Slack,
  and counting it could fake the exactly-one-verified-send inconclusive
  classification or suppress the deterministic markerless red. Fixture
  gains a failed marker-bearing preview, observed red before the fix.
- Align emit_results_json.py's bearer pattern with the scrub script's
  16-char floor so description prose in results.json is not mangled to
  "Bearer <REDACTED>"; prose-preservation regression added.
- Namespace the readback-inconclusive evidence per lane
  (vendor_evidence/deliver_evidence) — both dicts carry
  parse_error_count and the flat merge let one overwrite the other.
- Reuse the production root_filesystem schema helpers in the new test
  fixtures instead of a hand-written CREATE TABLE.
- Document why the deterministic content check keys on
  skipped/no_default_configured rather than the whole healthy_terminal
  class: `delivered` includes a fire parked on an approval gate whose
  run resumes — and may deliver — after the notice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(live-qa): pin the exact 16-char bearer floor on both redaction rules

The prose-preservation tests prove prose survives but not the threshold
itself — a {15,} regression would have passed both. Pin the 15/16
boundary explicitly in the emitter suite and the shell scrubber suite,
since the two rule sets are documented as kept in sync.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…p, stale events floor recapture (nearai#6966)

* fix(ci): use histogram diff in the changed-coverage gate

`reborn_changed_coverage.py` built its denominator from `git diff
--unified=0` with no `--diff-algorithm`, so it inherited git's default
(myers). Myers anchors greedily: on a deletion-shaped diff it shreds one
large removal into interleaved -/+ hunks and re-emits surviving, unchanged
text as added lines. The gate then demands 100% coverage for code the PR
never touched.

Discovered on nearai#6964 (deleting the verified-dead half of `llm::reasoning`),
where the gate saw 478 changed lines / 208 branch arms and failed at 83.89%
/ 65.87%. Measured on that same range with the gate's own parser:

  myers (old):      917 changed production lines
  histogram (new):   14 changed production lines

All 14 are doc comments and imports — zero executable — so the true
changed-testable denominator was zero and the 478 was entirely artifact.

The regression test asserts the invocation rather than re-staging a myers
pathology: the pathology depends on git's internal heuristics, so a fixture
built around one can quietly stop reproducing on a future git and leave a
vacuous green test. Verified red-then-green — removing the flag fails the
new case.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(deps): bump wasmtime 47.0.2 -> 47.0.3 (RUSTSEC-2026-0222, RUSTSEC-2026-0223)

Two Wasmtime advisories published today fail `cargo deny check advisories`
on every branch, which is the "Fast deterministic checks" red cascading
into the required Code Style aggregate:

  RUSTSEC-2026-0222 — stores can mix up type indices between engines
  RUSTSEC-2026-0223 — preemption/traps during bulk operations can break
                      internal VM state

Both name `>=47.0.3` as the fix for the 47.x line.

`cargo update -p wasmtime`. Cargo.lock-only. 28 packages move, every one of
them on Wasmtime's own lockstep release train — wasmtime* 47.0.2 -> 47.0.3,
cranelift* 0.134.2 -> 0.134.3 (its codegen backend), pulley* 47.0.2 ->
47.0.3 (its interpreter). Nothing outside that family changed; no package
added or removed.

Verified locally with cargo-deny 0.19.9: both advisories reproduce on the
old lock and `advisories ok` after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(coverage): recapture the stale ironclaw_events floor (inherited from nearai#6943)

The ironclaw_events floor was captured before nearai#6943 deleted
`events::{parse_jsonl, replay_jsonl}`, so main has been sitting under its
own covered-lines floor ever since: observed 1197 covered / 1486 total
against a floor of 1252 covered (effective 1232). Every branch that reaches
the coverage job fails on it — PR nearai#6958, which touches no events file,
fails with the byte-identical block.

Recaptured to the observed numbers per coverage-floor.toml's own
same-PR recapture workflow for legitimate deletion-driven shrinkage.

The entry is copied byte-for-byte from nearai#6964's commit 7ca468d, which
carries the same fix. Identical text on both branches means git merges them
cleanly in either order. nearai#6964's ironclaw_llm entry is deliberately NOT
brought along — that shrinkage is caused by that PR's own deletion and
belongs to it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…(WS8 closeout) (nearai#6964)

* refactor(llm): delete the verified-dead half of the reasoning module (WS8 closeout)

`ironclaw_llm::reasoning` was half-live. `ironclaw_runner`'s model gateway
calls `clean_response`, `contains_codex_text_tool_call_syntax`, and
`recover_codex_text_tool_calls_from_tool_names` on the live model-response
path; everything else in the module was a v1 engine remnant with no
production consumer. PR nearai#6943 excluded the module for exactly this reason
and left an enumeration to re-verify.

Re-verified all 20 enumerated names against this base: the module is private
(`mod reasoning;`), so its entire external surface is the two `pub use`
blocks in lib.rs, and no crate in the workspace imports any of the 20. All
20 deleted. Three private helpers — `truncate_at_tool_tags`,
`closing_tag_for`, `TOOL_TAG_PATTERNS` — were not enumerated but are
transitively dead: all 10 of their non-test call sites were inside the
deleted `impl Reasoning` block.

The three live helpers are untouched, byte-for-byte, and stay where they
are. Placement is deferred to the Wave 3 runner shed.

Un-masking: ironclaw_llm 1000 -> 884 (-116, all reasoning-module tests
belonging to deleted code, each classified); ironclaw_runner 467 -> 467 with
an identical roster. No surviving test was edited — the test-module diff has
zero added lines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(checklist): tick the WS8 llm::reasoning row as landed via nearai#6964

Records the outcome on that row only: 20/20 enumerated names re-verified dead
and deleted with no exclusions, three transitively-dead helpers found beyond
the enumeration, the un-masking counts, and the explicit "deferred to the
Wave 3 runner shed" answer to the row's placement question.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(coverage): recapture ironclaw_llm and ironclaw_events ratchet floors

Both floors are the documented legitimate-shrinkage case from
coverage-floor.toml's own header (a code+test deletion lowering covered
lines updates the entry in the same PR). Numbers copied verbatim from this
PR's coverage-report run; no tests were added to chase the old floors.

ironclaw_llm — caused by this PR. Deleting the verified-dead half of
`llm::reasoning` removed 1,871 instrumented lines (28,235 -> 26,364, a
material -6.63% denominator move). That dead half carried denser test
coverage than the crate average — 116 of the crate's tests exercised it —
so removing code and tests together lowered the percentage even though no
live path lost coverage. Recaptured to observed: 79.22% / 20,885 covered.

ironclaw_events — inherited from main, not caused by this PR. The previous
floor was captured before nearai#6943 deleted `events::{parse_jsonl,
replay_jsonl}`, so main has been sitting under its own floor
(-59 instrumented lines and their covered code); this branch is simply the
first the ratchet caught. PR nearai#6958, which touches no events file, fails
identically. Recaptured to observed: 80.55% / 1,197 covered.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* Instrument canary model and tool usage

* ci: unblock the queue — histogram diff gate fix, wasmtime RUSTSEC bump, stale events floor recapture (nearai#6966)

* fix(ci): use histogram diff in the changed-coverage gate

`reborn_changed_coverage.py` built its denominator from `git diff
--unified=0` with no `--diff-algorithm`, so it inherited git's default
(myers). Myers anchors greedily: on a deletion-shaped diff it shreds one
large removal into interleaved -/+ hunks and re-emits surviving, unchanged
text as added lines. The gate then demands 100% coverage for code the PR
never touched.

Discovered on nearai#6964 (deleting the verified-dead half of `llm::reasoning`),
where the gate saw 478 changed lines / 208 branch arms and failed at 83.89%
/ 65.87%. Measured on that same range with the gate's own parser:

  myers (old):      917 changed production lines
  histogram (new):   14 changed production lines

All 14 are doc comments and imports — zero executable — so the true
changed-testable denominator was zero and the 478 was entirely artifact.

The regression test asserts the invocation rather than re-staging a myers
pathology: the pathology depends on git's internal heuristics, so a fixture
built around one can quietly stop reproducing on a future git and leave a
vacuous green test. Verified red-then-green — removing the flag fails the
new case.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(deps): bump wasmtime 47.0.2 -> 47.0.3 (RUSTSEC-2026-0222, RUSTSEC-2026-0223)

Two Wasmtime advisories published today fail `cargo deny check advisories`
on every branch, which is the "Fast deterministic checks" red cascading
into the required Code Style aggregate:

  RUSTSEC-2026-0222 — stores can mix up type indices between engines
  RUSTSEC-2026-0223 — preemption/traps during bulk operations can break
                      internal VM state

Both name `>=47.0.3` as the fix for the 47.x line.

`cargo update -p wasmtime`. Cargo.lock-only. 28 packages move, every one of
them on Wasmtime's own lockstep release train — wasmtime* 47.0.2 -> 47.0.3,
cranelift* 0.134.2 -> 0.134.3 (its codegen backend), pulley* 47.0.2 ->
47.0.3 (its interpreter). Nothing outside that family changed; no package
added or removed.

Verified locally with cargo-deny 0.19.9: both advisories reproduce on the
old lock and `advisories ok` after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(coverage): recapture the stale ironclaw_events floor (inherited from nearai#6943)

The ironclaw_events floor was captured before nearai#6943 deleted
`events::{parse_jsonl, replay_jsonl}`, so main has been sitting under its
own covered-lines floor ever since: observed 1197 covered / 1486 total
against a floor of 1252 covered (effective 1232). Every branch that reaches
the coverage job fails on it — PR nearai#6958, which touches no events file,
fails with the byte-identical block.

Recaptured to the observed numbers per coverage-floor.toml's own
same-PR recapture workflow for legitimate deletion-driven shrinkage.

The entry is copied byte-for-byte from nearai#6964's commit 7ca468d, which
carries the same fix. Identical text on both branches means git merges them
cleanly in either order. nearai#6964's ironclaw_llm entry is deliberately NOT
brought along — that shrinkage is caused by that PR's own deletion and
belongs to it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* Render aggregate metrics in canary PR reports

* test(reborn): allow instrumented runtime paths to settle

* Map live QA harness in Reborn test planner

* Preserve selected Reborn coverage mode

* Ignore workspace MSRV in selected Reborn lanes

---------

Co-authored-by: Benjamin Kurrek <57506486+BenKurrek@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
)

* feat(reborn): enable progressive tool disclosure by default

* test(reborn): pin scripted QA tool disclosure

* test(reborn): pin composition tool surfaces

* test(reborn): pin hook runtime tool surface

* test(reborn): make flat tool fixtures explicit

* test(e2e): pin flat disclosure fixtures

* test(e2e): pin responses fixtures to flat tools

* chore(reborn): clarify tool disclosure wording
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…ure (nearai#7390)

A web-created routine asking for GitHub-issue summaries "in a Slack
message" was created with a stored prompt instructing the fire to use
the vendor send-message tool instead of a pinned
builtin__outbound_deliver step; an identical retry produced the correct
pinned step. Two compounding causes, both observed live:

1. builtin.outbound_deliver and builtin.outbound_delivery_targets_list
   were Discoverable-tier, so on a catalog past the defer threshold the
   bridged disclosure surface (default since nearai#6958) drops them from
   visible_capabilities — and the delivery guidance block renders only
   while both are visible (delivery_tools_visible). trigger_create is
   Core, so the model could create routines while blind on the delivery
   lane and without the "'Send it to me' is bot delivery via
   builtin__outbound_deliver" steering. Whether the steering existed
   depended on whether an earlier tool_search happened to disclose the
   pair. Both tools are now Core, restoring nearai#7157's guidance-iff-tools
   coupling as a deterministic fact. The wide-catalog reduction
   benchmark is unchanged (82.9%): its synthetic fixture carries no
   outbound tools.

2. The trigger_create description and its prompt-field schema said
   "never call builtin__outbound_deliver in a web-app-created routine".
   The clause is correct for the no-named-destination default, but
   creation turns over-apply it — the qa_8d canary creation verbatim
   reasoned "I'm in the web app, so there's no outbound delivery target
   to pin — let me use the Slack extension's tools for the send step"
   before recovering. Both texts now scope the no-delivery default to
   "no external destination named" and state the named-destination rule
   explicitly: reaching the user or anyone else on an external surface
   goes through builtin__outbound_deliver with a pinned target id,
   never through integration messaging tools (concrete extension names
   kept out per the specificity gate).

Regression tests: the core-name census pins both tools with their
capability ids; the description tests pin the scoped clause, the
absence of the categorical never-clause, and the named-destination
steering on both the tool description and the prompt-field schema.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…ntract (nearai#7389)

* fix(live-qa): verify triggered Slack delivery through the two-lane contract

Since nearai#7157 a triggered fire's result is never pushed by the completion
driver: the fire itself calls builtin.outbound_deliver, and the
background-run notifier's triggered-run-delivery record describes NOTICE
delivery only — a cleanly completed fire records `skipped`. The delivery
cases still required that record to say `delivered`, which no longer
exists for results, so qa_3d/qa_8d/qa_9b/qa_9d hard-failed every
scheduled run from the first post-nearai#7157 canary (2026-08-08 00:24 UTC)
even though all four live fires verifiably delivered (three had the
marker sitting in Slack history; the fourth was provider-confirmed).

The waiter now verifies what the product actually guarantees:

- success = the fire's durable outbound/deliveries model-delivery record
  for the exact run (delivered, expected DM) PLUS the independent Slack
  history read-back finding the marker;
- notifier records: `skipped`/`no_default_configured`/`delivered` are
  healthy terminals, only `failed`/`denied` fail the case, and unknown
  future vocabulary surfaces through timeout diagnostics;
- a completed outbound_deliver whose composed content lacks the marker
  fails deterministically (the qa_8d mode: the stale prompt bound the
  marker to the final answer, which is no longer the delivered payload);
- the readback-inconclusive flake classification accepts an
  exactly-one-verified-send through either lane.

Case prompts now bind the marker to the delivered Slack message itself
(and still to the final answer), via one shared prompt-requirement
helper.

Also fixes the QA 6D-6E strict-scrub false positive: progressive tool
disclosure (nearai#6958) records tool_search output in traces, and the
builtin.extension_register_hosted_mcp description's "bearer for a static
API token or PAT sent as a Bearer token" prose tripped the bearer
pattern, deleting the trace and failing the shard with all cases green.
The bearer pattern now requires 16+ token-alphabet characters.

All delivery-wait decision logic is pinned by new unit tests against the
production record shapes captured from the failing canary artifacts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(live-qa): close review findings on the two-lane delivery contract

- Gate marker_deliver_count on completed previews: a failed or in-flight
  outbound_deliver whose content carries the marker never reached Slack,
  and counting it could fake the exactly-one-verified-send inconclusive
  classification or suppress the deterministic markerless red. Fixture
  gains a failed marker-bearing preview, observed red before the fix.
- Align emit_results_json.py's bearer pattern with the scrub script's
  16-char floor so description prose in results.json is not mangled to
  "Bearer <REDACTED>"; prose-preservation regression added.
- Namespace the readback-inconclusive evidence per lane
  (vendor_evidence/deliver_evidence) — both dicts carry
  parse_error_count and the flat merge let one overwrite the other.
- Reuse the production root_filesystem schema helpers in the new test
  fixtures instead of a hand-written CREATE TABLE.
- Document why the deterministic content check keys on
  skipped/no_default_configured rather than the whole healthy_terminal
  class: `delivered` includes a fire parked on an approval gate whose
  run resumes — and may deliver — after the notice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(live-qa): pin the exact 16-char bearer floor on both redaction rules

The prose-preservation tests prove prose survives but not the threshold
itself — a {15,} regression would have passed both. Pin the 15/16
boundary explicitly in the emitter suite and the shell scrubber suite,
since the two rule sets are documented as kept in sync.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6958 — 34431dc9 Deployed Aug 7, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: S 10-49 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Make progressive tool disclosure default-on without degrading everyday tool use

2 participants