Skip to content

reborn: enable final-answer nudge for planned_default and scheduled_trigger - #5568

Closed
henrypark133 wants to merge 16 commits into
mainfrom
worktree-reborn-nudge-profile-enable
Closed

henrypark133 wants to merge 16 commits into
mainfrom
worktree-reborn-nudge-profile-enable

Conversation

@henrypark133

Copy link
Copy Markdown
Collaborator

Summary

  • Turns on Reborn's dormant final-answer nudge (SteeringPolicy.allow_driver_specific_nudges) for planned_default (real interactive/chat/CLI traffic) and scheduled_trigger (trigger-fired runs); subagent stays off, guarded by a regression test.
  • New RunProfileDefinition::with_driver_specific_nudges(bool) builder mirrors the existing with_personal_context_policy pattern; the shared planned_like_profile_definition helper is untouched, so each profile opts in individually.
  • Proven at 3 tiers: unit assertions on profile resolution, driver-tier tests (crates/ironclaw_reborn/tests/planned_driver_e2e.rs) driving the real PlannedDriver, and a product-level test (tests/reborn_integration_nudge_final_answer.rs) through the real submit_turn → product workflow → agent loop path.

Design notes

  • Original scoping targeted the literal RunProfileId::interactive_default() construct; corrected mid-design after finding real interactive traffic actually resolves to planned_default (requested_run_profile: None → production resolver default). See docs/superpowers/specs/2026-07-02-reborn-nudge-profile-enable-design.md.
  • Went through task-level review per commit, a final whole-branch review, and a thermo-nuclear maintainability pass (found + fixed one test-code duplication, extracted assert_nudge_fires_for_resolved_profile).

Test plan

  • cargo fmt clean
  • cargo clippy --all --benches --tests --examples --all-features clean (zero new warnings)
  • cargo test -p ironclaw_turns -p ironclaw_reborn -p ironclaw_agent_loop — all green
  • cargo test --test reborn_integration_nudge_final_answer — green
  • Full workspace cargo test — green (two environmental SIGABRT flakes under parallel load confirmed unrelated via isolated re-run)

🤖 Generated with Claude Code

henrypark133 and others added 15 commits July 2, 2026 09:40
Scopes allow_driver_specific_nudges to interactive_default and
scheduled_trigger via a builder method, avoiding a shared-base flip
that would leak into planned_default/subagent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…anned_default

Real production interactive/chat/CLI turns request no explicit run
profile (submit_user_turn passes requested_run_profile: None) and the
production resolver defaults that to planned_default, not the literal
interactive_profile() construct. Retargets the design accordingly and
simplifies the implementation (no shared-base change needed at all).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…trigger nudges

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…lear review

Tasks 6/7 previously copy-pasted the same scripted scenario and
completion assertion; extract no_progress_script()/
assert_completed_via_nudge() once, following the file's existing
run_request/run_context_for_driver helper-extraction pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
User asked for coverage similar to tests/reborn_group_extensions. The
group harness is for cross-thread scenarios; this is single-thread, so
the correct analog per tests/support/reborn/CLAUDE.md is a flat
reborn_integration_*.rs test using RebornIntegrationHarness. Feasibility
checked: submit_turn defaults to planned_default with no override
needed, and CapabilityProgress::NoChange is computed generically from
real capability output, not test-mock-only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…comments

Fix two wording-only doc-comment inaccuracies from review: the
scheduled-trigger profile doc now names both axes that diverge from
the shared base (capability surface AND driver-specific-nudges), and
the nudge integration test's module doc attributes builtin.echo's
exclusion to the test harness's own fixed capability grant list
(core_builtin_tools_from_runtime) rather than a nonexistent
production run-profile-level exclusion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The two *_profile_completes_via_final_answer_nudge tests duplicated
~20 lines of driver/host/request construction, differing only in the
resolved profile and a context label. Extract
assert_nudge_fires_for_resolved_profile to hold the shared run+assert
sequence — found during a pre-PR thermo-nuclear pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings July 2, 2026 19:40
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5568 July 2, 2026 19:40 Destroyed
@github-actions github-actions Bot added scope: docs Documentation size: S 10-49 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 2, 2026
@coderabbitai

coderabbitai Bot commented Jul 2, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added support for driver-specific nudges in planned run profiles, including a new configuration option to enable or disable them.
    • Improved no-progress handling so repeated identical tool results can now trigger a final-answer nudge and complete the run.
  • Bug Fixes

    • Updated planned-driver behavior to complete successfully in repeated no-progress scenarios.
    • Expanded end-to-end coverage for nudge-based completion in planned and scheduled profiles.

Walkthrough

Adds a with_driver_specific_nudges builder on RunProfileDefinition, enables allow_driver_specific_nudges for planned_default and scheduled_trigger profiles, updates a doc comment, and adds unit, e2e, and product-level integration tests validating final-answer nudge completion on no-progress detection.

Changes

Nudge enablement and tests

Layer / File(s) Summary
RunProfileDefinition builder
crates/ironclaw_turns/src/run_profile/resolver.rs
Adds with_driver_specific_nudges(enabled) builder toggling steering_policy.allow_driver_specific_nudges, with unit test asserting default-false and explicit-true resolution.
Profile wiring
crates/ironclaw_reborn/src/planned_driver_factory.rs, crates/ironclaw_agent_loop/src/executor/loop_exit.rs
planned_default_profile_definition and schedule_trigger_planned_profile_definition now chain with_driver_specific_nudges(true); existing resolver tests updated to assert enabled/disabled flag per profile; stale try_final_answer_nudge doc comment corrected to match current gating.
Driver-level e2e tests
crates/ironclaw_reborn/tests/planned_driver_e2e.rs
Adds scripted no-progress scenario helpers and two Tokio tests confirming PlannedDriver completes via the final-answer nudge for planned_default and scheduled_trigger profiles.
Product-level integration test
tests/reborn_integration_nudge_final_answer.rs
New integration test scripts four identical builtin.http calls plus a final reply through submit_turn, asserting completion via the final-answer nudge.

Estimated code review effort: 2 (Simple) | ~15 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Test as E2E/Integration Test
  participant Resolver as RunProfileResolver
  participant Driver as PlannedDriver
  participant Host as ScriptedHost

  Test->>Resolver: resolve planned_default / scheduled_trigger profile
  Resolver-->>Test: steering_policy.allow_driver_specific_nudges = true
  Test->>Driver: run(no-progress script)
  Driver->>Host: repeated identical capability calls
  Host-->>Driver: identical outcomes (no progress)
  Driver->>Driver: detect NoProgressDetected
  Driver->>Host: final-answer nudge call
  Host-->>Driver: scripted final reply
  Driver-->>Test: LoopExit::Completed via nudge
Loading

Possibly related PRs

  • nearai/ironclaw#4588: Introduces the final-answer nudge logic and allow_driver_specific_nudges flag that this PR enables in run profiles.
  • nearai/ironclaw#4837: Adds the try_final_answer_nudge/StopKind::NoProgressDetected gating that this PR's profile wiring and doc-comment update depend on.
  • nearai/ironclaw#4993: Modifies the same NoProgressDetected completion path and steering-policy gating exercised by this PR's new tests.
🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is detailed, but it omits required template sections like Change Type, Linked Issue, Security Impact, and several checklist fields. Add the missing template sections, especially Change Type, Linked Issue, Security Impact, trust-boundary checklist, rollback, and review-follow-through.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly matches the PR’s main change to enable final-answer nudges for planned_default and scheduled_trigger.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enables Reborn's final-answer nudge (allow_driver_specific_nudges) for the planned_default and scheduled_trigger run profiles, while keeping it disabled for the subagent profile. To support this, a new builder method with_driver_specific_nudges was added to RunProfileDefinition. The changes are accompanied by comprehensive unit tests, driver-tier end-to-end tests, and a product-level integration test proving that the nudge successfully fires and completes the run when no progress is detected. No review comments were provided, so there is no feedback to address.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Plan/design-spec docs were working artifacts for this session, not
meant to ship in the PR.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5568 July 2, 2026 19:41 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Enables Reborn’s final-answer nudge mechanism by allowing driver-specific nudges for the real interactive planned_default profile and scheduled_trigger, while keeping subagent opted out. This extends run-profile configuration in ironclaw_turns, wires the flag at the intended Reborn profile-definition call sites, and adds end-to-end tests proving the nudge fires both at the driver tier and through the real submit_turn path.

Changes:

  • Add RunProfileDefinition::with_driver_specific_nudges(bool) and unit-test it in ironclaw_turns.
  • Enable the nudge flag for planned_default and scheduled_trigger in ironclaw_reborn (with regression assertions that subagent remains disabled).
  • Add driver-tier and product-tier integration tests that exercise no-progress → final-answer-nudge completion.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
tests/reborn_integration_nudge_final_answer.rs New product-level integration test proving the nudge fires through the real submit_turn pipeline using scripted builtin.http calls.
docs/superpowers/specs/2026-07-02-reborn-nudge-profile-enable-design.md Design/spec write-up explaining scope correction and implementation approach.
docs/superpowers/plans/2026-07-02-reborn-nudge-profile-enable.md Implementation plan and task breakdown for the change (note: contains a few now-stale builtin.echo references).
crates/ironclaw_turns/src/run_profile/resolver.rs Adds the with_driver_specific_nudges builder + unit test verifying it propagates through resolution.
crates/ironclaw_reborn/tests/planned_driver_e2e.rs Adds driver-tier tests that resolve real profiles and prove the nudge causes completion (extra tool-free model call).
crates/ironclaw_reborn/src/planned_driver_factory.rs Opts planned_default and scheduled_trigger into driver-specific nudges; asserts subagent remains off.
crates/ironclaw_agent_loop/src/executor/loop_exit.rs Updates try_final_answer_nudge doc comment to reflect production enablement for select profiles.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +32 to +36
/// Gated by `SteeringPolicy.allow_driver_specific_nudges` (enabled for select
/// Reborn run profiles — see `ironclaw_reborn::planned_driver_factory`; off by
/// default elsewhere) and capped at one nudge per run. Returns `Ok(None)` when
/// disabled, capped, or the model still declines to answer — callers then keep
/// their existing behavior.
Copilot AI review requested due to automatic review settings July 2, 2026 19:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

github-actions Bot commented Jul 2, 2026

Copy link
Copy Markdown
Contributor

⚠️ 4 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_reborn_config, ironclaw_reborn_identity, ironclaw_reborn_traces, ironclaw_webui_v2

Reborn integration-tier coverage

Line coverage (Reborn crates): 15.06% — 9590 / 63697 lines

Per-crate breakdown (11 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_reborn_config 0% 0 / 1142
ironclaw_reborn_identity 0% 0 / 237
ironclaw_reborn_traces 0% 0 / 6662
ironclaw_webui_v2 0% 0 / 2785
ironclaw_reborn_event_store 0.73% 6 / 825
ironclaw_product_adapter_registry 5.62% 25 / 445
ironclaw_product_workflow 5.72% 562 / 9820
ironclaw_product_adapters 12.71% 283 / 2227
ironclaw_reborn_composition 21.8% 6726 / 30858
ironclaw_reborn 22.77% 1977 / 8682
ironclaw_product_context 78.57% 11 / 14

This signal is informational: coverage never gates the PR — not the percentage, not the per-crate holes, not the 0-coverage callout.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/reborn_integration_nudge_final_answer.rs`:
- Around line 56-87: The test
no_progress_repeated_http_call_completes_via_final_answer_nudge only checks the
synthesized reply text and does not verify that exactly four builtin.http tool
calls happened before finalization. Update this test to assert the recorded
tool-call count on the RebornIntegrationHarness (or its scripted-reply history)
after submit_turn, so it proves the no-progress path triggered on the expected
4th batch rather than incidentally producing the same final text.
- Around line 1-44: Gate the reborn integration test so it does not run under
plain cargo test by default. Update
tests/reborn_integration_nudge_final_answer.rs to use the integration gate
already used elsewhere, either with a cfg attribute or by registering it under
the integration test target in the test manifest. Make sure the test entry point
and its RebornIntegrationHarness-based setup are only enabled when the
integration feature is active, matching the existing integration-only pattern
used by the other product-stack tests.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: dfcdd839-92ef-48fc-a432-758388477102

📥 Commits

Reviewing files that changed from the base of the PR and between 0d2ddcf and 66f0540.

📒 Files selected for processing (5)
  • crates/ironclaw_agent_loop/src/executor/loop_exit.rs
  • crates/ironclaw_reborn/src/planned_driver_factory.rs
  • crates/ironclaw_reborn/tests/planned_driver_e2e.rs
  • crates/ironclaw_turns/src/run_profile/resolver.rs
  • tests/reborn_integration_nudge_final_answer.rs

Comment on lines +1 to +44
//! Product-level proof: the final-answer nudge fires through the real
//! `submit_turn` entry point (product workflow → turn coordinator →
//! scheduler → agent loop → real `LlmProviderModelGateway` decorator chain
//! → scripted model), one layer up from the executor/driver-tier proof in
//! `crates/ironclaw_reborn/tests/planned_driver_e2e.rs`.
//!
//! `RebornIntegrationHarness::test_default()` resolves `requested_run_profile:
//! None` to `planned_default` — the profile Task 2 enabled driver-specific
//! nudges for — with no special wiring. Four identical `builtin.http` calls
//! (same URL) drive the real no-progress detector: `RecordingRuntimeHttpEgress`
//! (installed by `.with_builtin_http_tools()`) always returns the same fixed
//! scripted body, so the first call's output digest is first-seen
//! (`MadeProgress`) and the next three repeat the same digest (`NoChange`) —
//! `trailing_no_progress_results` reaches the default
//! `typed_progress_run_threshold` (3) right after the 4th capability batch —
//! `DefaultStopConditionStrategy::should_stop_after_observed_turn` in
//! `crates/ironclaw_agent_loop/src/strategies/stop.rs`. The executor then
//! resolves that `NoProgressDetected` stop via `try_final_answer_nudge`
//! (`crates/ironclaw_agent_loop/src/executor/loop_exit.rs`), issuing one
//! extra tool-free model call that the 5th scripted reply satisfies.
//!
//! Deviation from the plan's starting shape: the brief scripted
//! `builtin.echo`, reasoning that a first-party capability with a stable
//! digest would drive the detector. `builtin.echo` IS registered as a
//! first-party handler with `CapabilityVisibility::Model`
//! (`crates/ironclaw_host_runtime/src/first_party_tools/mod.rs`), but the
//! resolved capability surface here (as observed via `RUST_LOG=debug` —
//! `visible_capability_sample`) does not include it — this test harness's
//! `core_builtin_tools_from_runtime` grants a fixed `capability_ids` list
//! (`tests/support/reborn/harness.rs`) that omits `builtin.echo`, not a
//! production run-profile-level exclusion; the model gateway rejects the
//! scripted call as "outside the visible capability surface"
//! (`ironclaw_reborn::model_gateway`), which surfaces as a terminal
//! `model_error`, not a no-progress signal. Swapping to `builtin.http`
//! (already proven visible + deterministic by
//! `tests/reborn_integration_tool_call.rs`) exercises the same digest-based
//! `NoChange` mechanism without depending on a capability outside this
//! harness's granted surface.

#[allow(dead_code)]
#[path = "support/reborn/mod.rs"]
mod reborn_support;
#[allow(dead_code)]
mod support;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Check whether this integration test target is gated via Cargo.toml required-features
rg -n 'reborn_integration' Cargo.toml
rg -n '\[\[test\]\]' -A5 Cargo.toml

Repository: nearai/ironclaw

Length of output: 997


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '\n== Cargo test target entries ==\n'
sed -n '340,390p' Cargo.toml

printf '\n== Presence of the reviewed file and neighboring test files ==\n'
git ls-files 'tests/**' | rg 'reborn_integration_nudge_final_answer|reborn_integration|reborn_group|e2e_thread_scheduling'

printf '\n== File header ==\n'
cat -n tests/reborn_integration_nudge_final_answer.rs | sed -n '1,80p'

Repository: nearai/ironclaw

Length of output: 8150


Gate tests/reborn_integration_nudge_final_answer.rs behind integration.
This is auto-discovered by plain cargo test, and there is no #[cfg(feature = "integration")] or [[test]] required-features = ["integration"] entry for it, so the full product-stack harness still runs on the default test path.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/reborn_integration_nudge_final_answer.rs` around lines 1 - 44, Gate the
reborn integration test so it does not run under plain cargo test by default.
Update tests/reborn_integration_nudge_final_answer.rs to use the integration
gate already used elsewhere, either with a cfg attribute or by registering it
under the integration test target in the test manifest. Make sure the test entry
point and its RebornIntegrationHarness-based setup are only enabled when the
integration feature is active, matching the existing integration-only pattern
used by the other product-stack tests.

Source: Path instructions

Comment on lines +56 to +87
async fn no_progress_repeated_http_call_completes_via_final_answer_nudge() {
let h = RebornIntegrationHarness::test_default()
.with_builtin_http_tools()
.script([
RebornScriptedReply::tool_call(
"builtin.http",
serde_json::json!({"url": REPEATED_URL}),
),
RebornScriptedReply::tool_call(
"builtin.http",
serde_json::json!({"url": REPEATED_URL}),
),
RebornScriptedReply::tool_call(
"builtin.http",
serde_json::json!({"url": REPEATED_URL}),
),
RebornScriptedReply::tool_call(
"builtin.http",
serde_json::json!({"url": REPEATED_URL}),
),
RebornScriptedReply::text("final answer synthesized via nudge"),
])
.build()
.await
.expect("harness builds");
h.submit_turn("fetch the same item four times")
.await
.expect("turn completes");
h.assert_reply_contains("final answer synthesized via nudge")
.await
.expect("reply finalized via the final-answer nudge");
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Assertion only checks final text, not that exactly 4 tool calls preceded it.

The test proves the nudge-completion text appears, but doesn't assert the harness actually made 4 builtin.http calls before the synthesized reply — i.e. that no-progress detection fired at the expected point rather than some other path incidentally producing the same final text. Given the doc comment's detailed claim about the exact mechanism (trailing_no_progress_results reaching threshold on the 4th batch), asserting the recorded call count would make this a tighter proof rather than relying solely on the doc comment's narrative.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/reborn_integration_nudge_final_answer.rs` around lines 56 - 87, The
test no_progress_repeated_http_call_completes_via_final_answer_nudge only checks
the synthesized reply text and does not verify that exactly four builtin.http
tool calls happened before finalization. Update this test to assert the recorded
tool-call count on the RebornIntegrationHarness (or its scripted-reply history)
after submit_turn, so it proves the no-progress path triggered on the expected
4th batch rather than incidentally producing the same final text.

@henrypark133

Copy link
Copy Markdown
Collaborator Author

/benchmark pinchbench --framework ironclaw-reborn

@github-actions

github-actions Bot commented Jul 2, 2026

Copy link
Copy Markdown
Contributor

🧪 Started pinchbench on ironclaw-reborn against ironclaw 66f0540910 — watch run.

@railway-app

railway-app Bot commented Jul 2, 2026

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5568 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 2, 2026 at 7:49 pm

@henrypark133

Copy link
Copy Markdown
Collaborator Author

Closing as stale — no activity in over three weeks. The branch is untouched; reopen if this is still needed.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5568 — 66f05409 Deployed Jul 2, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: S 10-49 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants