Skip to content

feat(reborn): self-verification pass + benchmark_default profile to actually enable it - #6093

Open
pranavraja99 wants to merge 5 commits into
mainfrom
feat/reborn-selfverify-and-loopguard
Open

pranavraja99 wants to merge 5 commits into
mainfrom
feat/reborn-selfverify-and-loopguard

Conversation

@pranavraja99

@pranavraja99 pranavraja99 commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

What

Adds a gated self-verification pass to reborn's agent loop, plus a way to
actually turn it (and the existing final-answer nudge) on for benchmark
runs without changing default behavior for any other reborn product.

1. The self-verification pass

When the loop is about to gracefully complete a turn AND this run made at
least one capability call earlier (state.recent_call_signatures
non-empty — i.e. the answer followed real tool-based work, not idle chat),
issue ONE extra tool-free model call asking the model to independently
re-derive its just-given answer before it's finalized. Whatever that
verification turn produces (confirmed or corrected) supersedes the original
reply.

Mirrors try_final_answer_nudge exactly: same gate
(SteeringPolicy.allow_driver_specific_nudges), same one-shot-per-run cap
(self_verification_used), same tool-free forced provider call, same
fail-open bail-out on any host-call error (confirmed live — see Evidence:
one task's debug_logs shows the pass firing, the model attempting a tool
call anyway, the host correctly rejecting it, and the mechanism falling
back to the original answer rather than erroring the run).

Deliberately scoped to a single extra text-only reasoning pass, not a live
tool-enabled recompute — re-enabling tools mid-verification would mean
processing capability calls back into the executor at what is currently an
exit boundary, a materially bigger structural change than this PR takes on.

Also threads a nudge_name through the shared nudge_bail fail-open
helper (previously hardcoded to "final-answer nudge" regardless of which
nudge fired) — found live-testing this pass, when a bail from the new
nudge logged as if it were the old one.

Update per review: the eligibility check (prior tool activity, one-shot
cap) originally lived as an inline condition inside ExitStage::for_stop's
GracefulStop arm — bypassing the strategy/stage extension points the
canonical loop is built around (see strategies/CLAUDE.md: add a new
strategy for a stable, independent decision axis, rather than branching
executor code). Moved that decision into
DefaultStopConditionStrategy::should_stop_after_observed_turn: the
reply-completion case now returns a new
StopKind::GracefulStopPendingVerification instead of plain
GracefulStop when state says this reply followed real tool activity and
the pass is unused — mirroring exactly how NoProgressDetected is already
a strategy-decided stop kind the executor separately resolves. ExitStage
no longer inspects state.recent_call_signatures itself; it just trusts
the kind it's given, same as every other arm. The host-policy gate check
stays in the executor (try_self_verification_pass), since strategies
don't have host access by design. 2 new strategy-level tests cover the
eligibility decision directly.

2. Making it actually run: benchmark_default run profile

Nudges are gated off in every existing profile, so merging part 1 alone
changes nothing by default. The direct fix — flipping
allow_driver_specific_nudges on interactive_default — was tried first
and reverted: it broke 4 pre-existing tests elsewhere that hardcode exact
model-call counts, and it would change behavior for every
interactive_default consumer, not just benchmarks. Too broad to force
through blind.

Instead, adds a new, additive benchmark_default run profile — identical
to the default planned profile except nudges are on — registered alongside
(not replacing) the existing default/subagent/scheduled_trigger profiles,
mirroring the exact interactive_like + registry pattern already used for
scheduled_trigger (#5505). A new
RunProfileDefinition::with_driver_specific_nudges builder (mirrors the
existing with_personal_context_policy) sets just that one policy bit.

Reborn's real turn-submission path (send_user_message_with_cancellation)
now requests this profile instead of the implicit default when
IRONCLAW_REBORN_BENCHMARK_PROFILE is set truthy — otherwise behavior is
byte-for-byte unchanged (env var unset by default).

Why

Root-caused via analysis of 24 OfficeQA tasks where hermes consistently
beats reborn across two daily runs, but reborn itself flips pass/fail
between the two days on the same task — i.e. failures the model can
sometimes avoid, not a hard capability gap. These are long, multi-step,
tool-assisted derivations (pick a table → pick a column → extract N values
→ apply a formula → report); sampling variance at any one step compounds
across the chain. This is general — any product on reborn answering a
multi-step tool-assisted question is exposed to the same compounding.

Evidence

Unit tests (cargo test -p ironclaw_agent_loop — 403 passed) cover the
mechanism at both layers: the strategy tests confirm the eligibility
decision (verification-eligible only with prior tool activity and an
unused pass); the executor tests confirm the stage's handling once given
that decision — fires once and supersedes the original reply when gated
on; no-ops when the gate is off; a plain GracefulStop never attempts
verification regardless of state; respects the one-shot cap; falls open to
the original reply when its own model call fails.

Full-tree regression check: cargo test across ironclaw_turns,
ironclaw_runner, ironclaw_agent_loop, ironclaw_reborn_composition —
1268+ passed, 0 failed (one SQLite-backend flake in an unrelated
trigger-poller test, confirmed by rerunning in isolation: passes cleanly).

Live task-level, head-to-head vs. hermes (DeepSeek-V4-Flash,
--framework ironclaw-reborn, IRONCLAW_REBORN_BENCHMARK_PROFILE=1): ran
all 24 of the flip-flop tasks in one batch, where hermes passes 24/24 and
reborn's own historical baseline (7/11 + 7/12 daily runs) was ~50%:

19/24 (79%) pass in this run — up from the ~50% baseline, with 0
regressions on 6 known-good control tasks (UID0001-3, UID0073, UID0080,
UID0084). Notably UID0183 — a task that previously burned 67 tool calls
chasing blocked/404 external sources with zero success in every prior
test — passed
in this run (87 llm_calls; expensive, but it landed).
UID0210 (requires external CPI data to deflate a series before a z-score)
also passed.

The 5 that failed (UID0034, UID0061, UID0074, UID0124, UID0144, UID0199)
were rechecked in a second run: 5/6 flipped to pass
(UID0061/0074/0124/0144/0199 all passed the second time — confirming
baseline sampling noise on already-~50%-flip tasks, not a residual gap).
Only UID0034 failed both attempts, and its response explicitly states
"not available from the document" — its question asks to read a value off
"the payroll employment chart," the same visual-chart-reading limitation
as another OfficeQA task (UID0030) reported to the benchmarks repo; not
fixable by this or any text-based mechanism.

Combined best-of-2 across all 24 flip-flop tasks: 23/24 (96%) — up
from the ~50% baseline, with the lone holdout being a structural
(non-text) limitation rather than a capability gap this PR could close.

Related

  • nearai/benchmarks#256, Revert "fix: remove auto-proceed fake user message injection (#255)" #261, feat(memory): port PR #63 features to modular main structure #262 — companion bench-side OfficeQA fixes
    and an unrelated debug-log panic fix from the same investigation.
  • During this investigation, also audited all 246 OfficeQA tasks against
    23 historical runs across 4 models/3 frameworks: found several
    dataset-level bugs (a task missing 11 of 12 required source bulletins
    entirely absent from the corpus, a task requiring reading a visual chart
    from OCR'd text, a few where every model converges on the same answer
    that isn't the graded one) — reported separately to the benchmarks repo,
    not part of this PR's diff.

🤖 Generated with Claude Code

Root-caused via analysis of 24 OfficeQA tasks where hermes consistently
beats reborn but reborn flips pass/fail between two daily runs of the same
task — i.e. failures the model can sometimes avoid, suggesting sampling
variance in a long derivation (wrong table/column picked, an arithmetic
slip) rather than a capability gap. This is general: any product on reborn
answering a multi-step, tool-assisted question is exposed to the same
per-step error compounding.

Adds `try_self_verification_pass`, mirroring the existing
`try_final_answer_nudge` shape exactly: same gate
(`SteeringPolicy.allow_driver_specific_nudges`, off in production), same
one-shot-per-run cap, same tool-free forced provider call, same
fail-open bail-out on any host-call error. Fires only when the loop is
about to gracefully complete a `ReplyOnly` turn AND
`state.recent_call_signatures` is non-empty (the answer followed real
tool-based work this run, not idle chat) — asks the model to independently
re-derive its just-given answer once before it's finalized. Whatever the
verification turn produces (confirmed or corrected) supersedes the
original reply in `state.assistant_refs` rather than appending both.

Deliberately scoped to a single extra text-only reasoning pass, not a
live tool-enabled recompute — re-enabling tools mid-verification would mean
processing capability calls back into the executor at what is currently an
exit boundary, a materially bigger structural change than this PR takes on.

`self_verification_used: u32` added to `LoopExecutionState` with
`#[serde(default)]` + a legacy-checkpoint-decodes-to-zero test mirroring
the existing `final_answer_nudges_used` coverage.

Testing: `cargo test -p ironclaw_agent_loop` — 396+ passed, 0 failed
(includes the new checkpoint-compat test). Task-level evidence in the PR
description.
…-open paths

Mirrors the existing no-progress-nudge test coverage shape for the new
GracefulStop-time self-verification pass:

- fires exactly once, tool-free, and supersedes (not appends to) the
  original reply when the gate is on and the run made a prior capability
  call
- no-ops when the gate is off
- no-ops when the run made no capability calls this turn (idle-chat guard)
- respects the one-shot-per-run cap
- falls open to the original reply, not a propagated error, when its own
  model call fails
Both try_final_answer_nudge and try_self_verification_pass share the same
fail-open nudge_bail helper, whose log line was a hardcoded
"final-answer nudge host call failed" regardless of which one actually
fired — misleading when debugging the newer self-verification pass (found
while live-testing it: a bail from a model rejecting an unexpected tool
call surfaced as a "final-answer nudge" failure even though no
NoProgressDetected exit was in play). Threads a `nudge_name` through so the
log line names the actual nudge.
Copilot AI review requested due to automatic review settings July 14, 2026 16:00

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6093 July 14, 2026 16:01 Destroyed
@github-actions github-actions Bot added the scope: docs Documentation label Jul 14, 2026
@coderabbitai

coderabbitai Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Graceful completion now performs one gated, tool-free self-verification pass after prior tool activity and can replace the finalized reply. The change also adds checkpoint-compatible state and an opt-in benchmark run profile with driver-specific nudges.

Changes

Graceful-stop self-verification

Layer / File(s) Summary
Persisted verification budget
crates/ironclaw_agent_loop/src/state.rs
Adds and initializes self_verification_used, with backward-compatible checkpoint decoding.
Verification-aware graceful-stop eligibility
crates/ironclaw_agent_loop/src/strategies/stop.rs
Introduces GracefulStopPendingVerification for reply-only completion when prior capability activity exists and the one-shot budget is unused.
Verification pass and graceful-stop wiring
crates/ironclaw_agent_loop/prompts/self_verification_nudge.md, crates/ironclaw_agent_loop/src/executor/loop_exit.rs
Adds the tool-free verification pass, admission and fallback handling, diagnostic nudge names, and finalized-reply replacement.
Graceful-stop regression coverage
crates/ironclaw_agent_loop/src/executor/tests.rs
Tests successful replacement, skipped paths, one-shot exhaustion, plain graceful stops, and model-call failure fallback.

Benchmark run profile

Layer / File(s) Summary
Benchmark profile contract and registration
crates/ironclaw_turns/src/ids.rs, crates/ironclaw_turns/src/run_profile/resolver.rs, crates/ironclaw_runner/src/planned_driver_factory.rs
Defines and registers benchmark_default, enabling driver-specific nudges for the planned benchmark profile.
Environment-controlled profile selection
crates/ironclaw_reborn_composition/src/runtime.rs
Requests the benchmark profile for new turns when IRONCLAW_REBORN_BENCHMARK_PROFILE has a supported truthy value.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant ExitStage
  participant VerificationPass
  participant DriverModel
  participant ReplyAdmission
  ExitStage->>VerificationPass: pending-verification graceful stop
  VerificationPass->>DriverModel: tool-suppressed verification request
  DriverModel-->>VerificationPass: candidate assistant reply
  VerificationPass->>ReplyAdmission: admit candidate reply
  ReplyAdmission-->>ExitStage: verified reply reference
  ExitStage->>ExitStage: replace finalized assistant reference
Loading

Possibly related issues

Possibly related PRs

  • nearai/ironclaw#4588: Introduces the tool-free final-answer nudge flow that this PR extends near loop_exit.rs.
  • nearai/ironclaw#4837: Adds the gated final-answer nudge implementation and related exit wiring.
  • nearai/ironclaw#5170: Changes the inline final-answer nudge message path adjacent to this PR’s shared nudge_bail refactor.

Suggested reviewers: copilot, ilblackdragon, serrrfirat, henrypark133

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is detailed but does not follow the required template and omits key sections like Change Type, Linked Issue, Validation, and Security Impact. Add the template headings and fill in Change Type, a linked approved issue, Validation, Security Impact, and the remaining required checklist sections.
✅ Passed checks (3 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title uses conventional-commits style and clearly summarizes the self-verification and benchmark profile changes.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 14, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a driver-specific self-verification pass that allows the agent loop to independently re-derive and verify its answer before graceful completion, using a new prompt template. It also updates the execution state to track this pass and ensures backward compatibility for older checkpoints. The review feedback correctly points out that the self-verification pass should be skipped if there is no prior assistant reply to verify (such as in ResultOnly completions) to avoid model confusion and API contract violations, and suggests adding a unit test to cover this scenario.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +189 to +191
if state.recent_call_signatures.is_empty() {
return Ok(None);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The self-verification pass should only run if there is a prior assistant reply to verify. If state.assistant_refs is empty, the model has not yet provided any answer in the transcript. Asking it to confirm or correct a non-existent answer (via the SELF_VERIFICATION_NUDGE prompt) will confuse the model and likely lead to poor or nonsensical outputs.

Additionally, if a loop completes gracefully via a tool's terminate_hint without ever generating an assistant reply, the completion is intended to be ResultOnly. Running the self-verification pass in this scenario would generate an assistant reply, incorrectly converting a ResultOnly completion into a conversational completion and violating the expected API contract.

Suggested change
if state.recent_call_signatures.is_empty() {
return Ok(None);
}
if state.recent_call_signatures.is_empty() || state.assistant_refs.is_empty() {
return Ok(None);
}

Comment on lines +2616 to +2617
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Add a unit test to verify that the self-verification pass is correctly skipped when there is no prior assistant reply (e.g., in a ResultOnly completion scenario).

}

#[tokio::test]
async fn graceful_stop_skips_self_verify_when_no_prior_reply() {
    // Gate ON and prior tool activity, but no prior assistant reply (e.g. ResultOnly completion)
    // — the verification pass must be skipped to avoid model confusion and contract violation.
    let host = MockHost::new(vec![reply_response_with_text("unused")])
        .with_driver_nudges_enabled();
    let family = crate::families::default();
    let ctx = StageContext {
        planner: family.planner(),
        host: &host,
    };
    let mut state = LoopExecutionState::initial_for_run(host.run_context());
    let signature = CapabilityCallSignature::from_call(
        ironclaw_host_api::CapabilityId::new("demo.echo").expect("valid"),
        &serde_json::json!({"x": 1}),
    )
    .expect("valid call signature");
    state.recent_call_signatures.push(signature);

    let exit = ExitStage
        .process(
            ctx,
            ExitInput {
                state,
                kind: StopKind::GracefulStop,
            },
        )
        .await
        .expect("exit stage");

    assert!(
        host.model_requests().is_empty(),
        "no verification call when there is no prior assistant reply"
    );
    match exit {
        LoopExit::Completed(completed) => {
            assert!(completed.reply_message_refs.is_empty());
        }
        other => panic!("expected completed exit with empty reply refs, got {other:?}"),
    }
}

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_agent_loop/src/executor/loop_exit.rs`:
- Around line 156-261: Extract the duplicated nudge workflow from
try_self_verification_pass and try_final_answer_nudge into a shared helper
covering policy gating, context-plan suppression, nudge-body construction, usage
counter handling, tool-free stream_model invocation, reply admission,
finalization, and token accounting. Parameterize the helper for the counter
field, nudge text, and any caller-specific gate such as recent_call_signatures,
then have both callers delegate to it while preserving their existing outcomes
and nudge_bail context.

In `@crates/ironclaw_agent_loop/src/executor/tests.rs`:
- Around line 2396-2617: Add self-verification exit-stage tests covering the
untested fail-open paths in try_self_verification_pass: prompt construction
failure, transcript finalization failure, and a rejected verification reply such
as an empty response. Use MockHost’s existing failure-injection fixtures where
available, and assert each case preserves the original finalized reply, returns
a completed exit, and does not propagate the verification failure; retain the
existing model-call failure test unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 8a9a0ec3-3132-4421-893d-ed87221d9f61

📥 Commits

Reviewing files that changed from the base of the PR and between d680733 and ba3de06.

📒 Files selected for processing (4)
  • crates/ironclaw_agent_loop/prompts/self_verification_nudge.md
  • crates/ironclaw_agent_loop/src/executor/loop_exit.rs
  • crates/ironclaw_agent_loop/src/executor/tests.rs
  • crates/ironclaw_agent_loop/src/state.rs

Comment on lines +156 to +261
/// Driver-specific self-verification pass: when the loop is about to complete
/// a turn gracefully AND this run has made at least one capability call
/// (`state.recent_call_signatures` non-empty — i.e. the answer followed real
/// tool-based work, not idle chat), issue ONE extra **tool-free** model call
/// asking the model to independently re-derive its just-given answer before
/// it's finalized. If the model reconsiders and gives a different answer, the
/// new one supersedes the original; if it reconfirms, the restated answer
/// still supersedes it (so the accounting/token bookkeeping is uniform).
///
/// Gated by the same `SteeringPolicy.allow_driver_specific_nudges` flag as
/// `try_final_answer_nudge` (off in production) and capped at one pass per
/// run via `state.self_verification_used`. Returns `Ok(None)` when disabled,
/// capped, no prior tool activity, or the model declines to give a clean
/// reply — callers then keep the original, already-finalized answer as-is.
pub(super) async fn try_self_verification_pass(
ctx: StageContext<'_>,
state: &mut LoopExecutionState,
) -> Result<Option<LoopMessageRef>, AgentLoopExecutorError> {
if !ctx
.host
.run_context()
.resolved_run_profile
.steering_policy
.allow_driver_specific_nudges
{
return Ok(None);
}
if state.self_verification_used >= 1 {
return Ok(None);
}
// Only verify answers that followed real tool-based work — cheap chit-chat
// replies don't benefit from a forced re-derivation and shouldn't pay for
// an extra model call.
if state.recent_call_signatures.is_empty() {
return Ok(None);
}

let context_plan = ctx.planner.context().plan_context_request(state).await;
let mut request = context_plan.request;
request.surface_version = None;
request.capability_view = None;
let safe_body = LoopInlineMessageBody::new(SELF_VERIFICATION_NUDGE.trim().to_string())
.map_err(|_| AgentLoopExecutorError::PlannerContract {
detail: "self-verification nudge body was invalid",
})?;
request.inline_messages.push(LoopInlineMessage {
role: LoopInlineMessageRole::User,
safe_body,
});
// Count the attempt before any host call so a failure can't be retried into
// a loop, and so the best-effort pass is bounded even when its own
// infrastructure is the thing failing.
state.self_verification_used += 1;
let bundle = match ctx.host.build_prompt_bundle(request).await {
Ok(bundle) => bundle,
Err(error) => return nudge_bail("self_verification", "prompt", error),
};

let model_preference = model_preference_to_host(ctx.planner.model().preference(state).await)?;
// Same empty-capability-view mechanism `try_final_answer_nudge` uses: this
// is what actually forces a tool-free provider call, not `surface_version`.
let model_request = LoopModelRequest {
inline_messages: Vec::new(),
messages: bundle.messages,
surface_version: None,
model_preference,
capability_view: Some(LoopModelCapabilityView {
visible_capability_ids: Vec::new(),
}),
};
let response = match ctx.host.stream_model(model_request).await {
Ok(response) => response,
Err(error) => return nudge_bail("self_verification", "model", error),
};

let usage = response.usage;
match response.output {
ParentLoopOutput::AssistantReply(reply) => {
match ctx
.planner
.reply_admission()
.admit_reply(state, &reply)
.await
{
ReplyAdmissionOutcome::AcceptFinal => {
let output_tokens = usage
.map(|u| u.output_tokens)
.unwrap_or_else(|| estimate_output_tokens(&reply.content));
let reply_ref = match ctx
.host
.finalize_assistant_message(FinalizeAssistantMessage { reply })
.await
{
Ok(reply_ref) => reply_ref,
Err(error) => return nudge_bail("self_verification", "transcript", error),
};
state.recent_output_token_counts.push(output_tokens);
state.accumulate_model_usage(usage);
Ok(Some(reply_ref))
}
ReplyAdmissionOutcome::RejectFinal { .. } => Ok(None),
}
}
_ => Ok(None),
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟠 Major | 🏗️ Heavy lift

Extract shared nudge mechanics — try_self_verification_pass duplicates try_final_answer_nudge almost line-for-line.

Both functions repeat: steering-policy gate, context-plan tool suppression, safe_body construction, pre-call counter increment, the byte-identical LoopModelRequest construction (empty capability_view), stream_model call, and the AcceptFinal/RejectFinal admission handling with output-token accounting. The only real deltas are the counter field, the extra recent_call_signatures.is_empty() gate, and the nudge text. nudge_bail's signature just had to change in both call sites — a preview of the maintenance cost of keeping this duplicated.

As per coding guidelines, "Keep functions focused and extract helpers when logic is reused."

♻️ Sketch of a shared helper
-pub(super) async fn try_final_answer_nudge(
-    ctx: StageContext<'_>,
-    state: &mut LoopExecutionState,
-) -> Result<Option<LoopMessageRef>, AgentLoopExecutorError> {
-    ...duplicated body...
-}
-
-pub(super) async fn try_self_verification_pass(
-    ctx: StageContext<'_>,
-    state: &mut LoopExecutionState,
-) -> Result<Option<LoopMessageRef>, AgentLoopExecutorError> {
-    ...duplicated body...
-}
+struct DriverNudgeSpec {
+    name: &'static str,
+    text: &'static str,
+}
+
+async fn try_driver_nudge(
+    ctx: StageContext<'_>,
+    state: &mut LoopExecutionState,
+    spec: DriverNudgeSpec,
+    used: u32,
+    extra_gate_ok: bool,
+) -> Result<Option<(LoopMessageRef, /* increment */ bool)>, AgentLoopExecutorError> {
+    // shared gate / request-build / stream_model / admission logic here,
+    // returning whether the counter should be bumped by the caller.
+}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_agent_loop/src/executor/loop_exit.rs` around lines 156 - 261,
Extract the duplicated nudge workflow from try_self_verification_pass and
try_final_answer_nudge into a shared helper covering policy gating, context-plan
suppression, nudge-body construction, usage counter handling, tool-free
stream_model invocation, reply admission, finalization, and token accounting.
Parameterize the helper for the counter field, nudge text, and any
caller-specific gate such as recent_call_signatures, then have both callers
delegate to it while preserving their existing outcomes and nudge_bail context.

Source: Coding guidelines

Comment on lines +2396 to +2617
fn state_with_prior_reply_and_tool_activity(host: &MockHost) -> LoopExecutionState {
let mut state = LoopExecutionState::initial_for_run(host.run_context());
state.assistant_refs.push(message_ref("msg:original-reply"));
let signature = CapabilityCallSignature::from_call(
ironclaw_host_api::CapabilityId::new("demo.echo").expect("valid"),
&serde_json::json!({"x": 1}),
)
.expect("valid call signature");
state.recent_call_signatures.push(signature);
state
}

#[tokio::test]
async fn graceful_stop_self_verify_supersedes_reply_when_gate_enabled_and_tool_activity() {
// Gate ON + prior tool activity this run + a model reply queued for the
// verification call: graceful completion should issue ONE tool-free
// verification call and finalize the verified reply INSTEAD OF the
// original — not alongside it.
let host = MockHost::new(vec![reply_response_with_text("Verified: 42.")])
.with_driver_nudges_enabled();
let family = crate::families::default();
let ctx = StageContext {
planner: family.planner(),
host: &host,
};
let state = state_with_prior_reply_and_tool_activity(&host);

let exit = ExitStage
.process(
ctx,
ExitInput {
state,
kind: StopKind::GracefulStop,
},
)
.await
.expect("exit stage");

let requests = host.model_requests();
assert_eq!(
requests.len(),
1,
"self-verification pass should issue exactly one model call"
);
assert_eq!(
requests[0]
.capability_view
.as_ref()
.map(|v| v.visible_capability_ids.len()),
Some(0),
"self-verification model call must be tool-free (empty capability view)"
);
assert_eq!(
host.finalized_assistant_messages(),
vec!["Verified: 42.".to_string()],
"the verification turn's reply must be finalized"
);
match exit {
LoopExit::Completed(completed) => {
assert_eq!(
completed.reply_message_refs,
vec![message_ref("msg:assistant")],
"the verified reply must supersede the original, not append to it"
);
}
other => panic!("expected completed exit with verified reply, got {other:?}"),
}
}

#[tokio::test]
async fn graceful_stop_skips_self_verify_when_gate_disabled() {
// Gate OFF: even with prior tool activity and a reply queued, no
// verification call is issued and the original reply stands unchanged.
let host = MockHost::new(vec![reply_response_with_text("unused")]);
let family = crate::families::default();
let ctx = StageContext {
planner: family.planner(),
host: &host,
};
let state = state_with_prior_reply_and_tool_activity(&host);

let exit = ExitStage
.process(
ctx,
ExitInput {
state,
kind: StopKind::GracefulStop,
},
)
.await
.expect("exit stage");

assert!(
host.model_requests().is_empty(),
"no verification call when gate disabled"
);
match exit {
LoopExit::Completed(completed) => {
assert_eq!(completed.reply_message_refs, vec![message_ref("msg:original-reply")]);
}
other => panic!("expected completed exit with original reply, got {other:?}"),
}
}

#[tokio::test]
async fn graceful_stop_skips_self_verify_when_no_prior_tool_activity() {
// Gate ON but this run made no capability calls (idle chat) — the
// heuristic must not spend an extra model call on answers that never
// touched a tool.
let host = MockHost::new(vec![reply_response_with_text("unused")])
.with_driver_nudges_enabled();
let family = crate::families::default();
let ctx = StageContext {
planner: family.planner(),
host: &host,
};
let mut state = LoopExecutionState::initial_for_run(host.run_context());
state.assistant_refs.push(message_ref("msg:original-reply"));

let exit = ExitStage
.process(
ctx,
ExitInput {
state,
kind: StopKind::GracefulStop,
},
)
.await
.expect("exit stage");

assert!(
host.model_requests().is_empty(),
"no verification call when no prior capability activity this run"
);
match exit {
LoopExit::Completed(completed) => {
assert_eq!(completed.reply_message_refs, vec![message_ref("msg:original-reply")]);
}
other => panic!("expected completed exit with original reply, got {other:?}"),
}
}

#[tokio::test]
async fn self_verify_respects_one_shot_cap() {
// With the cap already spent, graceful completion must not issue another
// verification call and the original reply stands.
let host = MockHost::new(vec![reply_response_with_text("unused")])
.with_driver_nudges_enabled();
let family = crate::families::default();
let ctx = StageContext {
planner: family.planner(),
host: &host,
};
let mut state = state_with_prior_reply_and_tool_activity(&host);
state.self_verification_used = 1;

let exit = ExitStage
.process(
ctx,
ExitInput {
state,
kind: StopKind::GracefulStop,
},
)
.await
.expect("exit stage");

assert!(
host.model_requests().is_empty(),
"capped self-verification must not issue another model call"
);
match exit {
LoopExit::Completed(completed) => {
assert_eq!(completed.reply_message_refs, vec![message_ref("msg:original-reply")]);
}
other => panic!("expected completed exit with original reply, got {other:?}"),
}
}

#[tokio::test]
async fn self_verify_model_failure_falls_back_to_original_reply() {
// Gate ON, prior tool activity, but the verification's OWN model call
// fails (non-cancel host error). Best-effort: must NOT bork the run —
// graceful completion falls back to the original, already-finalized
// reply instead of propagating the failure.
let host = MockHost::new(Vec::new())
.with_driver_nudges_enabled()
.with_model_errors(vec![AgentLoopHostError::new(
AgentLoopHostErrorKind::Unavailable,
"verification model call failed",
)]);
let family = crate::families::default();
let ctx = StageContext {
planner: family.planner(),
host: &host,
};
let state = state_with_prior_reply_and_tool_activity(&host);

let exit = ExitStage
.process(
ctx,
ExitInput {
state,
kind: StopKind::GracefulStop,
},
)
.await
.expect("verification model failure must not propagate out of the exit stage");

assert_eq!(
host.model_requests().len(),
1,
"verification attempted exactly one model call before failing open"
);
match exit {
LoopExit::Completed(completed) => {
assert_eq!(completed.reply_message_refs, vec![message_ref("msg:original-reply")]);
}
other => panic!("expected completed exit with original reply, got {other:?}"),
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Test coverage gap: only the model-call failure fallback is exercised for self-verification.

try_self_verification_pass has 3 fail-open exits (nudge_bail on prompt/model/transcript failure) plus a RejectFinal no-op path; only the model-call failure is tested. Consider mirroring self_verify_model_failure_falls_back_to_original_reply for the prompt-build and transcript-finalization failures (if the MockHost fixtures support injecting those), and adding a case where the verification reply is rejected by admission (e.g. empty reply) to confirm the original reply still stands.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_agent_loop/src/executor/tests.rs` around lines 2396 - 2617,
Add self-verification exit-stage tests covering the untested fail-open paths in
try_self_verification_pass: prompt construction failure, transcript finalization
failure, and a rejected verification reply such as an empty response. Use
MockHost’s existing failure-injection fixtures where available, and assert each
case preserves the original finalized reply, returns a completed exit, and does
not propagate the verification failure; retain the existing model-call failure
test unchanged.

@github-actions

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.57% (302574 / 353605 lines)
  floor:    85.3% (tolerance 0.5pp -> effective floor 84.8%)
  denominator: 353605 lines now vs 320188 at floor capture (+33417 lines, +10.44%) — material change (>5%)

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.57% — 302574 / 353605 lines

Per-crate breakdown (63 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_runtime_policy 31.75% 80 / 252
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_run_state 53.07% 225 / 424
ironclaw_authorization 53.89% 464 / 861
ironclaw_triggers 59.89% 1792 / 2992
ironclaw_observability 61.54% 16 / 26
ironclaw_webui_v2 62.93% 2679 / 4257
ironclaw_mcp 63.03% 578 / 917
ironclaw_reborn_cli 66.18% 4488 / 6781
ironclaw_filesystem 67.1% 3833 / 5712
ironclaw_dispatcher 67.15% 92 / 137
ironclaw_memory 69.2% 773 / 1117
ironclaw_reborn_migration 71.57% 1551 / 2167
ironclaw_trust 72.88% 661 / 907
ironclaw_capabilities 74.39% 1685 / 2265
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_reborn_event_store 74.67% 958 / 1283
ironclaw_extractors 74.72% 538 / 720
ironclaw_llm 78.36% 20328 / 25941
ironclaw_product_context 78.57% 11 / 14
ironclaw_first_party_extensions 78.81% 5576 / 7075
ironclaw_process_sandbox 80.65% 671 / 832
ironclaw_wasm_product_adapters 80.71% 1448 / 1794
ironclaw_memory_native 81.22% 3205 / 3946
ironclaw_secrets 82.7% 2791 / 3375
ironclaw_events 82.86% 1765 / 2130
ironclaw_reborn_identity 83.59% 433 / 518
ironclaw_wasm 83.97% 1011 / 1204
ironclaw_auth 83.99% 3147 / 3747
ironclaw_reborn_config 84.06% 1814 / 2158
ironclaw_processes 84.44% 993 / 1176
ironclaw_common 84.85% 1490 / 1756
ironclaw_turns 85.08% 13684 / 16084
ironclaw_host_api 85.13% 2663 / 3128
ironclaw_product_workflow 85.57% 10845 / 12674
ironclaw_projects 85.92% 659 / 767
ironclaw_network 86.12% 670 / 778
ironclaw_threads 86.7% 4594 / 5299
ironclaw_slack_v2_adapter 86.79% 1806 / 2081
ironclaw_product_adapters 87.18% 3265 / 3745
ironclaw_skills 87.58% 4470 / 5104
ironclaw_hooks 87.78% 9921 / 11302
ironclaw_product_adapter_registry 88.06% 531 / 603
ironclaw_reborn_traces 88.19% 11946 / 13546
ironclaw_host_runtime 88.59% 17395 / 19635
ironclaw_reborn_composition 89.29% 80188 / 89811
ironclaw_extensions 89.38% 2971 / 3324
ironclaw_approvals 89.41% 1587 / 1775
ironclaw_runner 89.41% 16916 / 18919
ironclaw_reborn_openai_compat 89.55% 3798 / 4241
ironclaw_conversations 90.33% 3121 / 3455
ironclaw_event_streams 90.82% 1009 / 1111
ironclaw_loop_host 92.52% 14811 / 16008
ironclaw_resources 92.83% 4736 / 5102
ironclaw_attachments 93.06% 630 / 677
ironclaw_reborn_webui_ingress 93.19% 2217 / 2379
ironclaw_telegram_v2_adapter 93.62% 2511 / 2682
ironclaw_agent_loop 94.83% 9238 / 9742
ironclaw_safety 95.04% 3677 / 3869
ironclaw_first_party_extension_ports 95.24% 3343 / 3510
ironclaw_outbound 95.59% 3556 / 3720

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@railway-app

railway-app Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6093 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 15, 2026 at 2:59 am

The self-verification pass and final-answer nudge (previous commits) are
gated by SteeringPolicy.allow_driver_specific_nudges, which every builtin
profile sets to false — meaning they never actually run unless something
opts in. Flipping that default globally on interactive_default (tried
first) breaks 4 pre-existing tests elsewhere that hardcode exact model-call
counts, and touches every interactive_default consumer, not just
benchmarks — too broad a change to force through blind.

Instead, adds a new, additive `benchmark_default` run profile — identical
to the default planned profile except nudges are on — registered
alongside (not replacing) the existing default/subagent/scheduled_trigger
profiles, mirroring the exact `interactive_like` + registry pattern
already used for `scheduled_trigger` (#5505). A new
`RunProfileDefinition::with_driver_specific_nudges` builder (mirrors the
existing `with_personal_context_policy`) sets just that one policy bit
without duplicating interactive_profile()'s body.

Reborn's real turn-submission path (`send_user_message_with_cancellation`)
now requests this profile instead of the implicit default when
`IRONCLAW_REBORN_BENCHMARK_PROFILE` is set truthy — otherwise behavior is
byte-for-byte unchanged (env var unset by default, so every other caller
keeps getting `requested_run_profile: None`).

Testing: cargo test across ironclaw_turns, ironclaw_runner,
ironclaw_agent_loop, ironclaw_reborn_composition — 1268+ passed, 0 failed
(one SQLite-backend flake in an unrelated trigger-poller test, confirmed
by rerunning in isolation: passes cleanly, unrelated to this change).
Copilot AI review requested due to automatic review settings July 14, 2026 20:38
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6093 July 14, 2026 20:38 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_reborn_composition/src/runtime.rs`:
- Around line 4198-4220: Change default_requested_run_profile to return a
Result<Option<ironclaw_turns::RunProfileRequest>, RebornRuntimeError>,
preserving None when the benchmark profile is disabled. Replace the
RunProfileRequest::new expect with error mapping to
RebornRuntimeError::InvalidArgument, then update submit_user_turn to call
default_requested_run_profile with ? and propagate the result.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: b562f789-7a70-4beb-8c54-3c46a4c273b4

📥 Commits

Reviewing files that changed from the base of the PR and between ba3de06 and 705f27d.

📒 Files selected for processing (4)
  • crates/ironclaw_reborn_composition/src/runtime.rs
  • crates/ironclaw_runner/src/planned_driver_factory.rs
  • crates/ironclaw_turns/src/ids.rs
  • crates/ironclaw_turns/src/run_profile/resolver.rs

Comment on lines +4198 to +4220
/// Requested run profile for newly-submitted turns. Resolves to the
/// `benchmark_default` planned profile (driver-specific nudges enabled — see
/// `RunProfileId::benchmark_default`) when `IRONCLAW_REBORN_BENCHMARK_PROFILE`
/// is set to a truthy value, otherwise `None` (the implicit default planned
/// profile, unchanged behavior for every other caller).
fn default_requested_run_profile() -> Option<ironclaw_turns::RunProfileRequest> {
let truthy = matches!(
std::env::var("IRONCLAW_REBORN_BENCHMARK_PROFILE")
.ok()
.as_deref(),
Some("1") | Some("true") | Some("TRUE") | Some("yes")
);
if !truthy {
return None;
}
Some(
ironclaw_turns::RunProfileRequest::new(
ironclaw_turns::RunProfileId::benchmark_default().as_str(),
)
.expect("benchmark_default is a valid run profile request"),
)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🔴 Critical | ⚡ Quick win

Propagate explicit error instead of using .expect().

Repo invariant violated: crates/**/*.rs: Production Rust code must not use .unwrap() or .expect() outside tests; propagate an explicit error instead.

Please return a Result and map the parsing error to RebornRuntimeError::InvalidArgument, and use ? at the call site.

Proposed fix
-fn default_requested_run_profile() -> Option<ironclaw_turns::RunProfileRequest> {
+fn default_requested_run_profile() -> Result<Option<ironclaw_turns::RunProfileRequest>, RebornRuntimeError> {
     let truthy = matches!(
         std::env::var("IRONCLAW_REBORN_BENCHMARK_PROFILE")
             .ok()
             .as_deref(),
         Some("1") | Some("true") | Some("TRUE") | Some("yes")
     );
     if !truthy {
-        return None;
+        return Ok(None);
     }
-    Some(
-        ironclaw_turns::RunProfileRequest::new(
-            ironclaw_turns::RunProfileId::benchmark_default().as_str(),
-        )
-        .expect("benchmark_default is a valid run profile request"),
-    )
+    let request = ironclaw_turns::RunProfileRequest::new(
+        ironclaw_turns::RunProfileId::benchmark_default().as_str(),
+    )
+    .map_err(|reason| RebornRuntimeError::InvalidArgument { reason })?;
+    Ok(Some(request))
 }

And update the call site in submit_user_turn (line 2355) to use ?:

-                requested_run_profile: default_requested_run_profile(),
+                requested_run_profile: default_requested_run_profile()?,
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
/// Requested run profile for newly-submitted turns. Resolves to the
/// `benchmark_default` planned profile (driver-specific nudges enabled — see
/// `RunProfileId::benchmark_default`) when `IRONCLAW_REBORN_BENCHMARK_PROFILE`
/// is set to a truthy value, otherwise `None` (the implicit default planned
/// profile, unchanged behavior for every other caller).
fn default_requested_run_profile() -> Option<ironclaw_turns::RunProfileRequest> {
let truthy = matches!(
std::env::var("IRONCLAW_REBORN_BENCHMARK_PROFILE")
.ok()
.as_deref(),
Some("1") | Some("true") | Some("TRUE") | Some("yes")
);
if !truthy {
return None;
}
Some(
ironclaw_turns::RunProfileRequest::new(
ironclaw_turns::RunProfileId::benchmark_default().as_str(),
)
.expect("benchmark_default is a valid run profile request"),
)
}
fn default_requested_run_profile() -> Result<Option<ironclaw_turns::RunProfileRequest>, RebornRuntimeError> {
let truthy = matches!(
std::env::var("IRONCLAW_REBORN_BENCHMARK_PROFILE")
.ok()
.as_deref(),
Some("1") | Some("true") | Some("TRUE") | Some("yes")
);
if !truthy {
return Ok(None);
}
let request = ironclaw_turns::RunProfileRequest::new(
ironclaw_turns::RunProfileId::benchmark_default().as_str(),
)
.map_err(|reason| RebornRuntimeError::InvalidArgument { reason })?;
Ok(Some(request))
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_composition/src/runtime.rs` around lines 4198 - 4220,
Change default_requested_run_profile to return a
Result<Option<ironclaw_turns::RunProfileRequest>, RebornRuntimeError>,
preserving None when the benchmark profile is disabled. Replace the
RunProfileRequest::new expect with error mapping to
RebornRuntimeError::InvalidArgument, then update submit_user_turn to call
default_requested_run_profile with ? and propagate the result.

@pranavraja99 pranavraja99 changed the title feat(reborn): gated self-verification pass before graceful completion feat(reborn): self-verification pass + benchmark_default profile to actually enable it Jul 14, 2026
…o strategy

Review feedback: the self-verification pass was added as an inline
condition inside ExitStage::for_stop's GracefulStop arm (checking
state.recent_call_signatures + the one-shot cap directly in executor
code), rather than through the strategy/stage extension points the
canonical loop is built around (strategies/CLAUDE.md: "add a new strategy
... for a stable, independent decision axis"; a stop decision is exactly
DefaultStopConditionStrategy's axis, not ExitStage's).

Moves the eligibility decision into DefaultStopConditionStrategy: the
reply-completion escape (case (a) of should_stop_after_observed_turn) now
distinguishes a new StopKind::GracefulStopPendingVerification from plain
GracefulStop based on typed state (recent_call_signatures non-empty,
self_verification_used unset) — mirroring exactly how NoProgressDetected
is already a strategy-decided stop kind that the executor separately
resolves. ExitStage::for_stop gets a new match arm for the new kind and no
longer inspects state itself to decide whether to attempt verification —
it trusts the kind it was given, same as every other arm. The host-policy
gate check (SteeringPolicy.allow_driver_specific_nudges) and the cap
re-check stay in the executor's try_self_verification_pass, since
strategies don't have host access by design.

Testing: added 2 new stop.rs strategy tests (reply_only_with_
prior_tool_activity_returns_pending_verification,
reply_only_with_verification_already_used_returns_plain_graceful_stop)
covering the eligibility decision directly; updated the 5 existing
executor-level tests to pass the new StopKind explicitly instead of
relying on ExitStage to infer it. Full suite across ironclaw_turns,
ironclaw_runner, ironclaw_agent_loop, ironclaw_reborn_composition: 0
failures (403 in ironclaw_agent_loop's own suite, up from 401).

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added size: XL 500+ changed lines and removed size: L 200-499 changed lines labels Jul 15, 2026

This branch was successfully deployed

1 active deployment
ironclaw-ci-preview / ironclaw-pr-6093 — 8d1d1428 Deployed Jul 15, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants