Skip to content

fix(runner): stop silently retrying model-stage failures that cannot succeed - #6824

Merged
serrrfirat merged 3 commits into
mainfrom
claude/close-the-epic-tails
Jul 29, 2026
Merged

serrrfirat merged 3 commits into
mainfrom
claude/close-the-epic-tails

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

Closing #6284 WS1's remaining mapping box. I went in expecting a naming tidy-up and found a live retry-burn.

model_stage_failure_category returned None for InvalidInvocation, Invalid, ScopeMismatch and PolicyDenied. All four fell through to host_stage_unavailable_model, which is_auto_retriable_category lists among "transient host / lease / store / provider / tool faults where re-running the identical request from the checkpoint is likely to succeed."

None of the four can succeed on an identical retry. Policy does not change between attempts, a malformed request stays malformed, a scope mismatch is configuration-shaped. So the run silently re-drove a call that could never work, and told the operator "host stage unavailable" instead of the real cause.

That list's own doc comment says "Conservative by design — anything not clearly transient falls to UserInitiated." The fallthrough defeated exactly that intent.

executor/mapping.rs documented this as handled — "the runner preserves the original kind when categorizing the failure." It did not. Corrected in place rather than left to mislead the next reader.

Change Type

  • Bug fix

Linked Issue

Related #6284 (WS1). Closes its model_failure_mapping box.

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all-targets --all-features — zero warnings
  • cargo build
  • Relevant tests: ironclaw_runner, ironclaw_agent_loop, ironclaw_turns — 1,826 passed, 0 failed
  • cargo test --features integration — not applicable: no persistence or DB behavior changed. The routing this affects is covered by the runner's own suite plus the two conformance tests below.
  • Manual testing: red-verified (below)

The fix

Three categories, none auto-retriable — an unlisted category falls to UserInitiated, which is the behavior we want:

kind category
InvalidInvocation, Invalid model_stage_request_invalid
PolicyDenied model_stage_policy_denied
ScopeMismatch model_stage_scope_mismatch

Two follow-on sites the suite caught, both worth naming

  1. failure_lane's canonical category list needed the additions. Its doc says "keep this in lockstep" and nothing enforces that except the test — worth knowing if you add categories.

  2. every_failure_category_is_explainable_and_classified rejected them for having no user-facing explanation. The generic fallback they would have used reads:

    "The run failed before producing a reply. Retry the run, and contact support if it keeps happening."

    That advice can never work for a refused or malformed request. Each new category now says plainly that retrying will not help and names what has to change.

    This is the same defect the epic fixes for the model, one layer up. The user was also being told to retry something that could not succeed. Good test — it caught a real gap rather than just a missing table entry.

Test Strategy

User behavior: a model call refused by policy, or malformed, now fails once with an accurate reason and an explanation that does not tell the user to retry. Previously it burned silent retries and reported a generic host outage.

Risk areas:

  • Model behavior
  • Cross-component behavior (driver reason_kind → retry disposition → failure lane → user summary)

New: permanent_model_stage_failures_are_not_categorized_as_transient_outages asserts each kind gets a category and that the category is not auto-retriable. The second half is the part that matters — a category alone would not have fixed the retry burn.

Two tests inverted, each with the reason recorded beside it:

  • model_stage_host_error_kind_category_matrix_is_exhaustive pinned all four at None. The epic anticipated this exactly: "a runner test pins the wrong behavior as expected — fix the mapping and invert the test."
  • text_only_model_reply_driver_sanitizes_model_failures_and_skips_transcript_write pinned the generic model_error reason kind for a model-stage PolicyDenied.

Red verification: the new test fails against the old mapping with "PolicyDenied has no model-stage category, so it falls through to the generic host-stage outage and is silently auto-retried."

Security Impact

None directly. A refused call is now reported as refused rather than as a host outage, which is more accurate to the operator and does not disclose anything the failure did not already carry.

Reborn Trust-Boundary Checklist

  • New/changed status/error variants — downstream audited: three new category strings. Traced every consumer: text_loop_driver (reason_kind), retry_disposition (falls to UserInitiated — intended), failure_lane (canonical list updated), failure_summary (explanations added). That chain is exactly what the two failing tests walked me through.
  • Trust-bearing types / untrusted content / hashes / bounds / sandbox naming: N/A — host-authored constants only.
  • serde(default): N/A.

Database Impact

None.

Blast Radius

ironclaw_runner failure categorization. A model-stage failure of these four kinds now surfaces a different reason_kind and is no longer auto-retried. Anything keying on model_error for these specific kinds will see the more specific value — that is the intent, and the audit above lists every consumer.

Worth flagging to whoever watches run-failure dashboards: some failures previously counted as host_stage_unavailable_model / model_error will move to the three new categories, and their auto-retry volume will drop. Same events, honest labels, fewer wasted calls.

Rollback Plan

Revert the commit. No migration and no persisted shape change; the categories are runtime strings.

Review Follow-Through

One judgment call worth a look: I mapped InvalidInvocation and Invalid to a single model_stage_request_invalid because the remediation is identical (fix the request). If they warrant separate operator-facing messages, splitting is a one-line change.

Scope note: I sized the other WS1/WS3 tails while here and deliberately left them out. LoopDiagnosticRef's deletion is ~70 references across four crates including public struct fields — a refactor, not a tail. The detail: None producer sweep is unbounded until someone surveys it. Neither belongs in this PR.


Review track: C (runtime)

🤖 Generated with Claude Code

…succeed

#6284 WS1's remaining mapping box. Went in expecting a naming tidy-up;
it is a live retry-burn.

`model_stage_failure_category` returned `None` for `InvalidInvocation`,
`Invalid`, `ScopeMismatch` and `PolicyDenied`, so all four fell through
to `host_stage_unavailable_model` — which `is_auto_retriable_category`
lists among "transient host / lease / store / provider / tool faults
where re-running the *identical* request from the checkpoint is likely to
succeed".

None of the four can succeed on an identical retry. Policy does not
change between attempts, a malformed request stays malformed, and a
scope mismatch is configuration-shaped. So the run silently re-drove a
call that could never work, and reported "host stage unavailable" to the
operator instead of the real cause. That same list's doc comment says
"Conservative by design — anything not clearly transient falls to
`UserInitiated`"; the fallthrough defeated exactly that intent.

`executor/mapping.rs` documented this as already handled — "the runner
preserves the original kind when categorizing the failure." It did not.
Comment corrected in place rather than left to mislead the next reader.

Three categories, none auto-retriable (an unlisted category falls to
`UserInitiated`, which is the wanted behavior):

  - `model_stage_request_invalid`  (InvalidInvocation | Invalid)
  - `model_stage_policy_denied`    (PolicyDenied)
  - `model_stage_scope_mismatch`   (ScopeMismatch)

Two follow-on sites the compiler and the suite caught, both worth naming:

  - `failure_lane`'s canonical category list needed the three additions;
    its doc says "keep this in lockstep" and nothing enforces that but
    the test.
  - `every_failure_category_is_explainable_and_classified` rejected them
    for having no user-facing explanation. The generic fallback they
    would have used reads "Retry the run, and contact support if it keeps
    happening" — advice that can never work for a refused or malformed
    request. Each new category now states plainly that retrying will not
    help and names what has to change. That is the same defect this epic
    fixes for the model, one layer up: the *user* was also being told to
    retry something that could not succeed.

Two tests inverted, each with the reason recorded beside it:

  - `model_stage_host_error_kind_category_matrix_is_exhaustive` pinned
    all four kinds at `None` — the epic anticipated this ("a runner test
    pins the wrong behavior as expected — fix the mapping and invert the
    test").
  - `text_only_model_reply_driver_sanitizes_model_failures_...` pinned
    the generic `model_error` reason kind for a model-stage
    `PolicyDenied`.

New regression test `permanent_model_stage_failures_are_not_categorized_
as_transient_outages` asserts each kind gets a category AND that the
category is not auto-retriable — the second half is the part that
actually matters, and it fails against the old code.

Gate: runner, agent_loop, turns — 1,826 tests pass, clippy clean, fmt
clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6824 July 28, 2026 19:44 Destroyed
@coderabbitai

coderabbitai Bot commented Jul 28, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@serrrfirat, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 1 minute

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 1eadddc7-8389-4c56-9e72-c1069a18fd0a

📥 Commits

Reviewing files that changed from the base of the PR and between 5f848e6 and 2ecb43d.

📒 Files selected for processing (1)
  • crates/ironclaw_runner/src/text_loop_driver.rs
📝 Walkthrough

Walkthrough

The runner now classifies permanent model-stage invalid, policy-denied, and scope-mismatch failures with dedicated non-retriable categories and summaries. Registries, mapping tests, product summary tests, host-loop expectations, and an explanatory comment were updated.

Changes

Model failure classification

Layer / File(s) Summary
Model-stage category classification
crates/ironclaw_runner/src/failure_categories.rs, crates/ironclaw_runner/src/model_failure_mapping.rs, crates/ironclaw_runner/src/retry_disposition.rs
Defines dedicated categories, maps permanent model errors to them, exposes the retry predicate within the crate, and updates exhaustive tests.
Category registration and summaries
crates/ironclaw_runner/src/failure_lane.rs, crates/ironclaw_runner/src/failure_summary.rs, crates/ironclaw_product/src/projection/tests/failure_explanation.rs
Registers the categories, adds specific failure summaries, and extends product summary coverage.
Runner integration and classification documentation
crates/ironclaw_runner/tests/loop_driver_host.rs, crates/ironclaw_agent_loop/src/executor/mapping.rs
Updates the policy-denied integration expectation and corrects the prior generic retry-classification comment.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

  • nearai/ironclaw#6449: Uses the same failure-category, retry, lane, and summary mappings.
  • nearai/ironclaw#6677: Classifies the same permanent model-stage error kinds separately from transient host outages.

Suggested reviewers: think-in-universe

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title is Conventional Commits-style and accurately summarizes the main fix in the PR.
Description check ✅ Passed The description is mostly complete and fills the required template sections with clear summary, validation, test strategy, and impact details.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 28, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_runner/src/model_failure_mapping.rs`:
- Around line 65-107: Extend the driver-level integration test that currently
covers PolicyDenied to exercise InvalidInvocation, Invalid, and ScopeMismatch
through the production caller path rather than calling
model_stage_failure_category directly. For each case, assert the emitted driver
failure category matches its specific non-transient category and verify
is_auto_retriable_category returns false.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 71e239f9-d455-4f6f-a389-1cbbdd145dbd

📥 Commits

Reviewing files that changed from the base of the PR and between 7086872 and deead5f.

📒 Files selected for processing (7)
  • crates/ironclaw_agent_loop/src/executor/mapping.rs
  • crates/ironclaw_runner/src/failure_categories.rs
  • crates/ironclaw_runner/src/failure_lane.rs
  • crates/ironclaw_runner/src/failure_summary.rs
  • crates/ironclaw_runner/src/model_failure_mapping.rs
  • crates/ironclaw_runner/src/retry_disposition.rs
  • crates/ironclaw_runner/tests/loop_driver_host.rs

Comment on lines +65 to +107
fn permanent_model_stage_failures_are_not_categorized_as_transient_outages() {
use crate::retry_disposition::is_auto_retriable_category;
use AgentLoopHostErrorKind as K;

for kind in [
K::InvalidInvocation,
K::Invalid,
K::ScopeMismatch,
K::PolicyDenied,
] {
let category = model_stage_failure_category(true, kind, None).unwrap_or_else(|| {
panic!(
"{kind:?} has no model-stage category, so it falls through to the generic \
host-stage outage and is silently auto-retried"
)
});
assert!(
!is_auto_retriable_category(category),
"{kind:?} -> {category:?} is auto-retriable, but an identical retry cannot succeed"
);
assert_ne!(
category,
crate::failure_categories::HOST_STAGE_UNAVAILABLE_MODEL_CATEGORY,
"{kind:?} must name its own cause, not a generic host outage"
);
}
}

#[test]
fn model_stage_host_error_kind_category_matrix_is_exhaustive() {
use AgentLoopHostErrorKind as K;

let expected_without_reason = |kind| match kind {
K::CredentialUnavailable => Some(MODEL_CREDENTIALS_UNAVAILABLE_CATEGORY),
K::BudgetAccountingFailed => Some(BUDGET_ACCOUNTING_FAILED_CATEGORY),
// Permanent for an identical retry — each names its own cause
// instead of falling through to the auto-retriable generic
// host-stage outage. Inverted from `None`, which pinned the bug
// `permanent_model_stage_failures_are_not_categorized_as_transient_outages`
// now guards.
K::InvalidInvocation | K::Invalid => Some(MODEL_STAGE_REQUEST_INVALID_CATEGORY),
K::PolicyDenied => Some(MODEL_STAGE_POLICY_DENIED_CATEGORY),
K::ScopeMismatch => Some(MODEL_STAGE_SCOPE_MISMATCH_CATEGORY),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Cover the remaining categories through the driver path.

This test calls the classifier directly; the integration test covers only PolicyDenied. Add caller-level cases for InvalidInvocation, Invalid, and ScopeMismatch that assert the emitted driver failure category and non-auto-retry behavior.

As per coding guidelines, “New or changed production-wired behavior must have a caller-level test”; as per path instructions, “Test through the caller” when a classifier gates a side effect.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_runner/src/model_failure_mapping.rs` around lines 65 - 107,
Extend the driver-level integration test that currently covers PolicyDenied to
exercise InvalidInvocation, Invalid, and ScopeMismatch through the production
caller path rather than calling model_stage_failure_category directly. For each
case, assert the emitted driver failure category matches its specific
non-transient category and verify is_auto_retriable_category returns false.

Sources: Coding guidelines, Path instructions

@ironloopai

ironloopai Bot commented Jul 28, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6824

🔴 Failed

Execution result is invalid

The structured result could not be verified.

Automatic · PR opened · attempt 1 of 3 · failed after 1m 56s

Failure details
  • Repository: nearai/ironclaw
  • Base: main at 7086872
  • Head: claude/close-the-epic-tails at deead5f
  • Created: Jul 28, 2026, 7:49 PM UTC
  • Updated: Jul 28, 2026, 7:51 PM UTC
  • Run: 27d8e440-e849-43ea-8163-8235c25e3fe3
  • Latest attempt: 1 · Completed · b24512b8-6070-4d24-a054-b1f200456c8a
  • Failed during: Verification
  • Retryable: No
  • Failure: c509deee-c037-4d7f-a425-b68eb6da10d6

…ary table

CI caught what my local gate missed: I ran the crates I edited
(`ironclaw_runner`, `ironclaw_agent_loop`, `ironclaw_turns`) rather than
the crates that CONSUME the value I changed.

`ironclaw_product` keeps a SECOND table of user-facing failure text, and
`failure_summary_covers_reborn_failure_category_constants` asserts it
covers exactly the public constants in `ironclaw_runner::failure_
categories`. Adding three constants without adding three entries fails
that test — correctly.

Added with the same wording used in the runner's table, so the two agree
for these categories.

Worth naming for whoever touches this next: the category CONSTANTS have a
single home (`ironclaw_product` imports them from `ironclaw_runner`), so
there is no vocabulary drift here — unlike the recovery-hint allowlist
fixed in #6792. What is duplicated is the user-facing TEXT, across
`runner::failure_summary` and `product::projection`. That duplication is
guarded by this test, which is why it surfaced in minutes instead of
silently. Whether the two tables should be one is a real question, but a
separate one: they may be deliberately different registers (operator vs
end-user), and I have not verified that either way.

Gate: ironclaw_product + ironclaw_event_projections — 526 tests pass,
including the guard that failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6824 July 28, 2026 20:36 Destroyed
@github-actions

github-actions Bot commented Jul 28, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.66% (312506 / 364802 lines)
  floor:    80.81% (tolerance 0.5pp -> effective floor 80.31%)
  denominator: 364802 lines now vs 377084 at floor capture (-12282 lines, -3.26%) — not a material change

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.66% — 312506 / 364802 lines

Per-crate breakdown (60 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_process_sandbox 33.91% 118 / 348
ironclaw_host_ingress 42.5% 17 / 40
ironclaw_event_projections 43.71% 684 / 1565
ironclaw_observability 61.54% 16 / 26
ironclaw_telegram_v2_adapter 62.35% 631 / 1012
ironclaw_authorization 62.98% 609 / 967
ironclaw_memory 64.41% 959 / 1489
ironclaw_trust 73.21% 664 / 907
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_filesystem 74.64% 4829 / 6470
ironclaw_extractors 74.72% 538 / 720
ironclaw_capabilities 75.41% 2879 / 3818
ironclaw_projects 76.48% 400 / 523
ironclaw_mcp 76.6% 779 / 1017
ironclaw_reborn_cli 78.22% 10693 / 13670
ironclaw_telegram_extension 78.59% 962 / 1224
ironclaw_wasm 79.72% 735 / 922
ironclaw_llm 79.98% 22014 / 27523
ironclaw_memory_native 81.02% 3299 / 4072
ironclaw_auth 81.88% 6679 / 8157
ironclaw_first_party_extensions 82.38% 6682 / 8111
ironclaw_events 82.56% 1600 / 1938
ironclaw_host_api 83.04% 9482 / 11418
ironclaw_processes 83.3% 933 / 1120
ironclaw_reborn_identity 83.8% 450 / 537
ironclaw_operator 84.37% 5558 / 6588
ironclaw_secrets 84.53% 2797 / 3309
ironclaw_reborn_config 85.23% 2101 / 2465
ironclaw_extension_host 85.27% 18975 / 22254
ironclaw_skills 85.27% 4493 / 5269
ironclaw_run_state 85.77% 458 / 534
ironclaw_reborn_composition 85.77% 24622 / 28706
ironclaw_triggers 85.92% 2783 / 3239
ironclaw_network 85.97% 913 / 1062
ironclaw_webui 86.1% 10918 / 12680
ironclaw_reborn_event_store 86.51% 1251 / 1446
ironclaw_hooks 86.72% 9949 / 11472
ironclaw_common 86.99% 1772 / 2037
ironclaw_approvals 87.07% 1542 / 1771
ironclaw_extensions 87.24% 4798 / 5500
ironclaw_threads 87.36% 4912 / 5623
ironclaw_product 87.74% 20169 / 22986
ironclaw_reborn_traces 88.13% 11986 / 13600
ironclaw_turns 88.44% 14453 / 16342
ironclaw_slack_extension 88.47% 1934 / 2186
ironclaw_host_runtime 88.93% 19947 / 22429
ironclaw_reborn_openai_compat 89.32% 3780 / 4232
ironclaw_conversations 90.01% 3164 / 3515
ironclaw_resources 90.84% 4474 / 4925
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_runner 91.34% 17354 / 19000
ironclaw_loop_host 91.9% 16572 / 18032
ironclaw_attachments 93.06% 630 / 677
ironclaw_outbound 93.91% 4101 / 4367
ironclaw_agent_loop 94.56% 9994 / 10569
ironclaw_safety 95.28% 3858 / 4049
ironclaw_first_party_extension_ports 95.62% 3672 / 3840
ironclaw_runtime_policy 96.56% 814 / 843

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@railway-app

railway-app Bot commented Jul 28, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6824 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 28, 2026 at 9:29 pm

Review feedback on #6824: the classifier had a direct unit test, but only
`PolicyDenied` was proven through the caller.

That distinction matters here rather than being a formality.
`map_host_error` reaches the category through an early return that
bypasses the kind match below it, so the classifier being correct does
not prove the driver emits what the classifier returns -- and the
emitted `reason_kind` is what `retry_disposition` keys on.

Covers all four permanent kinds -- `InvalidInvocation`, `Invalid`,
`ScopeMismatch`, `PolicyDenied` -- asserting each reaches the driver as
its own category and that `is_auto_retriable_category` rejects it. Before
the fix all four routed through `host_stage_unavailable_model`, which IS
auto-retriable, so a permanently-failing call was silently re-driven.

Sabotage-verified: removing the `InvalidInvocation | Invalid` arm from
the classifier now fails at the driver seam.

Refs #6524

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Fixed in 2ecb43d.

The distinction matters here rather than being a formality: map_host_error reaches the category through an early return that bypasses the kind match below it, so the classifier being correct does not prove the driver emits what the classifier returns — and the emitted reason_kind is what retry_disposition keys on.

Now covers all four permanent kinds (InvalidInvocation, Invalid, ScopeMismatch, PolicyDenied), asserting each reaches the driver as its own category and that is_auto_retriable_category rejects it.

Sabotage-verified: removing the InvalidInvocation | Invalid arm from the classifier fails at the driver seam.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6824 July 28, 2026 21:29 Destroyed
@serrrfirat
serrrfirat merged commit a136e81 into main Jul 29, 2026
62 checks passed
@serrrfirat
serrrfirat deleted the claude/close-the-epic-tails branch July 29, 2026 07:13
serrrfirat added a commit that referenced this pull request Jul 29, 2026
…g missing models (#6826)

* fix(llm): stop reading rate limits as auth failures, and stop retrying missing models

#6284 WS5's two live bugs. Both are permanent-vs-transient
misclassification, the same shape as #6824.

**1. A number containing 401 or 403 was read as an auth failure.**
`is_auth_error_message` matched `"401"`/`"403"` with a bare `contains`,
so any number holding those digits classified the failure as auth:

    "rate limited, retry after 4013 ms"   -> AuthFailed

A rate limit — the most common transient provider error there is — ended
the run immediately telling the user to fix their API key. `AuthFailed`
is deliberately neither retried nor circuit-broken, so the run had no
path back from a condition that would have cleared on its own.

Status codes now match on digit boundaries (`contains_status_code`).
Genuine `401`/`403` still classify; `4013`, `14030`, `1403`, `24019` and
`4010` no longer do.

**2. `LlmError::ModelNotAvailable` had zero producers.** Every 404 /
model-not-found fell into `RequestFailed`, which `retry::is_retryable`
treats as retryable, so a typo'd or decommissioned model id burned the
full 12-attempt budget before failing. Its consumers were all dead:
`retry::is_retryable`, `circuit_breaker::is_transient`, and the
gateway's `=> PolicyDenied` arm.

`is_model_not_available_message` now produces it. Deliberately
conservative: explicit phrases ("model not found", "does not exist",
"unknown model", ...) always match, but a bare `404` counts only when the
message also mentions the model — an unrelated 404 (a proxy path, a
health endpoint) still falls through to `RequestFailed`.

Ordering matters and is pinned by placement: context-length is checked
first (so a 413 is never read as auth), then auth (so a 403 naming a
model stays a permission problem), then missing-model.

Regression coverage, both red-verified against the old code:
  - `a_number_containing_401_is_not_an_auth_failure` fails with
    "rate limited, retry after 4013 ms" ... "the run would terminate
    telling the user to fix their API key".
  - `a_standalone_401_or_403_is_still_an_auth_failure` guards the fix
    from over-correcting.
  - `a_missing_model_is_not_retried_as_a_transient_failure` asserts the
    mapping AND that the result is not retryable — the second half is the
    part that fixes the burn.
  - `an_unrelated_404_still_falls_through` pins the conservative bound.

Gate: ironclaw_llm + ironclaw_runner — 1,671 tests pass, clippy clean,
fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(llm): restore the test attribute my insertion displaced

Clippy caught `duplicated attribute` at rig_adapter.rs. My test insertion
anchored on the `fn` line rather than its attribute, so the pre-existing
`#[test]` on `map_rig_error_unrelated_still_request_failed` ended up
stacked above my new doc comment, and that function lost its own.

Consequence worth naming: `map_rig_error_unrelated_still_request_failed`
STOPPED RUNNING. It compiled, the suite was green, and a test had quietly
been switched off. Only `-D warnings` on the full workspace caught it —
my local `cargo clippy -p ironclaw_llm` did not, because I scoped it to
the crate instead of running CI's actual command.

All five tests in that module now run and pass; clippy clean under
`cargo clippy --all --tests --examples --all-features -- -D warnings`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(llm): assert the auth boundary through map_rig_error, not the predicate

Review feedback on #6826: both boundary tests stopped at
`is_auth_error_message`, which is not what the run acts on. `map_rig_error`
turns the predicate into the error variant, and `retry::is_retryable` keys
off that variant -- so a predicate fix that failed to change the
classification would leave the bug exactly where it was.

Each case now also drives `map_rig_error` and asserts the contract:

- `4013 ms` and friends map to something retryable and NOT `AuthFailed`
  -- this is the rate limit the fix exists for, and it would have cleared
  on its own
- a standalone `401`/`403` maps to `AuthFailed` and is not retried,
  because retrying a bad credential cannot help

Sabotage-verified: reverting `contains_status_code` to the bare
`contains` fails on "rate limited, retry after 4013 ms".

Refs #6524

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…succeed (nearai#6824)

* fix(runner): stop silently retrying model-stage failures that cannot succeed

nearai#6284 WS1's remaining mapping box. Went in expecting a naming tidy-up;
it is a live retry-burn.

`model_stage_failure_category` returned `None` for `InvalidInvocation`,
`Invalid`, `ScopeMismatch` and `PolicyDenied`, so all four fell through
to `host_stage_unavailable_model` — which `is_auto_retriable_category`
lists among "transient host / lease / store / provider / tool faults
where re-running the *identical* request from the checkpoint is likely to
succeed".

None of the four can succeed on an identical retry. Policy does not
change between attempts, a malformed request stays malformed, and a
scope mismatch is configuration-shaped. So the run silently re-drove a
call that could never work, and reported "host stage unavailable" to the
operator instead of the real cause. That same list's doc comment says
"Conservative by design — anything not clearly transient falls to
`UserInitiated`"; the fallthrough defeated exactly that intent.

`executor/mapping.rs` documented this as already handled — "the runner
preserves the original kind when categorizing the failure." It did not.
Comment corrected in place rather than left to mislead the next reader.

Three categories, none auto-retriable (an unlisted category falls to
`UserInitiated`, which is the wanted behavior):

  - `model_stage_request_invalid`  (InvalidInvocation | Invalid)
  - `model_stage_policy_denied`    (PolicyDenied)
  - `model_stage_scope_mismatch`   (ScopeMismatch)

Two follow-on sites the compiler and the suite caught, both worth naming:

  - `failure_lane`'s canonical category list needed the three additions;
    its doc says "keep this in lockstep" and nothing enforces that but
    the test.
  - `every_failure_category_is_explainable_and_classified` rejected them
    for having no user-facing explanation. The generic fallback they
    would have used reads "Retry the run, and contact support if it keeps
    happening" — advice that can never work for a refused or malformed
    request. Each new category now states plainly that retrying will not
    help and names what has to change. That is the same defect this epic
    fixes for the model, one layer up: the *user* was also being told to
    retry something that could not succeed.

Two tests inverted, each with the reason recorded beside it:

  - `model_stage_host_error_kind_category_matrix_is_exhaustive` pinned
    all four kinds at `None` — the epic anticipated this ("a runner test
    pins the wrong behavior as expected — fix the mapping and invert the
    test").
  - `text_only_model_reply_driver_sanitizes_model_failures_...` pinned
    the generic `model_error` reason kind for a model-stage
    `PolicyDenied`.

New regression test `permanent_model_stage_failures_are_not_categorized_
as_transient_outages` asserts each kind gets a category AND that the
category is not auto-retriable — the second half is the part that
actually matters, and it fails against the old code.

Gate: runner, agent_loop, turns — 1,826 tests pass, clippy clean, fmt
clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(product): cover the new model-stage categories in the Tier-2 summary table

CI caught what my local gate missed: I ran the crates I edited
(`ironclaw_runner`, `ironclaw_agent_loop`, `ironclaw_turns`) rather than
the crates that CONSUME the value I changed.

`ironclaw_product` keeps a SECOND table of user-facing failure text, and
`failure_summary_covers_reborn_failure_category_constants` asserts it
covers exactly the public constants in `ironclaw_runner::failure_
categories`. Adding three constants without adding three entries fails
that test — correctly.

Added with the same wording used in the runner's table, so the two agree
for these categories.

Worth naming for whoever touches this next: the category CONSTANTS have a
single home (`ironclaw_product` imports them from `ironclaw_runner`), so
there is no vocabulary drift here — unlike the recovery-hint allowlist
fixed in nearai#6792. What is duplicated is the user-facing TEXT, across
`runner::failure_summary` and `product::projection`. That duplication is
guarded by this test, which is why it surfaced in minutes instead of
silently. Whether the two tables should be one is a real question, but a
separate one: they may be deliberately different registers (operator vs
end-user), and I have not verified that either way.

Gate: ironclaw_product + ironclaw_event_projections — 526 tests pass,
including the guard that failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(runner): cover every permanent model-stage kind through the driver

Review feedback on nearai#6824: the classifier had a direct unit test, but only
`PolicyDenied` was proven through the caller.

That distinction matters here rather than being a formality.
`map_host_error` reaches the category through an early return that
bypasses the kind match below it, so the classifier being correct does
not prove the driver emits what the classifier returns -- and the
emitted `reason_kind` is what `retry_disposition` keys on.

Covers all four permanent kinds -- `InvalidInvocation`, `Invalid`,
`ScopeMismatch`, `PolicyDenied` -- asserting each reaches the driver as
its own category and that `is_auto_retriable_category` rejects it. Before
the fix all four routed through `host_stage_unavailable_model`, which IS
auto-retriable, so a permanently-failing call was silently re-driven.

Sabotage-verified: removing the `InvalidInvocation | Invalid` arm from
the classifier now fails at the driver seam.

Refs nearai#6524

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…g missing models (nearai#6826)

* fix(llm): stop reading rate limits as auth failures, and stop retrying missing models

nearai#6284 WS5's two live bugs. Both are permanent-vs-transient
misclassification, the same shape as nearai#6824.

**1. A number containing 401 or 403 was read as an auth failure.**
`is_auth_error_message` matched `"401"`/`"403"` with a bare `contains`,
so any number holding those digits classified the failure as auth:

    "rate limited, retry after 4013 ms"   -> AuthFailed

A rate limit — the most common transient provider error there is — ended
the run immediately telling the user to fix their API key. `AuthFailed`
is deliberately neither retried nor circuit-broken, so the run had no
path back from a condition that would have cleared on its own.

Status codes now match on digit boundaries (`contains_status_code`).
Genuine `401`/`403` still classify; `4013`, `14030`, `1403`, `24019` and
`4010` no longer do.

**2. `LlmError::ModelNotAvailable` had zero producers.** Every 404 /
model-not-found fell into `RequestFailed`, which `retry::is_retryable`
treats as retryable, so a typo'd or decommissioned model id burned the
full 12-attempt budget before failing. Its consumers were all dead:
`retry::is_retryable`, `circuit_breaker::is_transient`, and the
gateway's `=> PolicyDenied` arm.

`is_model_not_available_message` now produces it. Deliberately
conservative: explicit phrases ("model not found", "does not exist",
"unknown model", ...) always match, but a bare `404` counts only when the
message also mentions the model — an unrelated 404 (a proxy path, a
health endpoint) still falls through to `RequestFailed`.

Ordering matters and is pinned by placement: context-length is checked
first (so a 413 is never read as auth), then auth (so a 403 naming a
model stays a permission problem), then missing-model.

Regression coverage, both red-verified against the old code:
  - `a_number_containing_401_is_not_an_auth_failure` fails with
    "rate limited, retry after 4013 ms" ... "the run would terminate
    telling the user to fix their API key".
  - `a_standalone_401_or_403_is_still_an_auth_failure` guards the fix
    from over-correcting.
  - `a_missing_model_is_not_retried_as_a_transient_failure` asserts the
    mapping AND that the result is not retryable — the second half is the
    part that fixes the burn.
  - `an_unrelated_404_still_falls_through` pins the conservative bound.

Gate: ironclaw_llm + ironclaw_runner — 1,671 tests pass, clippy clean,
fmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(llm): restore the test attribute my insertion displaced

Clippy caught `duplicated attribute` at rig_adapter.rs. My test insertion
anchored on the `fn` line rather than its attribute, so the pre-existing
`#[test]` on `map_rig_error_unrelated_still_request_failed` ended up
stacked above my new doc comment, and that function lost its own.

Consequence worth naming: `map_rig_error_unrelated_still_request_failed`
STOPPED RUNNING. It compiled, the suite was green, and a test had quietly
been switched off. Only `-D warnings` on the full workspace caught it —
my local `cargo clippy -p ironclaw_llm` did not, because I scoped it to
the crate instead of running CI's actual command.

All five tests in that module now run and pass; clippy clean under
`cargo clippy --all --tests --examples --all-features -- -D warnings`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(llm): assert the auth boundary through map_rig_error, not the predicate

Review feedback on nearai#6826: both boundary tests stopped at
`is_auth_error_message`, which is not what the run acts on. `map_rig_error`
turns the predicate into the error variant, and `retry::is_retryable` keys
off that variant -- so a predicate fix that failed to change the
classification would leave the bug exactly where it was.

Each case now also drives `map_rig_error` and asserts the contract:

- `4013 ms` and friends map to something retryable and NOT `AuthFailed`
  -- this is the rate limit the fix exists for, and it would have cleared
  on its own
- a standalone `401`/`403` maps to `AuthFailed` and is not retried,
  because retrying a bad credential cannot help

Sabotage-verified: reverting `contains_status_code` to the bare
`contains` fails on "rate limited, retry after 4013 ms".

Refs nearai#6524

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6824 — 2ecb43d2 Deployed Jul 28, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant