Skip to content

fix(query): warn before repeated tool failures stop - #1927

Merged
kevincodex1 merged 10 commits into
Twigpine:mainfrom
jatmn:fix/tool-failure-advisory-1926
Jul 12, 2026
Merged

kevincodex1 merged 10 commits into
Twigpine:mainfrom
jatmn:fix/tool-failure-advisory-1926

Conversation

@jatmn

@jatmn jatmn commented Jul 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Warn the model when each persistent tool-name/error-category failure signature reaches the penultimate configured threshold.
  • Deliver each warning only after the next API turn has passed continuation safety checks, while preserving the existing hard stop at the threshold.
  • Do not emit a warning when no next tool-capable/API turn can consume it.

Fixes #1926

Root cause

The tool-failure loop guard only returned hard-stop decisions. Models could repeatedly retry the same failing tool without being told that the next matching failure would terminate the query.

Validation

  • bun test src/query/toolFailureLoopGuard.test.ts
  • bun run build
  • bun run smoke

Limitations

The advisory is intentionally limited to persistent signature failures.

Final reviewed SHA: 5266ff8aa7f268ae6e8fda17d83e5689befb4ace

Summary by CodeRabbit

  • New Features

    • Added advance warnings when repeated tool failures are approaching the configured limit.
    • Warnings are shared with the next model step without interrupting the current run.
    • Warning details are sanitized to avoid exposing unsafe tool names or error text.
  • Bug Fixes

    • Prevented duplicate warnings and ensured warnings remain available after message compaction.
    • Improved handling of simultaneous and mixed tool-failure patterns.

@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 7 minutes

Your organization has reached its usage spending cap. Adjust your spending cap in the billing tab.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 40f752a3-4ad6-4253-842d-620951cceba2

📥 Commits

Reviewing files that changed from the base of the PR and between cb73acc and 9457a51.

📒 Files selected for processing (3)
  • src/query.ts
  • src/query/toolFailureLoopGuard.test.ts
  • src/query/toolFailureLoopGuard.ts
📝 Walkthrough

Walkthrough

The tool-failure loop guard now emits a near-threshold advisory, and queryLoop forwards it as a meta message before the next model turn. Tests cover persistent signatures, sanitization, threshold edge cases, compaction, logging, and forwarding conditions.

Changes

Tool failure advisory

Layer / File(s) Summary
Guard advisory contract and emission
src/query/toolFailureLoopGuard.ts, src/query/toolFailureLoopGuard.test.ts
The guard returns advisory metadata when a persistent failure reaches threshold - 1, sanitizes advisory content, preserves terminal trip behavior, and tests repeated, mixed, concurrent, and low-threshold cases.
Query-loop advisory forwarding
src/query.ts, src/query/toolFailureLoopGuard.test.ts
queryLoop defers advisories as meta user messages, preserves them through autocompaction with UUID de-duplication, logs their metadata, and forwards them only when another model turn is available.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Suggested labels: bug

Suggested reviewers: kevincodex1

🚥 Pre-merge checks | ✅ 5 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
No Hidden Policy Change ⚠️ Warning FAIL: the PR hides policy-sensitive changes in unrelated files—new NVIDIA NIM default routing, permission path resolution, and unknown-model cost policy—beyond the advisory feature. Split these into a separate PR or explicitly surface them and get maintainer signoff for routing-default, permission-policy, and cost/telemetry behavior changes.
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title is concise and accurately describes the main change to warn before repeated tool failures stop.
Description check ✅ Passed The description includes summary, validation, and limitations, and only omits the template's explicit Impact and Notes sections.
Linked Issues check ✅ Passed The change adds a near-threshold advisory, preserves the hard stop, and forwards it to the next model turn as requested.
Out of Scope Changes check ✅ Passed The diff stays focused on the guard logic, query-loop wiring, and tests, with no clear unrelated additions.
Risk Surface Disclosed ✅ Passed Only query-loop/tool-failure advisory plumbing changed; none of the flagged surfaces are touched, so no risk-surface blocker callout was required.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/query/toolFailureLoopGuard.ts (1)

158-197: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Preserve the advisory when a batch also contains success
hasSuccess returns early with { tripped: false }, so a near-threshold failure can drop its warning whenever the same batch includes an unrelated success. The persistent count stays at threshold - 1, but the model never sees the “one more failure will stop the query” advisory. Return the advisory from this branch too.

🐛 Proposed fix
   if (hasSuccess) {
     resetToolFailureLoopGuard(params.state, successfulMutationPaths)
-    return { tripped: false }
+    return advisory ? { tripped: false, advisory } : { tripped: false }
   }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/query/toolFailureLoopGuard.ts` around lines 158 - 197, Preserve any
pending advisory when the batch contains a successful mutation: update the
hasSuccess branch in the tool failure loop guard to reset state and return `{
tripped: false, advisory }` rather than discarding the advisory. Use the
existing advisory variable and resetToolFailureLoopGuard logic without changing
path-trip behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/query/toolFailureLoopGuard.test.ts`:
- Around line 94-153: Add a test alongside the existing tool failure loop guard
cases covering a single update containing one successful tool result and a
different tool’s matching failure that reaches threshold minus one. Assert the
decision is not tripped and includes the expected advisory for the failing tool,
including its tool name, error category, and remaining-failure message; use
update and the existing toolUse/toolResult helpers.
- Around line 854-871: Replace the source-text position assertions in the test
“query loop forwards an advisory to the next model turn” with a behavioral test
that invokes the query loop using mocked tool-failure/advisory decisions, then
verifies the advisory is forwarded as a meta message in the next model turn or
pushed tool result. Use the query generator’s existing test seams and mocks, and
ensure the scenario covers the success-handling path so advisory messages are
not dropped.

---

Outside diff comments:
In `@src/query/toolFailureLoopGuard.ts`:
- Around line 158-197: Preserve any pending advisory when the batch contains a
successful mutation: update the hasSuccess branch in the tool failure loop guard
to reset state and return `{ tripped: false, advisory }` rather than discarding
the advisory. Use the existing advisory variable and resetToolFailureLoopGuard
logic without changing path-trip behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 10065724-7312-4d36-8971-4ae9724b5f60

📥 Commits

Reviewing files that changed from the base of the PR and between 64d164d and ce95e4a.

📒 Files selected for processing (3)
  • src/query.ts
  • src/query/toolFailureLoopGuard.test.ts
  • src/query/toolFailureLoopGuard.ts
📜 Review details
⏰ Context from checks skipped due to timeout. (4)
  • GitHub Check: CodeRabbit / Review
  • GitHub Check: smoke-and-tests (22)
  • GitHub Check: smoke-and-tests (24.11.x)
  • GitHub Check: typecheck
🧰 Additional context used
📓 Path-based instructions (4)
**/*.{ts,tsx}

📄 CodeRabbit inference engine (AGENTS.md)

TypeScript code in this repository must use strict mode and ESM imports.

Files:

  • src/query.ts
  • src/query/toolFailureLoopGuard.test.ts
  • src/query/toolFailureLoopGuard.ts
**

⚙️ CodeRabbit configuration file

**: # AGENTS.md - AI Agent Coding Guide

This guide is for AI coding agents working in the OpenClaude repository. Read it before changing code, and also follow CONTRIBUTING.md for contributor policy, PR expectations, review follow-up, and project scope.

Project Snapshot

OpenClaude is a coding-agent CLI for cloud and local model providers. It supports OpenAI-compatible APIs, Anthropic, Gemini, DeepSeek, Ollama, MCP, local backends, slash commands, tools, agents, and a React/Ink terminal UI.

The installed CLI runs on Node.js >=22.0.0. Bun is used for source builds, scripts, dependency management, and tests.

Work Style

  • Keep changes focused on one problem.
  • Prefer existing patterns in the file or nearby module.
  • Avoid unrelated formatting, renames, dependency changes, or broad rewrites.
  • Add or update tests when behavior changes.
  • Update docs when setup, commands, provider behavior, or user-facing behavior changes.
  • For new features, larger refactors, dependencies, or runtime changes, follow the issue-first guidance in CONTRIBUTING.md.

Stack And Conventions

  • TypeScript with strict mode and ESM imports.
  • React + Ink for terminal UI.
  • Bun lockfile and Bun scripts for development workflows.
  • Node runtime for the built CLI.

Common libraries and patterns:

  • chalk for terminal color.
  • commander for CLI argument parsing.
  • execa for child processes.
  • Existing service, provider, settings, permission, and UI patterns over new abstractions.

Repository Map

  • src/commands/ - slash and CLI command implementations.
  • src/components/ - React/Ink UI components.
  • src/services/ - API, MCP, OAuth, wiki, voice, and other service integrations.
  • src/tools/ - tool implementations.
  • src/utils/ - shared utilities.
  • src/integrations/ - provider and model integration metadata.
  • src/entrypoints/ - CLI, MCP, SDK, and generated public types.
  • src/tasks/ - local, remote, workflow, and monitor tas...

Files:

  • src/query.ts
  • src/query/toolFailureLoopGuard.test.ts
  • src/query/toolFailureLoopGuard.ts
**/*

⚙️ CodeRabbit configuration file

**/*: Apply the OpenClaude maintainer review rubric from AGENTS.md. Review the current diff, not stale discussion context. Separate real blockers from suggestions. Do not request changes for vague style churn. Treat approval as merge-ready from CodeRabbit's side, pending required human review and GitHub Checks. If checks are failing or unavailable, say so clearly instead of implying the PR is fully ready.

Files:

  • src/query.ts
  • src/query/toolFailureLoopGuard.test.ts
  • src/query/toolFailureLoopGuard.ts
{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}

⚙️ CodeRabbit configuration file

{src/**/*.test.ts,src/**/*.test.tsx,tests/**,scripts/**/*.test.ts,vscode-extension/**/*.test.js}: Review tests for meaningful coverage of the changed behavior, isolation of global/env/config state, async cleanup, fake timers, provider profile leaks, and Windows-compatible assumptions. Block when risky runtime changes lack focused regression coverage or tests assert implementation details while missing the user-visible behavior.

Files:

  • src/query/toolFailureLoopGuard.test.ts
🔇 Additional comments (2)
src/query/toolFailureLoopGuard.ts (1)

29-41: LGTM!

Also applies to: 136-136, 239-239, 478-494

src/query.ts (1)

2806-2825: LGTM! Forwarding logic (meta message construction, yield, toolResults append, ordering before the recursive call) matches the guard's advisory contract.

One minor drive-by: hasToolName/hasErrorCategory in the debug log and logEvent payload are always true since advisory.toolName/advisory.errorCategory are required strings — logging the actual toolName/errorCategory values would be more useful than a constant flag, though this is inconsequential since logEvent is currently a no-op.

Comment thread src/query/toolFailureLoopGuard.test.ts
Comment thread src/query/toolFailureLoopGuard.test.ts
coderabbitai[bot]
coderabbitai Bot previously approved these changes Jul 10, 2026
@jatmn jatmn self-assigned this Jul 10, 2026
@jatmn jatmn added the enhancement New feature or request label Jul 10, 2026
@jatmn
jatmn marked this pull request as ready for review July 10, 2026 05:11

@jatmn jatmn left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@0xghost42 0xghost42 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice addition — a heads-up turn before the hard stop gives the model a real chance to change tack, and gating it on threshold > 1 && persistentSignatureCount === threshold - 1 makes it a clean single-shot warning at the penultimate failure rather than a repeated nag. The message is specific and actionable (names the tool, the error category, and the exact count), which is the right call.

One thing I'd verify: the "don't warn when no next turn is possible" case from the summary. The advisory is generated purely from the signature count in evaluateToolFailureLoop, and in query.ts it's yielded and pushed to toolResults unconditionally in the for (const advisoryDecision of ... advisories ?? []) loop. I don't see the guard for "this is already the last permitted turn (maxTurns / maxTokens reached)" in this diff — if that suppression lives upstream of this loop, great; if not, there's a path where the model gets "one more matching failure will stop the query" on a turn where it can't actually act, which is a slightly confusing dead-end. Worth a test that asserts no advisory is emitted on the final allowed turn.

Minor: the debug log computes hasToolName=${advisoryDecision.toolName !== undefined} while the logEvent right below hardcodes hasToolName: true, hasErrorCategory: true. Since both fields are required on the advisory type they're always defined, so true is accurate — but the two log sites disagree on whether the value is dynamic, which will read as a bug to the next person. I'd make them consistent (drop the !== undefined in the debug line, or compute both).

@kevincodex1

Copy link
Copy Markdown
Member

hello @jatmn please address some comments from @0xghost42

@jatmn
jatmn force-pushed the fix/tool-failure-advisory-1926 branch from af0bf85 to 5266ff8 Compare July 12, 2026 02:15
coderabbitai[bot]
coderabbitai Bot previously approved these changes Jul 12, 2026
@kevincodex1
kevincodex1 merged commit 2448ea9 into Twigpine:main Jul 12, 2026
5 checks passed
@jatmn
jatmn deleted the fix/tool-failure-advisory-1926 branch July 12, 2026 15:18
hotmanxp pushed a commit to hotmanxp/openclaude that referenced this pull request Jul 15, 2026
* fix(query): warn before repeated tool failures stop

* fix(query): preserve tool failure advisories

* fix(query): forward all tool failure advisories

* fix(query): harden tool failure advisories

* fix(query): preserve tool failure advisories

* fix(query): keep advisories one-shot

* fix(query): avoid duplicate advisories after compaction

* fix(query): compare advisory message IDs

* fix(query): cover advisory forwarding edges

* fix(query): retain advisories without tools
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Warn the model before "Stopped: repeated tool failures detected"

3 participants