Repository navigation
ci: give the fix agent Opus 4.8 (1M), more turns, and inspection tools - #239
Conversation
The fix run on PR #238 died with error_max_turns: 80 turns, 42 permission denials, $3.82, and zero output — no refutations, no fixes, no summary. The agent thrashed on denied tool calls and never converged; the loop then paused safely (ai-loop-paused). Root cause is a task/allowance mismatch on the FIX agent specifically. Its job is heavy — verify each finding against code, edit, test, reply in threads, commit, push — but its allowlist carried no read-only inspection utilities, so every rg/cat/sed/ls/grep a verifying agent reaches for was a denial. Compounded by a tight 80-turn cap and the weaker sonnet model. Mirror robobun's approach (scope tools to the task; stronger model; no tight turn starvation) for this heavy pass: - Add read-only inspection utils: rg, grep, cat, sed, ls, head, tail, wc, find — kills the denial source, all read-only. - Raise the cap 80 -> 120 (the implement agent already runs at 150). - Run the fixer on claude-opus-4-8[1m] instead of sonnet-5 — a stronger model wastes fewer turns on a multi-finding adversarial pass, and the deep reviewer already uses opus. Implement agent is unchanged (it works — it opened #238). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NVR5yWevceZEFbTJmpviiM
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
WalkthroughThis change updates the ChangesClaude PR Loop Workflow Configuration
Estimated code review effort: 1 (Trivial) | ~3 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Preview DeploymentPreview URL: https://9bb1711e.bestax.pages.dev |
|
🎉 This PR is included in version 5.2.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
…nt (#241) Two hardening changes ahead of stress-testing the loop on a real bug: 1. Mirror the read-only inspection utilities (rg/grep/cat/sed/ls/head/tail/ wc/find) that #239 added to the fix agent into the implement agent's allowlist. The implement agent has the same latent gap that thrashed the sonnet fixer (42 denials); it has coped on sonnet-5 + 150 turns + native Read/Grep tools, but a heavy issue could hit the same wall. Cheap insurance, all read-only. 2. Upload the action's execution-output.json as an artifact (if: always()) from both the implement and fix jobs. That file carries the full turn-by-turn stream — every tool call AND every permission denial — whereas the Actions log only streams init + final result. When the sonnet fixer thrashed we could see '42 denials' but not WHICH tools; this makes the next failure diagnosable instead of guesswork. Reuses the repo's existing pinned actions/upload-artifact and never fails the job (if-no-files-found: ignore). Models are deliberately left as-is: sonnet-5 implement, opus deep review, opus fix — the heterogeneity decorrelates implementer and reviewer blind spots. Claude-Session: https://claude.ai/code/session_01NVR5yWevceZEFbTJmpviiM Co-authored-by: Claude <noreply@anthropic.com>
|
🎉 This PR is included in version 3.2.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
|
🎉 This PR is included in version 1.0.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
|
🎉 This PR is included in version 1.0.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
What
Three changes to the fix job in
claude-pr-loop.yml(only the agent that failed):rg,grep,cat,sed,ls,head,tail,wc,find.--max-turns80 → 120 (the implement agent already runs at 150).--modelclaude-sonnet-5→claude-opus-4-8[1m](Opus 4.8, 1M context).Why
The fix run on PR #238 died with
error_max_turns: 80 turns, 42 permission denials, $3.82, and zero output — no refutations, no fixes, no summary comment. The agent thrashed on denied tool calls and never reached convergence; the loop then paused safely (ai-loop-paused).Root cause is a task/allowance mismatch on the fix agent. Its job is heavy — verify each finding against code, edit, test, reply in threads, commit, push — but its allowlist carried no read-only inspection utilities, so every
rg/cat/sed/ls/grepa verifying agent naturally reaches for was a denial. Compounded by a tight 80-turn cap and the weaker sonnet model.This mirrors robobun's approach (scope tools to the task, use a stronger model, don't starve it on turns) for this heavy adversarial pass. The implement agent is unchanged — it works (it opened #238).
Scope
One
claude_argsblock; three values changed. No prompt or logic changes.Caveat
claude-opus-4-8[1m]requests the 1M-context beta. If it's not available on the subscription token path, the next fix run will surface it and we drop the[1m]suffix — the Opus model itself is the main win.Validation
After merge: on PR #238 remove
ai-loop-pausedand re-addai-loopto retry the fix pass on the same commit, and confirm the agent now posts refutations instead of thrashing.Generated by Claude Code
Summary by CodeRabbit