Repository navigation
fix: preserve non-transient cooldown reasons - #15885
Conversation
Exercise non-transient failure precedence across Claude and native account cooldown stores, including capacity consumers. Co-Authored-By: Codex <noreply@openai.com>
Share cooldown deadline and failure-code precedence across Claude and native account stores. Co-Authored-By: Codex <noreply@openai.com>
|
Warning Review limit reachedNext included review available in 6 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
The comment explaining why the cooldown write is one SQL statement was lost when the two call sites moved to a shared helper. It records why a late provider error must not shorten a longer cooldown another request already stored, which is the reason this cannot be a read-modify-write. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ReviewReviewed at Blocking: the new precedence has no time bound, and the Claude table can never clear the flagThe rule "a stored non-transient reason outranks a new transient reason" is right while the account is still cooling for that reason. It is wrong once that cooldown has expired, and nothing bounds it. On The sequence that breaks, with every link checked in the code rather than assumed:
Before this PR the later transient write overwrote the stale reason, so step 4 held correctly. This is a regression on the capacity-hold path the change exists to serve. The fix is to bound the precedence to a live cooldown: a stored non-transient reason outranks a new transient one only while Two smaller items going in with it:
A test for the regressing sequence is going in first, red, before the fix. What I verified, and found correct
AlsoThe rationale comment explaining why this write is a single statement was lost when both call sites moved to the helper. Restored in Filed as a consequence, not a blocker: #15876. Even with the time bound, the Claude pool still has no operator action that clears a cooldown and its reason after a credential is repaired, and the dashboard renders |
Pin replacement of stale non-transient reasons and NULL-reason cooldown handling for both account stores. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Let fresh transient failures replace expired credential reasons and harden the shared cooldown tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |
Review resolvedNew head Fixed
Verified, by me, at the new head
Left
Enabling auto-merge. This is a fix, so it does not go to the team design tracker. |
|
Merge receipt for |
5eda931 fix(ios): keep auth operations from missing token store (manaflow-ai#14302) 4d9bec3 fix: restore per-label Blacksmith macOS capacity 572beb6 fix: preserve non-transient cooldown reasons (manaflow-ai#15885) 747aa96 Prevent stale Cloud agent-chat reconnects (manaflow-ai#15920) 33ad0b1 Hibernate only agents whose wake can relaunch the original launcher (manaflow-ai#15287)
Summary
PR #15856 fixed Claude cooldown deadlines with
GREATEST, but regressed the failure reason: a shorterinvalid_credentialcooldown could no longer replace a longerrate_limitedreason. This PR fixes that regression and the pre-existing native asymmetry in one shared SQL helper.Both account tables now use the same rule in one
UPDATEstatement:GREATEST(COALESCE(existing, new), new).invalid_credentialoutranks a transient reason, even when its deadline is shorter.updated_atis always bumped.rate_limited, then 401 15minvalid_credentialrate_limited(wrong). Native: 5h,invalid_credentialinvalid_credentialinvalid_credential, then 503 20supstream_unavailableinvalid_credential. Native: 15m,upstream_unavailable(wrong)invalid_credentialrate_limited, then 503 20supstream_unavailablerate_limited. Native: 5h,upstream_unavailable(wrong)rate_limitedupstream_unavailable, then 429 5hrate_limitedrate_limitedrate_limitedThe reason is load-bearing, not just display metadata. Claude's
capacityRetryAfterexcludesinvalid_credentialaccounts so the proxy does not hold a request for a credential that needs human intervention. NativenextCapacityAvailableAtapplies the same exclusion when deciding whether capacity is coming back. The sharedcooldownWrite.tsmodule now owns the taxonomy and precedence rule, and the native query uses the same predicate instead of a separate hardcoded string.Tests
The focused command was red before the fix at commit
b6bad0dd81e:The failures included Claude losing
invalid_credentialafter a longerrate_limitedwall and native overwriting a surviving non-transient reason withupstream_unavailable.The same command is green at commit
11c369f2b0b:Additional validation:
bun x tsc --noEmitpassed.bun run lint:complexitypassed with 42 findings matched to the unchanged baseline.web / web-db-migrationsjob throughbun run test:db:behavior.web/scripts/run-db-behavior-tests.shdiscovers every DB-gated test file, runs them serially, and fails if a file runs zero tests or skips one. That job passed on fix: preserve longest Claude upstream cooldown #15856, so the merged test ran in CI. The local real-Postgres run is the faster loop for this follow-up.There is no fleet dogfood because PR #8029 disabled Vercel branch previews, so an unmerged web change has no preview URL.
Changelog
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Summary by cubic
Fixes the cooldown failure-reason regression where a shorter
invalid_credentialcooldown could no longer replace a longerrate_limitedreason. Claude and native account stores now share one SQL helper (cooldownWrite) that keeps the longest deadline and gives non-transient reasons precedence over transient ones while their deadline is live; once expired, a fresh transient failure replaces them. It also fixes the native asymmetry whereupstream_unavailablecould overwrite a surviving non-transient reason. The stored reason now always describes the live cooldown, so capacity consumers (Claude'scapacityRetryAfterand nativenextCapacityAvailableAt) excludeinvalid_credentialcorrectly, and a NULL reason on a live deadline stays untouched. Tests cover all sequence orders, expired-reason recovery, NULL-reason handling, and the capacity paths.Written for commit b20de6b. Summary will update on new commits.