Repository navigation
CodeRouter: hold capacity errors on the same model instead of failing fast - #15310
Conversation
|
Warning Review limit reachedNext included review available in 2 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: manaflow-ai/cmux/.coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (17)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
Adds tests for issue #15301: a capacity storm (529, 429, transport, overloaded SSE events) shorter than the hold budget must reach the client as a success on the same model, with the wait recorded in telemetry. These fail until the proxies hold instead of failing fast. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…iling fast A capacity blip ended autonomous agent turns: after four accounts, or when every account was cooling, the proxies answered 503 and the agent CLI stopped. The Claude and Codex proxies now hold the request when a routing round ends on a transient failure (capacity, 429, 5xx/529, transport), wait with jittered exponential backoff that honors the soonest account cooldown, and replay the same body and model. The hold lasts up to CODEROUTER_CAPACITY_HOLD_MS (20 minutes), bounded by the header budget, never starts after output reached the client, and ends at once when no account recovers in time (revoked credential, later quota reset). route_events gains held_ms and hold_count (ClickHouse migration 006), and the request outcome carries them to PostHog traces. Refs #15301 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
f5d72f5 to
b95d040
Compare
CI failure attributionCI passes on Written by |
|
Merge receipt for |
4e0f7d2 fix(bash): keep $? for PROMPT_COMMAND hooks after cmux's (manaflow-ai#15255) ae49bf5 fix(examples): show custom description in Project Worktrees sidebar (manaflow-ai#15256) a9a229d Add cross-provider token usage accounting for agent transcripts (manaflow-ai#15332) 860619f Add a .worktreeinclude reader for seeding new worktrees (manaflow-ai#15413) 3edbd83 Clear the stale Needs input badge when Claude's permission is decided in the terminal (manaflow-ai#15170) 9ed9294 CodeRouter: hold capacity errors on the same model instead of failing fast (manaflow-ai#15310) 56d4547 docs: add a front door for outside contributors (manaflow-ai#15263) 799f906 fix(ci): recognize GUI token acquisition failures (manaflow-ai#15449) f118d43 ci: age parked builds by measured reuse distance (manaflow-ai#15616) 1f6744d ci: harden overflow switch recovery (manaflow-ai#15617) 9987778 Predicted echo: remote terminals only, withdraw on pasted and sent input (manaflow-ai#15211) d9e199b Subtle selection follow-ups: group header hairline, no focus re-render for legacy rows, cmux.json test (manaflow-ai#15195) c13afe1 test: cover UTF-8 workspace create commands (manaflow-ai#15622) e76a660 fix: preserve Claude remote-control names on restore (manaflow-ai#15619) 900f248 feat: expose cmux-owned scratch metadata in session listing (manaflow-ai#15615) b5604fa ci: say why compiled-product reuse refused an artifact (manaflow-ai#15553) # Conflicts: # .github/workflows/ci-cloud-overflow-probe.yml
A capacity blip ended autonomous agent turns. When every Claude or Codex account was cooling, or four accounts had failed, CodeRouter answered 503
overloaded_error. Codex and Claude Code show that as a final error and stop until a human types "continue" (#15301).CodeRouter now holds the request and retries the same model instead of failing fast:
retry-after.CODEROUTER_CAPACITY_HOLD_MS(default 20 minutes). It is capped by the existing 25-minute header budget, so the function still answers before its 30-minutemaxDuration.invalid_credentialcooldowns when computing the recovery time; Codex gets a smallnextCapacityAvailableAtquery for the same purpose.route_eventsgainsheld_msandhold_count, and the PostHog trace carriescoderouter_held_ms/coderouter_hold_count, so capacity pain shows up as latency.Deploy order: ClickHouse migration
web/db/clickhouse/006_route_events_capacity_hold.sql(additiveADD COLUMN IF NOT EXISTS ... DEFAULT 0) must be applied tocoderouter_devandcoderouterbefore merge. Otherwise route-event inserts carry unknown columns.Part 2 of #15301 (agents auto-resume after a retryable failure) is a separate PR.
Validation
bun test tests/coderouter-*.test.ts*: 506 pass, 0 fail (rebased on main). CI: commit 1 red (web typecheck and web tests), commit 2 green (59/59). New tests: a 529 → 429 → transport → 529 → 200 storm on one Claude account returns 200 with four holds and the same model on every replay; the same for Codex, with each wait honoring the recorded cooldown; the budget cap; no hold for a revoked credential or an hour-long quota; no replay after output started.tsc --noEmitclean;lint:complexityclean.next build+next start; this machine cannot reach cmux's Vercel previews). A throwaway, unpushed branch swaps in stub auth, one stub account with in-memory cooldowns, and a mock upstream, then drives the real/v1/messagesand/v1/responsesroutes. Every storm shorter than the budget reached the client as a 200 on the same model; the same build withCODEROUTER_CAPACITY_HOLD_MS=0returns the 529 in 2.3 s, which is the old behavior.Transcript
🤖 Generated with Claude Code