fix(kanban): dispatch fallback skips rungs whose credential is cooling - #1198
Conversation
A worker_route_pin_refused event with rate_limited=true (e.g. 'Codex credential is in cooldown.') marks that provider as cooling for kanban.credential_cooldown_seconds (default 1800, 0 disables). A cooling provider is never picked as a capped-pool fallback rung, and a card whose own route is cooling is held (reason credential_cooldown). When every rung is capped or cooling the card stays deferred: no spawn, no run, so no rate_limited close and no backoff stamp. The deferred payload, the dispatch_provider_fallback event and the route source name every skipped rung and why. openai-codex is not pool-bound, so the rate-limit circuit never covered it; the cooling set is keyed by provider name and read run-scoped (idx_events_run), not by scanning task_events. Verified: new test_kanban_fallback_credential_cooldown.py 8 passed (4 fail on base), plus lane_fallback_composition + pool_ratelimit_gates: 109 passed. Card: t_6445986b
|
🤖 merged-by: apollo · lane: discord · gate: BYPASS: FR paused by Ace 2026-09-22; Argus off card review (Ace 09-24 13:08); gate = CI green + Apollo read · why: kanban dispatcher: dispatch-fallback skips a rung whose credential recorded a rate-limited pin refusal / rate-limit circuit within the cooldown; holds the card instead of spawning into a dead rung (cost t_c7a20f1b 4 spawns + 2h backoff). Tests 37/0 (t_6445986b) |
FleetReviewReview: post-merge · head Post-merge review ( Confidence: 3/5 Findings
FleetReview provenance · models: D=grok-4.6 · cost: $0.00 · duration: 15m 59s · rounds: 1 · files examined: 4 |
…1032, #1043, #1095, #1198, #1254:121) (t_7da6cadf) - #942: desktop hydration accepts a tool-call confab notice on an empty system row (mirrors notice_from_display_row); per-kind label. - #970: switch_model(probe_catalog=False) passes allow_network through get_label / determine_api_mode; cold models.dev cache opens no socket. - #976: _build_hyg_agent / hygiene get_session use the _hyg_old_sid snapshot. - #1043: post-turn clear_resume_pending keeps a mark written during the turn. - #1095: an early-imported bundled provider module is re-registered in the bundled discovery step (filesystem over pip precedence). - #1198: cooling_providers ignores runtime-stage pin refusals. - #1254:121: process_env_files overlay pre-values ride to child agent processes (HERMES_PROCESS_ENV_OVERLAY) so strip_overlay works there. - #1032: store-level model resolution that changes the provider re-runs the base_url exfil guard at write time. Verified: new tests fail on fork/main (8/8), pass on head; neighbouring suites pass (see PR).
Card t_6445986b.
Incident 2026-09-25: lane pool claude-bpr was budget-capped, so the dispatch fallback picked openai-codex while its credential was in cooldown. Every worker died at auth (worker_route_pin_refused rate_limited=true, exit 75), and each death stamped a rate_limited backoff, up to 45 min. The same thing was still happening live on the board on 09-26 (runs 11412-11419).
dispatch-fallback(capped X; skipped openai-codex:credential_cooldown)) now name each skipped rung and the reason.Tests: tests/hermes_cli/test_kanban_fallback_credential_cooldown.py. 8 pass on this branch; 4 fail on base.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.