Skip to content

fix(kanban): dispatch fallback skips rungs whose credential is cooling - #1198

Merged
ang-fleet-lander[bot] merged 1 commit into
mainfrom
daedalus/t_6445986b-fallback-cooldown
Sep 26, 2026
Merged

ang-fleet-lander[bot] merged 1 commit into
mainfrom
daedalus/t_6445986b-fallback-cooldown

Conversation

@Kyzcreig

@Kyzcreig Kyzcreig commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Card t_6445986b.

Incident 2026-09-25: lane pool claude-bpr was budget-capped, so the dispatch fallback picked openai-codex while its credential was in cooldown. Every worker died at auth (worker_route_pin_refused rate_limited=true, exit 75), and each death stamped a rate_limited backoff, up to 45 min. The same thing was still happening live on the board on 09-26 (runs 11412-11419).

  • (1) A rate-limited worker_route_pin_refused event marks the provider as cooling for kanban.credential_cooldown_seconds (default 1800, 0 disables). A cooling rung is skipped the same way an open rate-limit circuit is. A card whose own route is cooling is held with reason credential_cooldown.
  • (2) If every rung is capped or cooling, the card is deferred: no spawn, no run, so no backoff stamp.
  • (3) Ordering is priority DESC, created_at ASC. That was already the case; a test now pins it for the freed pool slot.
  • (4) The deferred payload (fallback_skipped), the dispatch_provider_fallback event (skipped) and the route source string (dispatch-fallback(capped X; skipped openai-codex:credential_cooldown)) now name each skipped rung and the reason.

Tests: tests/hermes_cli/test_kanban_fallback_credential_cooldown.py. 8 pass on this branch; 4 fail on base.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

A worker_route_pin_refused event with rate_limited=true (e.g. 'Codex
credential is in cooldown.') marks that provider as cooling for
kanban.credential_cooldown_seconds (default 1800, 0 disables). A cooling
provider is never picked as a capped-pool fallback rung, and a card whose
own route is cooling is held (reason credential_cooldown). When every rung
is capped or cooling the card stays deferred: no spawn, no run, so no
rate_limited close and no backoff stamp. The deferred payload, the
dispatch_provider_fallback event and the route source name every skipped
rung and why.

openai-codex is not pool-bound, so the rate-limit circuit never covered it;
the cooling set is keyed by provider name and read run-scoped
(idx_events_run), not by scanning task_events.

Verified: new test_kanban_fallback_credential_cooldown.py 8 passed (4 fail
on base), plus lane_fallback_composition + pool_ratelimit_gates: 109 passed.

Card: t_6445986b
@ang-fleet-lander

Copy link
Copy Markdown

🤖 merged-by: apollo · lane: discord · gate: BYPASS: FR paused by Ace 2026-09-22; Argus off card review (Ace 09-24 13:08); gate = CI green + Apollo read · why: kanban dispatcher: dispatch-fallback skips a rung whose credential recorded a rate-limited pin refusal / rate-limit circuit within the cooldown; holds the card instead of spawning into a dead rung (cost t_c7a20f1b 4 spawns + 2h backoff). Tests 37/0 (t_6445986b)

@ang-fleet-lander
ang-fleet-lander Bot added this pull request to the merge queue Sep 26, 2026
Merged via the queue into main with commit 838bd59 Sep 26, 2026
56 checks passed
@ang-fleet-lander
ang-fleet-lander Bot deleted the daedalus/t_6445986b-fallback-cooldown branch September 26, 2026 05:41
@Kyzcreig Kyzcreig added the fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent) label Sep 26, 2026
@ang-prism

ang-prism Bot commented Sep 27, 2026

Copy link
Copy Markdown

FleetReview

Review: post-merge · head 838bd59948d0 · duration 11m 05s
Profile: full recipe · policy: below-size-and-path-gates
Roster: B-assert-ctx → gpt-6-sol (openai), B-state → gpt-6-sol (openai), C-assert-xhigh → claude-code-opus-5-5 (anthropic), F → gpt-6-sol (openai), G → grok-4.6 (xai), L6 → gpt-6-sol (openai)

Post-merge review (fleetreview:post-merge override): this reviewed the merge commit against its first parent — the bytes that already shipped. It is not a pre-merge gate pass.

Confidence: 3/5

Findings

  • P1 hermes_cli/kanban_db.py:14659 — One runtime 429 on a pinned pool provider holds every card on that provider for 30 min · agreed: B-assert-ctx,C-assert-xhigh,F (openai, anthropic)

FleetReview provenance · models: D=grok-4.6 · cost: $0.00 · duration: 15m 59s · rounds: 1 · files examined: 4

Kyzcreig pushed a commit that referenced this pull request Sep 28, 2026
…1032, #1043, #1095, #1198, #1254:121) (t_7da6cadf)

- #942: desktop hydration accepts a tool-call confab notice on an empty
  system row (mirrors notice_from_display_row); per-kind label.
- #970: switch_model(probe_catalog=False) passes allow_network through
  get_label / determine_api_mode; cold models.dev cache opens no socket.
- #976: _build_hyg_agent / hygiene get_session use the _hyg_old_sid snapshot.
- #1043: post-turn clear_resume_pending keeps a mark written during the turn.
- #1095: an early-imported bundled provider module is re-registered in the
  bundled discovery step (filesystem over pip precedence).
- #1198: cooling_providers ignores runtime-stage pin refusals.
- #1254:121: process_env_files overlay pre-values ride to child agent
  processes (HERMES_PROCESS_ENV_OVERLAY) so strip_overlay works there.
- #1032: store-level model resolution that changes the provider re-runs the
  base_url exfil guard at write time.

Verified: new tests fail on fork/main (8/8), pass on head; neighbouring
suites pass (see PR).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant