Skip to content

fix(combo): bound the pre-dispatch unavailable skip so a stale label cannot dark a pool (#12168) - #12285

Merged
diegosouzapw merged 1 commit into
release/v3.8.51from
fix/12168-predispatch-lazy-cooldown
Sep 1, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.51from
fix/12168-predispatch-lazy-cooldown

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

What

getPersistedConnectionCooldownSkipReason() skipped any connection whose testStatus was unavailable, with no elapsed-cooldown check:

if (status === "unavailable") return `Skipping ...`;

This is the raw-label anti-pattern AGENTS.md explicitly warns about ("check whether code is reading raw state instead of using getStatus()/canExecute()") — the resilience layers are designed to recover lazily. The sibling helper immediately above it, getConnectionStatusQuotaCutoffReason(), gets this right: it requires hasFutureRateLimitUntil() before treating unavailable as blocking.

Why the original justification does not hold

The in-code comment argued:

Lazy recovery is unaffected: clearAccountError() resets the status on first success.

But this gate runs before dispatch — so it blocks the very successful request that would call clearAccountError(). Chicken-and-egg.

And the out-of-band recovery job cannot rescue it either: hasElapsedCooldown() in src/lib/quota/connectionRecovery.ts requires a rateLimitedUntil to be present to consider a cooldown elapsed. A row with a stale unavailable label and no timestamp is unreachable by both paths.

Net effect, exactly as reported in #12168: a whole combo pool answering ALL_TARGETS_SKIPPED with recordedAttempts === 0 — zero upstream attempts made, and no route back to healthy.

Regression introduced by #11360, shipped in v3.8.50.

How

The original intent is preserved — do not burst into a connection AUTH just retired, before its timestamp lands — but bounded: the bare label is honoured only while lastErrorAt is inside a grace window, mirroring ERROR_LABEL_GRACE_MS in connectionRecovery.ts so the two never disagree about whether a label is still meaningful. Past that window the request goes through, and one real attempt either succeeds (clearing the status) or re-arms the cooldown with a fresh timestamp.

A missing/unparseable lastErrorAt is treated as stale, not blocking — an unbounded skip is precisely the reported failure, and one extra upstream attempt is far cheaper than a permanently dark pool.

Validation (TDD, Hard Rule #18)

Two assertions in repro-combo-persisted-cooldown-preskip.test.ts encoded the buggy behavior as intended ("skips an unavailable connection whose cooldown already expired"). They are realigned to the corrected contract — not weakened: the recent-failure case still asserts the skip fires, it just now requires lastErrorAt to make the label meaningful.

  • RED before the fix: the two new #12168 cases fail against origin/release/v3.8.51 (verified by restoring the base file and re-running) — 12 pass / 2 fail
  • GREEN after: 14/14
  • Sibling suites (combo-cooldown-retry, combo-provider-cooldown, combo-provider-cooldown-sibling, chatcore-compression-combo-predicates): 40/40
  • typecheck:core clean · eslint (with project suppressions) clean

Scope note

Refs #12168 rather than Closes. The validation also flagged a plausible but unproven second path that can produce the same orphan state: markAccountUnavailable()'s zero-cooldown branch (src/sse/services/auth.ts) nulls rateLimitedUntil while updateProviderConnection merges, leaving a pre-existing unavailable label in place. This PR makes such a row recoverable, but hardening that write so the state is never produced deserves its own change with its own repro.

Refs #12168

…cannot dark a pool (#12168)

getPersistedConnectionCooldownSkipReason() returned a skip for ANY connection
whose testStatus was `unavailable`, with no elapsed-cooldown check:

    if (status === "unavailable") return `Skipping ...`;

That is the raw-label anti-pattern AGENTS.md warns about ("check whether code
is reading raw state instead of using getStatus()/canExecute()") — the
resilience layers are meant to recover lazily. The sibling helper directly
above it, getConnectionStatusQuotaCutoffReason(), does require
hasFutureRateLimitUntil() before treating `unavailable` as blocking.

Its stated justification — "Lazy recovery is unaffected: clearAccountError()
resets the status on first success" — does not hold on this path. This gate
runs BEFORE dispatch, so it prevents the very successful request that would
call clearAccountError(). And a row whose rateLimitedUntil is absent cannot be
rescued by the out-of-band recovery job either, because hasElapsedCooldown()
there requires a timestamp to be present.

Net effect reported in #12168: an entire combo pool answering
ALL_TARGETS_SKIPPED with recordedAttempts === 0 — zero upstream attempts, no
path back to healthy.

The original intent (do not burst into a connection AUTH just retired, before
the timestamp lands) is preserved, but bounded: the bare label is honoured only
while lastErrorAt is inside a grace window, mirroring ERROR_LABEL_GRACE_MS in
src/lib/quota/connectionRecovery.ts so the two never disagree about whether a
label is still meaningful. Past the window the request goes through, and one
real attempt either succeeds (clearing the status) or re-arms the cooldown with
a fresh timestamp.

Regression introduced by #11360, shipped in v3.8.50.

Two assertions in repro-combo-persisted-cooldown-preskip.test.ts encoded the
buggy behavior as intended ("skips an unavailable connection whose cooldown
already expired") and are realigned to the corrected contract, plus a case for
the orphan state (unavailable with no timestamps at all).
@diegosouzapw
diegosouzapw merged commit eeba382 into release/v3.8.51 Sep 1, 2026
20 of 21 checks passed
@diegosouzapw
diegosouzapw deleted the fix/12168-predispatch-lazy-cooldown branch September 1, 2026 12:22
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…cannot dark a pool (diegosouzapw#12168) (diegosouzapw#12285)

getPersistedConnectionCooldownSkipReason() returned a skip for ANY connection
whose testStatus was `unavailable`, with no elapsed-cooldown check:

    if (status === "unavailable") return `Skipping ...`;

That is the raw-label anti-pattern AGENTS.md warns about ("check whether code
is reading raw state instead of using getStatus()/canExecute()") — the
resilience layers are meant to recover lazily. The sibling helper directly
above it, getConnectionStatusQuotaCutoffReason(), does require
hasFutureRateLimitUntil() before treating `unavailable` as blocking.

Its stated justification — "Lazy recovery is unaffected: clearAccountError()
resets the status on first success" — does not hold on this path. This gate
runs BEFORE dispatch, so it prevents the very successful request that would
call clearAccountError(). And a row whose rateLimitedUntil is absent cannot be
rescued by the out-of-band recovery job either, because hasElapsedCooldown()
there requires a timestamp to be present.

Net effect reported in diegosouzapw#12168: an entire combo pool answering
ALL_TARGETS_SKIPPED with recordedAttempts === 0 — zero upstream attempts, no
path back to healthy.

The original intent (do not burst into a connection AUTH just retired, before
the timestamp lands) is preserved, but bounded: the bare label is honoured only
while lastErrorAt is inside a grace window, mirroring ERROR_LABEL_GRACE_MS in
src/lib/quota/connectionRecovery.ts so the two never disagree about whether a
label is still meaningful. Past the window the request goes through, and one
real attempt either succeeds (clearing the status) or re-arms the cooldown with
a fresh timestamp.

Regression introduced by diegosouzapw#11360, shipped in v3.8.50.

Two assertions in repro-combo-persisted-cooldown-preskip.test.ts encoded the
buggy behavior as intended ("skips an unavailable connection whose cooldown
already expired") and are realigned to the corrected contract, plus a case for
the orphan state (unavailable with no timestamps at all).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant