Skip to content

fix(combo): clear LKGP pin when its target fails, not only set it on success - #10034

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
hartmark:fix/lkgp-clear-on-target-failure
Aug 13, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
hartmark:fix/lkgp-clear-on-target-failure

Conversation

@hartmark

Copy link
Copy Markdown
Contributor

Summary

setLKGP() (Last Known Good Provider, src/lib/db/settings/lkgp.ts) was only
ever called on success — nothing invalidated a pin once that provider started
failing, so a separate subsequent request kept re-selecting the same
just-failed target via applyStrategyOrdering.ts's LKGP reordering.

Live incident: an OpenClaw request to combo default
(routerStrategy: lkgp) got a real reasoning + apply_patch tool call from
opencode-zen/big-pickle. Three separate follow-up requests over the next
~2 minutes each independently re-selected the same big-pickle target and
each timed out with 504 Stream produced no non-ping SSE event within 95000ms before the client gave up — instead of failing over to any of the
combo's other 12 models.

Root cause: confirmed via code read that circuit breaker and model
lockout deliberately don't react to this failure class —
isStreamReadinessFailureErrorBody exempts STREAM_READINESS_TIMEOUT/
combo_target_timeout 504s from tripping the provider breaker, and
REQUEST_SCOPED_UPSTREAM_ERROR_CODES suppresses model-lockout recording for
the same class (both intentional — avoid poisoning a healthy provider on
request-specific timing). Nothing else in the system was clearing the stale
LKGP pin, so it kept winning target-selection ordering for every new
top-level request.

⚠️ base-red inherited: #9985

Fix

  • Add clearLKGP(comboName, modelId) to src/lib/db/settings/lkgp.ts,
    exported through settings.ts/localDb.ts.
  • Call it (mirroring the existing setLKGP-on-success call pattern exactly,
    same two keys — target.executionKey and combo.id || combo.name) in
    both of combo.ts's per-target failure paths: handleComboChat's "Done
    retrying this model" block and handleRoundRobinCombo's structurally
    identical twin — right where a target is finally given up on and the loop
    moves to the next one.

Test plan

  • TDD: new regression test in tests/unit/combo-routing-engine.test.ts
    ("clears LKGP after the last-known-good target fails") reproduces the
    exact live scenario — confirmed failing against the pre-fix code, passing
    after.
  • Added direct unit coverage for clearLKGP itself in
    tests/unit/db-settings-crud.test.ts (deletes only the targeted key,
    sibling keys survive; no-op on an unset key doesn't throw) and registered
    the new export in db-settings-split.test.ts's public API surface
    characterization test.
  • Full combo/LKGP-related suite (combo-routing-engine, db-settings-crud,
    db-settings-split, combo-strategy-fallbacks,
    combo-selected-connection-success,
    delete-provider-connection-invalidates-lkgp-8887, db-read-cache) —
    183/183 passing.
  • npx tsc --noEmit — clean for all changed files (pre-existing unrelated
    errors elsewhere in the same test files confirmed identical against a
    pristine upstream/release/v3.8.50 checkout, zero diff at those lines).
  • npm run lint — clean (new test's any usage properly typed instead of
    inflating the file's frozen any-budget suppression).

@hartmark
hartmark requested a review from diegosouzapw as a code owner August 10, 2026 16:45
@hartmark
hartmark force-pushed the fix/lkgp-clear-on-target-failure branch 2 times, most recently from d635c55 to 54f491e Compare August 12, 2026 17:21
@hartmark
hartmark force-pushed the fix/lkgp-clear-on-target-failure branch from 54f491e to 0fdb6b0 Compare August 12, 2026 19:11
hartmark added a commit to hartmark/OmniRoute that referenced this pull request Aug 12, 2026
…success

setLKGP() was only ever called on success — nothing invalidated a "last
known good provider" pin once that provider started failing, so a
*separate* subsequent request kept re-selecting the same just-failed
target via applyStrategyOrdering.ts's LKGP reordering.

Live incident: an OpenClaw request to combo "default" (routerStrategy:
lkgp) got a real reasoning + apply_patch tool call from
opencode-zen/big-pickle, then 3 separate follow-up requests over the
next ~2 minutes each independently re-selected the same big-pickle
target and each timed out with "504 Stream produced no non-ping SSE
event within 95000ms" before the client gave up — instead of failing
over to any of the combo's other 12 models.

Root cause confirmed via code read: circuit breaker and model lockout
deliberately don't react to this failure class (isStreamReadinessFailureErrorBody
exempts STREAM_READINESS_TIMEOUT/combo_target_timeout 504s from tripping
the provider breaker, and REQUEST_SCOPED_UPSTREAM_ERROR_CODES suppresses
model-lockout recording for the same class — both intentional, to avoid
poisoning a healthy provider on request-specific timing). Nothing else
in the system was clearing the stale LKGP pin, so it kept winning
target-selection ordering for every new top-level request.

Fix: add clearLKGP(comboName, modelId) to src/lib/db/settings/lkgp.ts,
export it through settings.ts/localDb.ts, and call it (mirroring the
existing setLKGP-on-success call pattern exactly, same two keys) in both
combo.ts's per-target failure paths -- handleComboChat's "Done retrying
this model" block and handleRoundRobinCombo's structurally identical
twin -- right where a target is finally given up on and the loop moves
to the next one.

TDD: new regression test in tests/unit/combo-routing-engine.test.ts
("clears LKGP after the last-known-good target fails") reproduces the
exact live scenario -- confirmed failing against the pre-fix code,
passing after. Added direct unit coverage for clearLKGP itself in
tests/unit/db-settings-crud.test.ts (deletes only the targeted key,
sibling keys survive; no-op on an unset key doesn't throw) and
registered the new export in db-settings-split.test.ts's public API
surface characterization test.

Test plan:
- Full combo/LKGP-related suite (combo-routing-engine, db-settings-crud,
  db-settings-split, combo-strategy-fallbacks,
  combo-selected-connection-success,
  delete-provider-connection-invalidates-lkgp-8887, db-read-cache) --
  183/183 passing.
- npx tsc --noEmit -- clean for all changed files (pre-existing unrelated
  errors elsewhere in the same test files confirmed identical against a
  pristine upstream/release/v3.8.50 checkout, zero diff at those lines).
- npm run lint -- clean (new test's any usage properly typed, not left
  to inflate the file's frozen any-budget suppression).

⚠️ base-red inherited: diegosouzapw#9985
@hartmark
hartmark force-pushed the fix/lkgp-clear-on-target-failure branch from 0fdb6b0 to e5cc749 Compare August 12, 2026 19:24
hartmark added a commit to hartmark/OmniRoute that referenced this pull request Aug 12, 2026
@diegosouzapw

Copy link
Copy Markdown
Owner

Verified via TDD in an isolated probe: reverted combo.ts/settings.ts/settings/lkgp.ts/localDb.ts to base and confirmed the new regression test in combo-routing-engine.test.ts fails cleanly (LKGP pin stays set after the target's only retry fails); restored the PR's changes and the full combo/LKGP-related suite passes (88+53+42 tests across the touched and sibling files). clearLKGP() follows the existing db/settings domain-module and re-export conventions exactly, and the two call sites in combo.ts mirror the pre-existing setLKGP-on-success pattern (same fire-and-forget shape, same two keys, same non-fatal error handling). eslint (with the frozen suppressions) and typecheck:core are clean on all touched files. One minor, non-blocking observation: the protectedPriorityTarget early-return branch in handleComboChat's failure path (priority-strategy targets with fallbackOnlyOnQuotaExhaustion) skips the new clearLKGP call — worth a follow-up if that combination is ever expected to interact with LKGP pins, but it's outside the scope of the reported incident (lkgp strategy) and doesn't block this merge. Merge-ready as-is.


Note: the base-red that was failing this PR's CI (#9985) was drained today — base-reds PR #10213 just merged into release/v3.8.50. A rebase/sync onto the current tip should bring your checks green.

@diegosouzapw
diegosouzapw merged commit c9daf99 into diegosouzapw:release/v3.8.50 Aug 13, 2026
5 checks passed
@diegosouzapw

Copy link
Copy Markdown
Owner

Merged — thanks @hartmark! LKGP pins now clear on failure; the new settings module follows the db domain-module pattern cleanly.

fenix007 pushed a commit to fenix007/OmniRoute that referenced this pull request Aug 20, 2026
…success (diegosouzapw#10034)

setLKGP() was only ever called on success — nothing invalidated a "last
known good provider" pin once that provider started failing, so a
*separate* subsequent request kept re-selecting the same just-failed
target via applyStrategyOrdering.ts's LKGP reordering.

Live incident: an OpenClaw request to combo "default" (routerStrategy:
lkgp) got a real reasoning + apply_patch tool call from
opencode-zen/big-pickle, then 3 separate follow-up requests over the
next ~2 minutes each independently re-selected the same big-pickle
target and each timed out with "504 Stream produced no non-ping SSE
event within 95000ms" before the client gave up — instead of failing
over to any of the combo's other 12 models.

Root cause confirmed via code read: circuit breaker and model lockout
deliberately don't react to this failure class (isStreamReadinessFailureErrorBody
exempts STREAM_READINESS_TIMEOUT/combo_target_timeout 504s from tripping
the provider breaker, and REQUEST_SCOPED_UPSTREAM_ERROR_CODES suppresses
model-lockout recording for the same class — both intentional, to avoid
poisoning a healthy provider on request-specific timing). Nothing else
in the system was clearing the stale LKGP pin, so it kept winning
target-selection ordering for every new top-level request.

Fix: add clearLKGP(comboName, modelId) to src/lib/db/settings/lkgp.ts,
export it through settings.ts/localDb.ts, and call it (mirroring the
existing setLKGP-on-success call pattern exactly, same two keys) in both
combo.ts's per-target failure paths -- handleComboChat's "Done retrying
this model" block and handleRoundRobinCombo's structurally identical
twin -- right where a target is finally given up on and the loop moves
to the next one.

TDD: new regression test in tests/unit/combo-routing-engine.test.ts
("clears LKGP after the last-known-good target fails") reproduces the
exact live scenario -- confirmed failing against the pre-fix code,
passing after. Added direct unit coverage for clearLKGP itself in
tests/unit/db-settings-crud.test.ts (deletes only the targeted key,
sibling keys survive; no-op on an unset key doesn't throw) and
registered the new export in db-settings-split.test.ts's public API
surface characterization test.

Test plan:
- Full combo/LKGP-related suite (combo-routing-engine, db-settings-crud,
  db-settings-split, combo-strategy-fallbacks,
  combo-selected-connection-success,
  delete-provider-connection-invalidates-lkgp-8887, db-read-cache) --
  183/183 passing.
- npx tsc --noEmit -- clean for all changed files (pre-existing unrelated
  errors elsewhere in the same test files confirmed identical against a
  pristine upstream/release/v3.8.50 checkout, zero diff at those lines).
- npm run lint -- clean (new test's any usage properly typed, not left
  to inflate the file's frozen any-budget suppression).

⚠️ base-red inherited: diegosouzapw#9985

(cherry picked from commit c9daf99)
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…success (diegosouzapw#10034)

setLKGP() was only ever called on success — nothing invalidated a "last
known good provider" pin once that provider started failing, so a
*separate* subsequent request kept re-selecting the same just-failed
target via applyStrategyOrdering.ts's LKGP reordering.

Live incident: an OpenClaw request to combo "default" (routerStrategy:
lkgp) got a real reasoning + apply_patch tool call from
opencode-zen/big-pickle, then 3 separate follow-up requests over the
next ~2 minutes each independently re-selected the same big-pickle
target and each timed out with "504 Stream produced no non-ping SSE
event within 95000ms" before the client gave up — instead of failing
over to any of the combo's other 12 models.

Root cause confirmed via code read: circuit breaker and model lockout
deliberately don't react to this failure class (isStreamReadinessFailureErrorBody
exempts STREAM_READINESS_TIMEOUT/combo_target_timeout 504s from tripping
the provider breaker, and REQUEST_SCOPED_UPSTREAM_ERROR_CODES suppresses
model-lockout recording for the same class — both intentional, to avoid
poisoning a healthy provider on request-specific timing). Nothing else
in the system was clearing the stale LKGP pin, so it kept winning
target-selection ordering for every new top-level request.

Fix: add clearLKGP(comboName, modelId) to src/lib/db/settings/lkgp.ts,
export it through settings.ts/localDb.ts, and call it (mirroring the
existing setLKGP-on-success call pattern exactly, same two keys) in both
combo.ts's per-target failure paths -- handleComboChat's "Done retrying
this model" block and handleRoundRobinCombo's structurally identical
twin -- right where a target is finally given up on and the loop moves
to the next one.

TDD: new regression test in tests/unit/combo-routing-engine.test.ts
("clears LKGP after the last-known-good target fails") reproduces the
exact live scenario -- confirmed failing against the pre-fix code,
passing after. Added direct unit coverage for clearLKGP itself in
tests/unit/db-settings-crud.test.ts (deletes only the targeted key,
sibling keys survive; no-op on an unset key doesn't throw) and
registered the new export in db-settings-split.test.ts's public API
surface characterization test.

Test plan:
- Full combo/LKGP-related suite (combo-routing-engine, db-settings-crud,
  db-settings-split, combo-strategy-fallbacks,
  combo-selected-connection-success,
  delete-provider-connection-invalidates-lkgp-8887, db-read-cache) --
  183/183 passing.
- npx tsc --noEmit -- clean for all changed files (pre-existing unrelated
  errors elsewhere in the same test files confirmed identical against a
  pristine upstream/release/v3.8.50 checkout, zero diff at those lines).
- npm run lint -- clean (new test's any usage properly typed, not left
  to inflate the file's frozen any-budget suppression).

⚠️ base-red inherited: diegosouzapw#9985
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants