Skip to content

Fix/codex quota header leak - #13638

Closed
fenix007 wants to merge 84 commits into
diegosouzapw:release/v3.8.51from
fenix007:fix/codex-quota-header-leak
Closed

fenix007 wants to merge 84 commits into
diegosouzapw:release/v3.8.51from
fenix007:fix/codex-quota-header-leak

Conversation

@fenix007

Copy link
Copy Markdown
Contributor

Summary

  • Describe the user-facing or operational change.

Related Issues

  • Closes #
  • Related to #

Validation

Choose the change type and focused loop from the
Contribution Golden Path. The full unit suite,
Vitest, the 60% coverage gate, and the production build all run in CI on this PR (#8329):

  • Change type: provider / routing / UI / i18n / CLI / DB / build-deploy / other
  • Focused tests and category gates from the golden path
  • npm run lint
  • Reconciled with the current active release base; focused checks rerun afterward
  • Production-code changes include a new or updated automated test in this PR
  • SonarQube is temporarily opt-in while the private project has no quota; it is not a PR gate.

Tests Added Or Updated

  • List every changed or added automated test file.
  • If no production code changed, state that here.

Coverage Notes

  • If this PR changes src/, open-sse/, electron/, or bin/, explain which tests cover the change.
  • If coverage moved down in any touched file, explain why and what follow-up task will recover it.

Reviewer Notes

  • Call out any risky areas, migrations, feature flags, or manual validation that reviewers should know about.

fenix007 and others added 30 commits August 20, 2026 11:39
…osouzapw#7171)

* chore(ci): add .mergify.yml to main — Mergify only reads config from the default branch (diegosouzapw#7168)

* fix(executors): disable parallel tools for Codex Responses Lite

* docs(changelog): add Responses Lite fix fragment

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: alexey.nazarov@softmg.ru <alexey.nazarov@softmg.ru>
(cherry picked from commit 1636a8e)
* chore(ci): add .mergify.yml to main — Mergify only reads config from the default branch (diegosouzapw#7168)

* fix(ci): add the auto-enqueue pull_request_rule to the Mergify config (queue_conditions alone are eligibility-only) (diegosouzapw#7179)

* fix(ci): migrate Mergify auto-enqueue to merge_protections_settings.auto_merge_conditions (rules-based path is EOL 2026-07-16) (diegosouzapw#7216)

* fix(ci): drop Mergify batch settings (batching is a paid-tier feature; free plan queue is serial) (diegosouzapw#7220)

* fix(ci): merge queue tolerates the advisory dast-smoke failure (its GH-hosted build hang dequeued every attempt) (diegosouzapw#7225)

* test(ci): make the diegosouzapw#6634 selfref guard hermetic — main's copy hard-fails every PR (diegosouzapw#7341)

main's copy of this test still does git I/O inside a unit test:

    const baseSrc = git(['show', 'origin/main:' + FILE]);

Runners check out a shallow single ref, so origin/main does not resolve and the
test dies with 'fatal: invalid object name origin/main'. Every PR into main
fails Unit Tests (7/8) on it — today that is diegosouzapw#7313, diegosouzapw#7315, diegosouzapw#7316, diegosouzapw#7334, diegosouzapw#7336
and diegosouzapw#7337, six PRs red on a defect none of them introduced. diegosouzapw#7313 has no other
red at all.

release/v3.8.49 already carries a fix (2e42b8e, diegosouzapw#7174: try/catch, fetch
origin/main on demand, t.skip() when unreachable), but it only reaches main at
release time — so main stays broken for the whole cycle. Cherry-picking it would
also import a new problem: PR Test Policy classifies t.skip() as a silenced
assertion, which we watched it correctly catch on diegosouzapw#7300 today.

This is the hermetic version instead (ported from diegosouzapw#7327, which does the same for
the release branch): read the file straight off disk, compare against an empty
base so baseTaut/baseExtTaut are 0 — the strictest possible comparison point —
and call evaluateMasking() directly. No git ref, no fetch, no skip, nothing the
runner's checkout depth can break.

The diegosouzapw#6634 regression stays covered: the guard's logic lives in
SELF_TEST_FIXTURE_RE (check-test-masking.mjs:337), not in the test. Proven both
ways on main before committing — neutralise SELF_TEST_FIXTURE_RE to /$^/ and
the test FAILS; restore it and it passes 2/2, with check-test-masking.mjs left
byte-identical.

Co-authored-by: growab <nekron@icloud.com>

* chore(quality): tighten main's coverage baseline to the CI's real numbers (diegosouzapw#7347)

main's ratchet had been failing --require-tighten on every PR: 11 metrics
improved but the baseline was never tightened. Same class as the diegosouzapw#6634
selfref guard — an infra fix that lands only on the release branch leaves
main red for the whole cycle, and every PR into main pays for it.

Values are the merged-coverage numbers from a run on main itself (a local
run measures ~68% vs CI's ~80%; the baseline's own note warns about that
gap). Only the 11 coverage values change — gitleaks and semgrepFindings
keep main's own state.

No changelog fragment: diegosouzapw#7326 carries it on release/v3.8.49, and a second
one here would double the entry at release time.

* feat(providers): add xAI OAuth PKCE

* docs(changelog): note xAI OAuth provider

* test(xai): assert OAuth refresh client id

* refactor(oauth): rebaseline OAuthModal wiring note (file-size cap)

Correct the file-size-baseline.json annotation for the xai-oauth PKCE
provider-switch branch in OAuthModal.tsx (993->998, +5) to match the
modal's existing historical-progression annotation style (969->989->
993->998; structural shrink tracked in diegosouzapw#3501). The frozen value (998)
stays unchanged — only the annotation text is corrected.

tests/unit/oauth-providers-config.test.ts already sits exactly at its
frozen cap (845) after registering xai-oauth in the shared provider
enumerations (import, EXPECTED_PROVIDER_KEYS, EXPECTED_CONFIG_BY_PROVIDER,
REQUIRED_FIELDS_BY_PROVIDER). check:file-size reports 0 violations for
it, so no test move or baseline bump was needed.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* test(oauth): compact required-field arrays (file-size budget on frozen oauth-providers-config suite)

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: growab <nekron@icloud.com>
Co-authored-by: alexey.nazarov@softmg.ru <alexey.nazarov@softmg.ru>
Co-authored-by: Alex <4217955+fenix007@users.noreply.github.com>
(cherry picked from commit 9db5377)
* fix(api): enforce image generation API key auth

* fix(api): align image route auth guard with clientApiPolicy

The route-level guard added for image generation was stricter than the
authz middleware that already fronts /api/v1/* (src/proxy.ts →
clientApiPolicy), so requests the pipeline admits were 401'd by the
handler:

- A cookie-authenticated dashboard session was rejected under
  REQUIRE_API_KEY=true. The dashboard Media page
  (dashboard/cache/media) and the Playground call these routes with a
  session and no Bearer — the same mismatch already fixed for
  /api/playground/presets.
- A presented invalid key was rejected even with REQUIRE_API_KEY=false,
  where clientApiPolicy (diegosouzapw#2257) and the sibling /v1/embeddings and
  /v1/web/fetch routes degrade a stale CLI key to anonymous instead.

Extract the shared guard into shared/utils/clientApiRouteAuth so both
image routes (and future /v1 handlers) mirror the middleware contract
instead of re-deriving it, and drop the now-dead auth imports.

Also switch the call-log attribution fallback back to `||`: with `??`,
an empty-string apiKeyId/apiKeyName would be persisted verbatim and
would block the request-scoped context, which the previous
`entry.apiKeyId || null` never did.

Tests: cover the dashboard-session and keyless-mode-invalid-key
branches, and split the auth/attribution cases into
image-generation-route-auth.test.ts to stay under the 800-line
new-test-file cap.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: alexey.nazarov@softmg.ru <alexey.nazarov@softmg.ru>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 53f8284)
…out (diegosouzapw#10016)

A combo target that stalls past comboTargetTimeoutMs is aborted by
buildTargetTimeoutRunner, which swallows the resulting rejection behind its
synthetic 524. Nothing marks the account unavailable — correctly, since a stall
is not a quota/auth failure — so the diegosouzapw#6219 eviction on the generic
markAccountUnavailable -> shouldFallback path in chat.ts never ran. The session
pin therefore survived its full TTL and every following request in that session
was handed straight back to the account that had just stalled.

Seen in production on combo "coding" [priority]: one codex account pinned for a
30-minute TTL, four consecutive requests, four 120s timeouts, "all targets
exhausted" each time, while four sibling codex accounts stayed healthy and
unused.

Classify the abort reason (new dependency-free leaf comboAbortReasons.ts) and
evict the connection-matched pin. Only a genuine per-model timeout evicts: a
client disconnect or a hedge cancellation says nothing about account health, so
those keep the pin and its prompt-cache locality. Eviction is best-effort and
never breaks the dispatch path.

The dispatch itself moves into a new seam, chatDispatch.ts, which merges the
per-model abort signal into the outgoing request, runs executeChatWithBreaker,
and owns the eviction on both the rejection and failed-result paths. Keeping
that logic out of the frozen god-file leaves chat.ts one line SHORTER than
before (1844 -> 1843).

Co-authored-by: alexey.nazarov@softmg.ru <alexey.nazarov@softmg.ru>
Co-authored-by: fenix007 <fenix007@users.noreply.github.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
(cherry picked from commit 20fcb8d)
Port of upstream PR diegosouzapw#8307 (open, targets v3.8.50) onto the 3.8.48 base:
some ChatGPT accounts can use Codex but lack the requested image model
entitlement. Lock the model on that connection and retry one sibling
account for that exact upstream response; ordinary client 400s remain
single-attempt failures.

Adapted for 3.8.48: the codex handler here surfaces the upstream error
body as a raw JSON string (no sanitizeImageProviderError parsing), so
the model-access matcher unwraps JSON strings before comparing. The
3.8.50-only pieces (handleCodexImageEdit context, diegosouzapw#6928 no-auth branch,
release CI/docs/stryker changes) are not included.

(ported from upstream PR diegosouzapw#8307, head becd2be)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Trigger on the stable branch and on *-fork.* tags (never v*, so the
inherited upstream release workflows stay quiet). Mirrors the upstream
docker-publish.yml main image: target runner-base, linux/amd64+arm64.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Dockerfile's shared apt cache mounts deadlock when amd64 and arm64
build concurrently in one buildx invocation (apt list lock held by the
sibling platform). Mirror upstream docker-publish.yml: one job per
platform pushing by digest (arm64 on a native arm runner instead of
QEMU), then a merge job that assembles the manifest list and applies
the tags.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
npm run build SIGABRTs in the Next.js build worker on current node:24
minors with better-sqlite3 12.11.1 (this branch's lockfile), on both
amd64 and arm64. Upstream's working v3.8.48 images were built with
NODE_VERSION=24.18.0 (verified from the published image config), so pin
that exact base. Upstream escaped the same crash on 3.8.50 by moving to
better-sqlite3 ^13.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…iegosouzapw#6710)

* fix(codex): surface capacity errors embedded in 200-OK SSE streams

Codex sometimes answers with HTTP 200 and a text/event-stream body whose
payload carries a transient error mid-stream (e.g. "Selected model is at
capacity...", server_is_overloaded, service_unavailable_error). Because the
outer HTTP status was 200, this looked like a successful response to every
caller — no retry, no circuit breaker, and no combo/account fallback ever
engaged, so a healthy account sat idle while the request silently failed or
truncated.

Add peekCodexSseTransientError() to open-sse/executors/codex.ts: it peeks the
first bytes of a text/event-stream Codex response, pattern-matches the known
transient-error signatures, and converts a match into a real 503 Response via
errorResponse() (Hard Rule diegosouzapw#12 — sanitized, never raw upstream text). A 503 is
already a recognized provider-failure status in accountFallback.ts, so combo
routing and connection cooldown pick it up automatically. When no error
signature is found, the peeked prefix is prepended back onto the remaining
upstream body so the passthrough stays byte-identical to the unmodified
response.

Regression guard: tests/unit/codex-sse-capacity-fallback.test.ts — a
model-at-capacity payload and a server_is_overloaded/service_unavailable_error
payload both convert to 503; a normal single-chunk SSE stream and one split
across multiple network chunks both reassemble byte-for-byte unchanged.

Inspired-by: decolua/9router#2452 (sub-bug #3 only —
OmniRoute already covers PR diegosouzapw#2452's other two sub-bugs: service_tier "fast"
normalization and reasoning_effort "max" normalization).

Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>

* chore(6710): re-sync onto release tip; CHANGELOG → changelog.d fragment (fragments-first)

---------

Co-authored-by: ryanngit <74137224+ryanngit@users.noreply.github.com>
(cherry picked from commit 3f457cc)
…le-reader peek) (diegosouzapw#7526)

Every non-streaming Codex chat request for a ChatGPT-account connection failed
instantly with [502]: Response body is already used (reset after 1m). The
streaming/playground path was unaffected.

Root cause: peekCodexSseTransientError (open-sse/executors/codex.ts) peeked the
SSE prefix with response.body.getReader(), then called reader.releaseLock() and
response.body.getReader() a SECOND time on the same already-disturbed body to
build the replacement stream. Re-acquiring a reader on a disturbed body throws
on undici ('Response body is already used'); chatCore's generic upstream-error
handling then stamped the TypeError with a default 60s cooldown, masking a pure
code defect as a rate limit (and tripping the codex circuit breaker).

Fix: keep the single reader already held; never touch response.body again.

TDD: a getReader spy that throws on the 2nd acquire reproduces the exact hazard
— 1 test RED against the release code, 2/2 GREEN with the fix; the replacement
body stays byte-identical to the upstream SSE. No regression across the codex
unit suite. Reproduced live on the VPS 2026-07-16.

(cherry picked from commit a06ddb3)
…onse.body in peek (diegosouzapw#7570)

Non-stream Codex (ChatGPT account) chat 502'd with "Response body is already
used". On the wreq-js TLS-fingerprint transport the Response is backed by a
native body handle, and merely accessing response.body disturbs it so a later
.text() throws. The Codex non-stream upstream response has an empty content-type,
so peekCodexSseTransientError early-returns — but its guard evaluated
!response.body (touching .body) before the content-type check, consuming the
body; chatCore's readNonStreamingResponseBody then re-read it and 502'd.
Streaming was unaffected. Reorder the guard to check content-type first.

Validated live on the VPS (192.168.0.15): codex/gpt-5.5 and codex/gpt-5.6-terra
non-stream now return 200; streaming still works. Regression test drives the real
peek with a destructive-.body mock.

(cherry picked from commit 6e48903)
…odex Responses Lite (diegosouzapw#7821) (diegosouzapw#7957)

* fix(sse): preserve parallel_tool_calls for GPT-5.6 ultra/max delegation under Codex Responses Lite (diegosouzapw#7821)

* fix(codex): drop over-broad parallel_tool_calls allowlist entry — keep diegosouzapw#2608 stripping intact (diegosouzapw#7821)

The static RESPONSES_API_ALLOWLIST addition made parallel_tool_calls survive for
ALL models, breaking the diegosouzapw#2608 non-passthrough stripping guarantee for gpt-5.5.
The real diegosouzapw#7821 fix (isCodexDelegationDependentModel gating in
enforceCodexResponsesLiteParallelToolCalls) is model/effort-scoped and does not
need the allowlist entry — native Codex traffic returns before the allowlist runs.

(cherry picked from commit 0d4fbfe)
…pw#8020) (diegosouzapw#8043)

peekCodexSseTransientError() ran before chatCore's normal
readiness/idle-timeout pipeline and read the first SSE chunk with a
bare reader.read() — no timeout wrapper. A 200 text/event-stream body
that never emitted a byte hung for ~15min (901399ms observed) before
the platform killed the connection and surfaced a generic 502.

Wrap the peek loop's read and the re-assembled passthrough body's
pull() in readStreamChunkWithTimeout, bounded PER READ (not a total
deadline) so a long-but-alive reasoning stream keeps resetting the
window on every chunk it emits. On timeout the reader is cancelled and
the request now fails fast with a 504 instead of hanging.

New small module open-sse/executors/codex/bodyTimeout.ts holds the
wrapping helpers to keep codex.ts within its frozen size baseline.

(cherry picked from commit 5e234d5)
)

Validated in local merge-train (devbox-vm-06-dev002) @ combined-tip (FAST gates green: static + changed tests + vitest — only pre-existing audit.test.ts flake). Evidence: /home/diegosouzapw/dev/proxys/OmniRoute/.claude/worktrees/merge-train-20260805-213228-suite.log

(cherry picked from commit 3f9507f)
…ion on model-unsupported 400 (diegosouzapw#10525)

* fix(settings,auth): default debugMode to false and skip account rotation on model-unsupported 400

* fix(auth): disambiguate model-unsupported from auth-credential 400

The model-unsupported guard used MODEL_ACCESS_DENIED_PATTERNS directly,
which also matches auth-credential errors like 'invalid api key for
model X'. Add the AUTH_CREDENTIAL_ERROR_PATTERNS exclusion (same as
checkFallbackError) and use provider_model_unsupported log reason.

Addresses maintainer feedback on PR diegosouzapw#10525

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* fix(auth): narrow model-unsupported guard to avoid misclassifying account-scoped entitlement 400s

The diegosouzapw#10460 guard reused MODEL_ACCESS_DENIED_PATTERNS directly, which also
matches ambiguous "access"/"permission" phrasing (e.g. "does not have
permission to access this model") that commonly signals an ACCOUNT-scoped
entitlement gap (PRO vs free tier) rather than a genuinely provider-wide
unsupported model — a different account of the same provider may still
have access, so those must keep rotating normally instead of being
short-circuited.

Extract isProviderModelUnsupported400() in accountFallback.ts: reuses the
same AUTH_CREDENTIAL_ERROR_PATTERNS exclusion checkFallbackError's 400
branch already applies, narrowed to a strict subset of unambiguous
"provider does not serve this model at all" phrasings. auth.ts now calls
this shared helper instead of testing the broader patterns in isolation,
and exposes the sanitized reason ("provider_model_unsupported") on the
returned result, not just in the log line.

Also fix DATA_DIR test-isolation ordering in
account-fallback-service.test.ts: it was assigned after the first
dynamic import of accountFallback.ts, which transitively imports
src/lib/db/core.ts (DATA_DIR is captured once at module-load time), so
the intended isolated test directory was silently never used. Move the
assignment before any transitive DB import, and add regression tests for
the 3-account rotation contract: exactly one upstream call for an
unambiguous provider-wide 400 with the combo advancing to the next
target, continued rotation for account-scoped 401/403/429 and for the
permission/entitlement 400 case that motivated this narrowing.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
(cherry picked from commit ebf0bf9)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
diegosouzapw#9207)

The combo success path called recordProviderSuccess (cooldown-only)
without notifying the circuit breaker. When a provider breaker entered
HALF_OPEN after repeated failures, successful probe requests never
transitioned it back to CLOSED -- the breaker stayed stuck indefinitely.

Production evidence: agy breaker HALF_OPEN with 699 requests at 98%
success rate, never recovering.

Root cause: combo.ts calls recordProviderSuccess from
providerCooldownTracker.ts (resets cooldown failureCount only) but
never calls breaker._onSuccess(). The failure path in accountFallback.ts
calls breaker._onFailure(), creating an asymmetry.

Fix: add recordProviderSuccess to accountFallback.ts as the symmetric
counterpart of recordProviderFailure. Uses getProviderBreaker (not
configureProviderBreaker) to avoid overwriting the breaker's resetTimeout
with default profile values. Calls breaker._onSuccess() for all non-OPEN
states (CLOSED/DEGRADED/HALF_OPEN), matching execute()'s behavior.

(cherry picked from commit f10dca4)
…egosouzapw#9342)

* fix(combo): keep queue/network timeouts out of the provider breaker

A single-model network error (ECONNREFUSED / proxy_unreachable) means we never
reached the provider — the provider may be healthy while only the network path
is broken. OmniRoute's own rate-limit queue timeouts are backpressure we
applied, not an upstream failure. Neither should trip the whole-provider
breaker.

- chatPredicates: the single-model path excludes proxy_unreachable and
  RATE_LIMIT_QUEUE_* from the provider-breaker trip.
- accountFallback.recordProviderFailure: isQueueTimeout short-circuits before
  the breaker ever counts (combo.ts already flags it from errorText).
- chat.ts: the queue/network guard on the allRateLimited _onFailure trip.

Deliberately leaves the combo same-provider dead-proxy leg (diegosouzapw#8376) intact:
there a proxy_unreachable on the next same-provider target must still be able
to open the breaker, or a dead proxy burns every attempt until the 503
max-retry limit.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* fix(resilience): dedup same-provider network errors per event

Same-provider combo targets can all fail the same single network event (a VPN
blip) within one request. Without a dedup each target counts once toward the
provider breaker, so one transient blip opens the whole-provider breaker while
the provider is healthy — the antigravity outage this branch originally chased.

recordProviderFailure now keeps a short per-provider window (10s) for
proxy_unreachable failures: the first network error in a window counts, the rest
of that window are the same event and return. A genuinely dead proxy keeps
failing across requests (past the window) and still accumulates to its
threshold, so the diegosouzapw#8376 dead-proxy protection is not weakened.

Covered by tests/unit/breaker-network-error-guard.test.ts: same-window errors
dedup to one, cross-window errors still open the breaker.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
(cherry picked from commit 47c819d)
…zapw#9164)

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
(cherry picked from commit 3898305)
Upstream diegosouzapw#9342 and diegosouzapw#9164 assert through chatPredicates.ts and the
request-scoped predicate machinery (isRequestScopedUpstreamFailure,
shouldSkipConnDisable), none of which exist on the 3.8.48 base — the
guard lives inline in the frozen chat.ts and the classifier is the only
predicate available here. Assert the guard at its real dispatch site and
drop the assertions for machinery this base does not have.

Ported subset therefore covers the provider-breaker and combo-fallback
legs plus the local-capacity classifier; the connection-cooldown leg
rides on the later 3.8.50 predicate refactor and is intentionally out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…success (diegosouzapw#10034)

setLKGP() was only ever called on success — nothing invalidated a "last
known good provider" pin once that provider started failing, so a
*separate* subsequent request kept re-selecting the same just-failed
target via applyStrategyOrdering.ts's LKGP reordering.

Live incident: an OpenClaw request to combo "default" (routerStrategy:
lkgp) got a real reasoning + apply_patch tool call from
opencode-zen/big-pickle, then 3 separate follow-up requests over the
next ~2 minutes each independently re-selected the same big-pickle
target and each timed out with "504 Stream produced no non-ping SSE
event within 95000ms" before the client gave up — instead of failing
over to any of the combo's other 12 models.

Root cause confirmed via code read: circuit breaker and model lockout
deliberately don't react to this failure class (isStreamReadinessFailureErrorBody
exempts STREAM_READINESS_TIMEOUT/combo_target_timeout 504s from tripping
the provider breaker, and REQUEST_SCOPED_UPSTREAM_ERROR_CODES suppresses
model-lockout recording for the same class — both intentional, to avoid
poisoning a healthy provider on request-specific timing). Nothing else
in the system was clearing the stale LKGP pin, so it kept winning
target-selection ordering for every new top-level request.

Fix: add clearLKGP(comboName, modelId) to src/lib/db/settings/lkgp.ts,
export it through settings.ts/localDb.ts, and call it (mirroring the
existing setLKGP-on-success call pattern exactly, same two keys) in both
combo.ts's per-target failure paths -- handleComboChat's "Done retrying
this model" block and handleRoundRobinCombo's structurally identical
twin -- right where a target is finally given up on and the loop moves
to the next one.

TDD: new regression test in tests/unit/combo-routing-engine.test.ts
("clears LKGP after the last-known-good target fails") reproduces the
exact live scenario -- confirmed failing against the pre-fix code,
passing after. Added direct unit coverage for clearLKGP itself in
tests/unit/db-settings-crud.test.ts (deletes only the targeted key,
sibling keys survive; no-op on an unset key doesn't throw) and
registered the new export in db-settings-split.test.ts's public API
surface characterization test.

Test plan:
- Full combo/LKGP-related suite (combo-routing-engine, db-settings-crud,
  db-settings-split, combo-strategy-fallbacks,
  combo-selected-connection-success,
  delete-provider-connection-invalidates-lkgp-8887, db-read-cache) --
  183/183 passing.
- npx tsc --noEmit -- clean for all changed files (pre-existing unrelated
  errors elsewhere in the same test files confirmed identical against a
  pristine upstream/release/v3.8.50 checkout, zero diff at those lines).
- npm run lint -- clean (new test's any usage properly typed, not left
  to inflate the file's frozen any-budget suppression).

⚠️ base-red inherited: diegosouzapw#9985

(cherry picked from commit c9daf99)
Co-authored-by: Bryan Nathan <bryan@users.noreply.github.com>
(cherry picked from commit 06f41cd)
…ry (diegosouzapw#10217)

* fix(deps): bump nanoid, dompurify for 2 new Dependabot alerts (diegosouzapw#189, diegosouzapw#190)

Bumps: nanoid ^3.3.17 (was transitive, now overridden), dompurify ^3.4.13
(with monaco-editor scoped override). Closes Dependabot diegosouzapw#189, diegosouzapw#190.

Remaining diegosouzapw#182-diegosouzapw#188 (js-yaml + mermaid) already closed by diegosouzapw#9651 merge —
awaiting Dependabot re-scan.

npm audit → 0 vulnerabilities.

* fix(repo): harden .gitignore to also ignore a _tasks symlink (/_tasks)

_tasks is a SEPARATE nested git repo (gitignored). The pattern _tasks/ (trailing
slash) ignores only a directory, not a SYMLINK named _tasks. A self-referential
_tasks symlink can slip in via git add -A and, once pulled, checkout materializes
it over the real _tasks repo (destroying plans/specs/hands-off). Anchored /_tasks
ignores the symlink too, preventing re-capture.

* fix(combo): make failoverBeforeRetry actually skip the same-model retry

Both same-target retry loops (priority/auto and round-robin) checked
isTransient/maxRetries/providerExhausted but never consulted
config.failoverBeforeRetry, so a rate-limited model still got
maxRetries+1 back-to-back attempts on itself before falling back to a
sibling — the config option (diegosouzapw#2417) was only ever wired into
skipUpstreamRetry, a separate lower-level mechanism. Now the same-model
retry is skipped when failoverBeforeRetry is set AND a sibling target
is actually available; with no sibling left, it still retries same-model
since skipping would just burn the last attempt for nothing.

---------

Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@outlook.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
(cherry picked from commit d2fd88d)
…austed (diegosouzapw#10116)

* fix(account-fallback): classify 'insufficient credits' as credits-exhausted

Command Code returns 400 'You have insufficient credits to make this
request...' when an account's billing credits run out. The phrase was
missing from CREDITS_EXHAUSTED_SIGNALS, so the error stayed unclassified
(errorType=null) and the connection was never marked credits_exhausted —
getProviderCredentials kept re-selecting the same dead account on every
request instead of rotating to a healthy one.

Add 'insufficient credits'/'insufficient credit' to the signal list
(already used by antigravity429Engine.ts) so the error classifies as
QUOTA_EXHAUSTED and the account is skipped on subsequent selections.

* fix(account-fallback): harden insufficient-credit matching and preserve chatanywhere

Add the common 'insufficient credit balance' variation to
CREDITS_EXHAUSTED_SIGNALS alongside the Command Code 'insufficient
credits'/'insufficient credit' signals, and restore the consolidated
ChatAnywhere gateway entry that the stale snapshot removal would have
deleted when merging into release/v3.8.50.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com>
(cherry picked from commit be6f18b)
…iegosouzapw#10534)

* fix(sse): clear quota_exhausted cooldown when real window recovers

The claude-token-fallback combo was not auto-returning to Sonnet/Opus
after a subscription 429 recovered. maybeClearRecoveredQuotaState()
was honoring the synthetic 1h cooldown (SUBSCRIPTION_QUOTA_COOLDOWN_MS,
persisted when no upstream reset was parseable) instead of the REAL
per-window resetAt returned by the scheduled quota poller, so the
connection stayed locked long past the actual quota reset.

Add windowStillExhaustedAfterRealReset() and use it to decide recovery
per-quota-window: a quota_exhausted connection now clears as soon as no
governing window is still exhausted with a future-or-unknown real
reset, instead of waiting out the synthetic cooldown. Falls back to the
previous synthetic-cooldown guard when the fetch has no quota object at
all (degraded/failed shape) so existing behavior is unchanged there.

Preserves the existing kimi-coding partial-refresh semantics: an
exhausted window with no parseable resetAt still blocks recovery.

* fix(sse): preserve Claude extra-usage block from general quota recovery

maybeClearRecoveredQuotaState()'s new per-window recovery check (added in
this branch) only inspected usage.quotas, so a Claude connection blocked by
the extra-usage guard (lastErrorSource: "extra_usage") could be released
just because the session/weekly quota windows looked recovered, even while
extraUsage.queued was still true. Extra-usage blocking is orthogonal to
quota-window exhaustion and must only be released by
syncClaudeExtraUsageStateIfNeeded (buildClaudeExtraUsageConnectionUpdate).

Add a guard that keeps the connection locked when lastErrorSource is
"extra_usage", the blockExtraUsage policy is still enabled, and the fresh
usage snapshot still reports extraUsage.queued === true.

Add an integration test walking the real
fetchLiveProviderLimitsWithOptions -> syncClaudeExtraUsageStateIfNeeded ->
maybeClearRecoveredQuotaState call chain with recovered quota windows but
extraUsage.queued=true, asserting the connection stays unavailable with
lastErrorSource still "extra_usage".

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

---------

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
(cherry picked from commit 276b3df)
…e account (diegosouzapw#9708) (diegosouzapw#10792)

Merged via merge-train (release/v3.8.50, batch1 2026-08-20) — static gates (typecheck/file-size/complexity/cognitive/changelog) green on the combined tree; test:unit reds observed in the boarded run were verified pre-existing on the pure release tip (unrelated flake), not caused by this PR. Thanks for the contribution!

(cherry picked from commit aa32d2e)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
alexey.nazarov@softmg.ru and others added 22 commits September 5, 2026 13:54
Astra was catalogued but unusable: the Codex OAuth backend gates new models
by client version and rejected our 0.144.1 identity with HTTP 400 "requires
a newer version of Codex" (upstream issue diegosouzapw#12761). Advertise 0.153.4 from
both the shared inference identity and the Codex CLI profile, which have to
move together, and refresh the golden headers.

Correct the limits from the live OAuth catalog rather than the third-party
mirror: Astra reports max_context_window=872000, so 272000 was only the
pricing tier and under-advertised the usable window by ~3x. Astra therefore
shares GPT-5.6's Codex and public-API capability sets outright.

Add the `-ultra` tier: `ultra` is an OmniRoute-side alias that goes out as
wire effort `max` while keeping parallel tool calls for sub-agent
delegation, so Astra needs it for the same reason Sol and Terra do. Raise
its effort ceiling to `ultra` and bill Codex Fast at Astra's 2.5x rate.

Ported from upstream PR diegosouzapw#12759 (day-1 Astra support). The model-level
`timeoutMs` and `supportedThinkingEfforts` fields it also sets do not exist
on RegistryModel at this version and are deliberately left out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VjgvPkNiyaWNGmw1sutGHB
…odel id

The VS Code suffix matcher only knew the GPT-5.6 family, so `gpt-6-astra-max`
and `-ultra` never lost their alias and discovery expanded them a second time
into ids that do not exist upstream, such as `gpt-6-astra-max-high`. `max` and
`ultra` are catalog entries rather than standard effort suffixes, so the
standard pattern cannot catch them either.

Also update the last Codex fingerprint expectation left on 0.144.1, which the
integration suite asserts separately from the unit tests.

Both found by a codex review of the Astra work. Its third finding — Codex
`max` being normalized to `xhigh` by sanitizeReasoningEffortForProvider on the
HTTP path — is left alone: it predates this work, hits the whole GPT-5.6
family identically, and the production WebSocket path bypasses that sanitizer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VjgvPkNiyaWNGmw1sutGHB
Restrict scoped token credential access, authenticate dashboard ingress checks, trust verified proxy IPs, isolate reasoning caches, and confine outbound and WebDAV paths. Update vulnerable production dependencies.

Co-Authored-By: GPT-6 <noreply@openai.com>
Extract pure header and SSE status helpers, separate test fixtures and replay/WebDAV suites, and retain all 290 existing test cases. Correct the stale combo fallback fixture to assert the bounded same-account retry before provider fallback. Reduce explicit-any suppressions without changing size limits.

Co-Authored-By: GPT-6 <noreply@openai.com>
Adapt upstream PR diegosouzapw#12863 to v3.8.48: retain inclusive internal prompt totals and subtract cache only in Anthropic message_delta usage. Address Opus 5 review findings with full SSE regression coverage for cache-only usage, metadata, call logs and pricing. Record upstream review decisions and validation in FORK.md.

Co-Authored-By: GPT-6 <noreply@openai.com>
Adapt upstream PR diegosouzapw#13059 to the existing v3.8.48 sanitizer phases. Preserve tool arguments and response-schema fields while retaining actual constraint cleanup. Add regression coverage for OpenAI and Anthropic request translation and record upstream maintenance decisions.

Co-Authored-By: zcrew0x <zcrew0x@users.noreply.github.com>

Co-Authored-By: GPT-6 <noreply@openai.com>
Co-Authored-By: GPT-5.6 <noreply@openai.com>
Populate the sidebar build identifier for stable and manual branch builds while preserving release tag labels.

Co-Authored-By: GPT-6 <codex@openai.com>
Check successful fork image publications and compare commit ancestry before advertising updates. Link to the fork build, cache checks, and block the upstream updater in fork builds.

Co-Authored-By: GPT-6 <codex@openai.com>
Co-Authored-By: GPT-5.6 <noreply@openai.com>
Co-Authored-By: gpt-5.6-terra <noreply@openai.com>
workspacePlanType was written only by the OAuth code exchange, so a ChatGPT
subscription that lapsed after the account was connected kept reporting the
tier captured at connect time — indefinitely. The refreshed id_token carries
the current plan claim, so refreshCodexToken now re-derives it and returns a
providerSpecificDataPatch. Only the plan is patched: re-running the
team-vs-personal workspace heuristic could silently re-point an established
connection, so the workspace binding stays as the user selected it.

The patch is merged, not assigned, at both persist sites, so unrelated state
sharing the column (rate-limit windows, refresh circuit breaker) survives.

resolvePlanValue also discarded a live "free" outright and fell back to the
persisted tier, which hid exactly this case. That fallback exists because
antigravity returns a literal "Free" when it cannot read the tier at all, so
it is now scoped to antigravity/agy; every other provider — Codex included,
whose usage API reports "unknown" when it does not know — has its live plan
believed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MHHkZZNi3zvm9UpSQ7Sg7P
Port safe upstream schema and logging compatibility fixes for the frozen v3.8.48 base.

Co-Authored-By: GPT-6 <codex@openai.com>
Co-Authored-By: gpt-5.6 <noreply@openai.com>
Co-Authored-By: gpt-5.6 <noreply@openai.com>
Co-Authored-By: gpt-5.6 <noreply@openai.com>
Co-Authored-By: gpt-5.6 <noreply@openai.com>
Co-Authored-By: gpt-5.6 <noreply@openai.com>
…ility

Classify empty output consistently across retries, provider exhaustion and call-log summaries. Reject proven-incompatible combo targets, including pinned and nested dispatch, while retaining unknown capabilities.

Co-Authored-By: GPT-6 <noreply@openai.com>
Prefer known-fitting targets when a candidate appears too small, while retaining unknown and stale catalog entries for runtime fallback. Separate input and total context limits and honor effort-specific context overrides without relaxing hard capability checks.

Co-Authored-By: GPT-6 <noreply@openai.com>
Prevent Codex clients behind the multi-account router from treating the selected upstream account's quota as the client's own allowance, while preserving protocol turn state.

Co-Authored-By: GPT-6 Astra <noreply@openai.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

The isolated commit here (hide backend account quota headers) is clean and well-tested, but I
need to flag a direct tension with #10315: the current tip deliberately gives
x-codex-*-used-percent/reset/window/credits/plan-type headers forwarding priority specifically
because they used to be evicted by mistake and that broke real Codex CLI quota display for
single-account users. Stripping them unconditionally (as this PR does) would re-introduce that
exact problem for the common case, while fixing a real concern for the multi-account/combo
pooling case you describe.

I'd like to merge the underlying idea, but gated to only strip when the request actually went
through combo/multi-account routing — could you scope it that way? Also, to keep review
tractable, I'd cherry-pick just this one commit onto the tip rather than the full branch (which
carries a lot of history not related to this fix) — happy to do that with full credit to you.

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for raising this — the concern is real: when pool/combo routing picks an account that isn't the caller's, the x-codex-* quota headers expose that account's quota. I've opened #14116 to track exactly that, crediting you.

This PR itself can't merge as-is, for two reasons that aren't about the idea:

  1. The unconditional strip regresses fix(backend): deduplicate repeated upstream-header budget warnings #10315 — chatCore/responseHeaders.ts forwards those headers on purpose on the direct path (dropping them there was a previous bug), and the existing test middleware-header-strip-5849.test.ts › "streaming path keeps Codex quota headers…" fails with this change (8/9). The fix has to be conditional: strip only when the selected connection differs from the caller's own.
  2. The branch carries 83 unrelated commits, and the original commit has an AI co-author trailer, which the repo doesn't accept in history (Hard Rule fix(ci): fix npm publish auth — support vars.NPM_TOKEN #16 in AGENTS.md).

Closing this one; if you'd like to take #14116 on a fresh branch off release/v3.8.51 with the conditional strip, you'd have my full support — the design sketch is in the issue.

diegosouzapw added a commit to shubhayu-dev/OmniRoute that referenced this pull request Sep 19, 2026
…r strip

buildStreamingResponseHeaders already had the conditional strip
(isCodexQuotaHeader + the isForeignAccount option), but the sole
production caller (assembleStreamingResponseHeaders in
chatCore/streamingResponseHeaders.ts, called from chatCore.ts) never
passed it, so every response always took the default (falsy) path and
the diegosouzapw#13638/diegosouzapw#14116 quota-header leak stayed live in production — only
the unit tests exercised the strip, by hand-setting the flag on the
helper directly.

- Thread isForeignAccount through assembleStreamingResponseHeaders and
  set it from chatCore.ts's existing isCombo flag: pool/combo routing
  is exactly the signal for "this account was chosen by pool/combo
  routing, not requested directly" (see resolveComboTargets()/
  handleSingleModel in combo.ts); the direct path always has
  isCombo=false, so its headers are unaffected.
- Extract the header-drop predicate in responseHeaders.ts into
  shouldDropStreamingHeader() and drop the options default parameter
  value (accessed via `options?.` instead) — the added default
  parameter, not the extra branch itself, pushed
  buildStreamingResponseHeaders 1 point over the complexity ratchet;
  this keeps the ratchet at the pre-existing base count (2) instead of
  regressing it.
- Add a red-first-proven test locking the assembleStreamingResponseHeaders
  -> buildStreamingResponseHeaders wiring so the strip cannot go dead
  again silently.
- Fix the changelog.d fragment (it did not start with "- ", which fails
  check:changelog-integrity) and credit both @fenix007 (original design,
  diegosouzapw#13638, per the diegosouzapw#14116 mandate) and @shubhayu-dev.

Refs diegosouzapw#13638, diegosouzapw#14116

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: fenix007 <4217955+fenix007@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants