Skip to content

fix(combo): add cooldown cache for 429/503 targets and concurrent pre-screening - #3169

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.18from
pizzav-xyz:fix/combo-cooldown-prescreen
Jun 9, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.18from
pizzav-xyz:fix/combo-cooldown-prescreen

Conversation

@pizzav-xyz

Copy link
Copy Markdown
Contributor

Apologies for the PR spam — this is one of several small, focused PRs split from a larger batch to make review easier.

Summary

Adds an in-memory cooldown cache for combo targets that return 429/503 and introduces concurrent pre-screening to skip unavailable providers before dispatch.

Changes

  • Cooldown cache: When a target returns 429 or 503, it's cached as unavailable for 60s–300s (respects Retry-After header)
  • Retry-After parsing: Handles relative (30s), absolute (2026-01-01T...), and numeric (seconds) formats
  • Concurrent pre-screening: Before combo dispatch, pre-screens targets concurrently to skip unavailable providers early
  • Exports ProviderProfile type from accountFallback.ts

Files Changed

  • open-sse/services/combo.ts — cooldown cache, retry-after parsing, pre-screening
  • open-sse/services/comboConfig.ts — PRE_SCREEN_CONCURRENCY export
  • tests/unit/combo-provider-cooldown.test.ts — updated cooldown tests
  • tests/unit/combo-prescreen.test.ts — new pre-screen tests

@pizzav-xyz
pizzav-xyz requested a review from diegosouzapw as a code owner June 4, 2026 20:33

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces parallel pre-screening of targets for priority strategy combos to reduce latency, and implements an in-memory cooldown cache for targets returning 429/503 errors. It also updates and adds corresponding unit tests. Feedback on these changes includes a recommendation to reorder parsing checks in parseRetryAfterMs to prevent numeric strings from being incorrectly parsed as calendar years by Date.parse(), and a suggestion to make isCooldownExpired self-cleaning to avoid unbounded growth of the cooldown cache.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread open-sse/services/combo.ts Outdated
Comment on lines +109 to +117
const parsed = Date.parse(trimmed);
if (Number.isFinite(parsed)) {
const deltaMs = parsed - Date.now();
return deltaMs > 0 ? deltaMs : null;
}
const numeric = Number(trimmed);
if (Number.isFinite(numeric) && numeric >= 0) {
return numeric < 10_000_000_000 ? numeric * 1000 : numeric;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Pure numeric strings (e.g., '2026') representing a relative delay in seconds will be incorrectly parsed as calendar years (e.g., Year 2026) by Date.parse(), resulting in a massive, incorrect cooldown delta. Reordering the checks to parse pure numeric strings before attempting Date.parse() resolves this correctness issue.

Suggested change
const parsed = Date.parse(trimmed);
if (Number.isFinite(parsed)) {
const deltaMs = parsed - Date.now();
return deltaMs > 0 ? deltaMs : null;
}
const numeric = Number(trimmed);
if (Number.isFinite(numeric) && numeric >= 0) {
return numeric < 10_000_000_000 ? numeric * 1000 : numeric;
}
const numeric = Number(trimmed);
if (Number.isFinite(numeric) && numeric >= 0) {
return numeric < 10_000_000_000 ? numeric * 1000 : numeric;
}
const parsed = Date.parse(trimmed);
if (Number.isFinite(parsed)) {
const deltaMs = parsed - Date.now();
return deltaMs > 0 ? deltaMs : null;
}

Comment thread open-sse/services/combo.ts Outdated
Comment on lines +1475 to +1478
export function isCooldownExpired(executionKey: string): boolean {
const expiresAt = cooldownCache.get(executionKey);
return expiresAt ? Date.now() > expiresAt : true;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The cooldownCache map can grow indefinitely over time as various combo targets fail, since expired cooldown entries are never removed. We can make isCooldownExpired self-cleaning by deleting the expired entry from the map when it is checked.

Suggested change
export function isCooldownExpired(executionKey: string): boolean {
const expiresAt = cooldownCache.get(executionKey);
return expiresAt ? Date.now() > expiresAt : true;
}
export function isCooldownExpired(executionKey: string): boolean {
const expiresAt = cooldownCache.get(executionKey);
if (!expiresAt) return true;
if (Date.now() > expiresAt) {
cooldownCache.delete(executionKey);
return true;
}
return false;
}

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @pizzav-xyz! Heads up: release/v3.8.10 just merged #3156, which reworked the same same-provider fallback path in open-sse/services/combo.ts (now sequential + spaced for OAuth). This PR conflicts with that. Could you rebase onto the current release/v3.8.10 so the cooldown-cache change sits on top of the new sequencing? I want to make sure they compose rather than clobber each other. 🙏

@kilo-code-bot

kilo-code-bot Bot commented Jun 4, 2026

Copy link
Copy Markdown

Code Review Summary

Status: 3 Issues Found | Recommendation: Address before merge

Overview

Severity Count
CRITICAL 1
WARNING 2
Issue Details (click to expand)

CRITICAL

File Line Issue
open-sse/utils/proxyFallback.ts 267-280 Race condition: When one proxy succeeds in raceProxies, remaining pending probe requests are not cancelled. This wastes resources (outstanding undici connections) and could cause memory/connection leaks under high load. The AbortController in testSingleProxy only times out after 3s, but there's no mechanism to cancel in-flight requests when another proxy wins the race.

WARNING

File Line Issue
tests/unit/combo-provider-cooldown.test.ts 39 Test only verifies large value handling without testing the actual threshold boundary. parseRetryAfterMs(9999999999) (just below threshold) would be treated as seconds and multiplied by 1000, returning 9999999999000ms (115+ days), while parseRetryAfterMs(10000000000) would be treated as ms. Consider adding boundary tests.
tests/unit/combo-provider-cooldown.test.ts 43-46 Missing test for 'd' (days) unit format despite being implemented in parseRetryAfterMs. Test gap for parseRetryAfterMs("1d") → 86400000ms.
Other Observations (not in diff)

Issues found in unchanged code that cannot receive inline comments:

File Line Issue
open-sse/utils/proxyFallback.ts 267-280 The raceProxies function continues running all proxy probes even after finding a working one. Consider using an AbortController shared across all probes to cancel remaining requests when one succeeds.
open-sse/services/combo.ts 117-121 parseRetryAfterMs for ISO date strings returns null if the parsed date is in the past (deltaMs <= 0). Consider adding a test case for a past date to verify this behavior.
Files Reviewed (5 files)
  • open-sse/services/combo.ts - 1 critical issue (race condition in proxyFallback)
  • open-sse/services/comboConfig.ts - No issues
  • open-sse/utils/proxyFallback.ts - 1 critical race condition issue
  • src/lib/db/providerStats.ts - No issues (new file, clean implementation)
  • tests/unit/combo-provider-cooldown.test.ts - 2 warning level test gaps
  • tests/unit/combo-prescreen.test.ts - No issues

Comment thread open-sse/utils/proxyFallback.ts Outdated
// Find working proxy (with caching)
// ---------------------------------------------------------------------------

async function raceProxies(candidates: string[], targetUrl: string): Promise<string> {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CRITICAL: Race condition in raceProxies - When one proxy succeeds, remaining pending probe requests continue running without cancellation. This wastes resources (outstanding undici connections) and could cause memory/connection leaks. Consider using a shared AbortController to cancel remaining requests when one succeeds.

if (!targetHostname) return null;

// Check cache first
const cached = PROXY_FALLBACK_CACHE.get(targetHostname);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

WARNING: Cache key uses hostname but cache lookup happens before candidate collection. If getProxyCandidates throws, the cache entry from a previous successful lookup won't be used. Consider checking cache first before calling getProxyCandidates.

return buildClaudeResponse("claude fallback handled it");
}
test("parseRetryAfterMs handles numeric milliseconds (large values)", () => {
assert.equal(parseRetryAfterMs(1700000000000), 1700000000000);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

WARNING: Test verifies large value handling but doesn't test the threshold boundary. Consider adding a test for parseRetryAfterMs(9999999999) which would be interpreted as seconds (multiplying by 1000) vs parseRetryAfterMs(10000000000) which is treated as milliseconds. This is important because values near the boundary have different interpretations.

throw new Error(`unexpected upstream headers: ${JSON.stringify(headers)}`);
};
test("parseRetryAfterMs handles string duration formats", () => {
assert.equal(parseRetryAfterMs("30s"), 30_000);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

WARNING: Missing test coverage for 'd' (days) unit format. The parseRetryAfterMs function handles d in multipliers but there's no test case like parseRetryAfterMs("1d") → 86400000ms.


import { fetch as undiciFetch } from "undici";
import { createProxyDispatcher, normalizeProxyUrl } from "./proxyDispatcher.ts";
import { resolveProxyForScopeFromRegistry, listProxies, listOneproxyProxies } from "@/lib/localDb";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CRITICAL: resolveProxyForScopeFromRegistry is imported from @/lib/localDb but is NOT exported. The localDb.ts re-exports resolveProxyForConnectionFromRegistry and resolveProxyForProvider but NOT resolveProxyForScopeFromRegistry. This will cause a runtime/build error. Either add the export to localDb.ts or import directly from the proxies module.

@kilo-code-bot

kilo-code-bot Bot commented Jun 4, 2026 •

Copy link
Copy Markdown

Code Review Summary

Status: 3 Issues Found | Recommendation: Address before merge

Overview

Severity Count
CRITICAL 2
WARNING 1
Issue Details (click to expand)

CRITICAL

File Line Issue
open-sse/utils/proxyFallback.ts 12 resolveProxyForScopeFromRegistry is imported from @/lib/localDb but is NOT exported by that module. The localDb.ts re-exports resolveProxyForConnectionFromRegistry and resolveProxyForProvider but NOT resolveProxyForScopeFromRegistry. This will cause a runtime/build error. Either add the export to localDb.ts or import directly from ../db/proxies.
open-sse/utils/proxyFallback.ts 267-280 Race condition in raceProxies: When one proxy succeeds, remaining pending probe requests are NOT cancelled. This wastes resources (outstanding undici connections) and could cause memory/connection leaks under high load. The AbortController in testSingleProxy only times out after 3s, but there's no mechanism to cancel in-flight requests when another proxy wins the race.

WARNING

File Line Issue
tests/unit/combo-provider-cooldown.test.ts 39 Test verifies large value handling but doesn't test the threshold boundary. Consider adding a test for parseRetryAfterMs(9999999999) (treated as seconds) vs parseRetryAfterMs(10000000000) (treated as ms).
Other Observations (not in diff)

Issues found in unchanged code that cannot receive inline comments:

File Line Issue
tests/unit/combo-provider-cooldown.test.ts 43-46 Missing test for 'd' (days) unit format despite being implemented in parseRetryAfterMs. Test gap for parseRetryAfterMs("1d") → 86400000ms.
open-sse/services/combo.ts 117-121 parseRetryAfterMs for ISO date strings returns null if the parsed date is in the past. Consider adding a test case for a past date to verify this behavior.
Files Reviewed (6 files)
  • open-sse/services/combo.ts - No issues in changed lines
  • open-sse/services/comboConfig.ts - No issues
  • open-sse/utils/proxyFallback.ts - 2 critical issues (missing export + race condition)
  • src/lib/db/providerStats.ts - No issues (new file, clean implementation)
  • tests/unit/combo-provider-cooldown.test.ts - 1 warning (test gaps)
  • tests/unit/combo-prescreen.test.ts - No issues

Reviewed by laguna-m.1-20260312:free · 3,528,487 tokens

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for the work here @pizzav-xyz — combo resilience is valuable, but I hit two blockers during review that I'd like to resolve with you before this can land.

1. It removes existing behavioral coverage (the main blocker). This PR rewrites tests/unit/combo-provider-cooldown.test.ts, deleting the end-to-end test combo failover skips the cooled provider target on the next request (which drives the real chat pipeline via createChatPipelineHarness and asserts the second request skips the cooled target) and replacing it with isolated unit tests of the new helpers (setCooldown / isCooldownExpired / parseRetryAfterMs). The original test still passes on release/v3.8.11, so we'd be trading strong behavioral coverage for primitive-level coverage. Please keep the existing test intact.

2. The cooldown cache largely duplicates the existing connection-cooldown. The repo already cools a connection down on 429/503 (via markAccountUnavailable / rateLimitedUntil) and pre-screens unavailable targets before dispatch — that's exactly what the deleted test proves. Adding a second, parallel in-memory cache (setCooldown/isCooldownExpired keyed by target in combo.ts) means two cooldown systems to keep in sync. In my TDD repro the existing mechanism already short-circuits a same-provider cascade after the first failure (see #3194).

The genuinely-new piece that could be worth keeping is parseRetryAfterMs (relative/absolute/ISO Retry-After parsing) if it improves on the existing Retry-After handling for API-key 429s. Would you be up for: (a) restoring the original behavioral test, and (b) dropping the duplicate cache so the PR narrows to just the Retry-After parsing (or whatever the existing cooldown genuinely lacks)? Also note this branch is currently CONFLICTING against the release branch. Leaving it open for your revision — thanks again.

@diegosouzapw
diegosouzapw changed the base branch from main to release/v3.8.11 June 5, 2026 11:55
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @pizzav-xyz! The pre-screening idea (skip known-unavailable targets before dispatch) is appealing. Two things blocking a merge right now:

  1. Overlap with the existing connection cooldown. This adds a 4th in-memory cooldownCache (keyed by executionKey) on top of the three documented resilience layers (provider circuit breaker, per-account connection cooldown via rateLimitedUntil, model lockout — see docs/architecture/RESILIENCE_GUIDE.md). The per-account cooldown already skips 429/503 targets and is proven to short-circuit the same-provider cascade (test(combo): guard same-provider cascade is handled by connection cooldown (#3200) #3194). Four independent "unavailable" stores risk divergence. Could you explain how the combo cache coordinates with rateLimitedUntil so they can't disagree — or whether the pre-screen can read the existing cooldown state instead of maintaining its own?

  2. Stale base. The diff re-creates src/lib/db/providerStats.ts, which already shipped in v3.8.10 (feat(dashboard): provider stats API endpoint and dashboard page #3175). The branch predates it — a rebase onto current release/v3.8.11 is needed (and will drop that duplicate).

Leaving open pending a rebase + the coordination note. Happy to merge the pre-screen optimization once it reuses the existing cooldown state rather than adding a parallel one. 🙏

@diegosouzapw

Copy link
Copy Markdown
Owner

Detailed guidance to land this (keeping it open)

@pizzav-xyz — I want this to merge; the pre-screening idea (skip known-unavailable targets before dispatch) is a genuine latency win. Two concrete things to adjust so it composes with what's already in the tree:

1. Reuse the existing cooldown state instead of a 4th parallel store

Right now this adds an in-memory cooldownCache (keyed by executionKey). The codebase already tracks per-connection unavailability authoritatively — please read from that rather than maintaining a separate map that can drift:

  • Per-account cooldown: rateLimitedUntil / testStatus / backoffLevel on the provider connection. Written by src/sse/services/auth.ts::markAccountUnavailable(), cleared by clearAccountError(). A connection is "unavailable" while new Date(rateLimitedUntil).getTime() > Date.now().
  • Provider circuit breaker: src/shared/utils/circuitBreaker.ts — use getStatus() / canExecute() (they lazily refresh OPEN→HALF_OPEN, so don't read raw state).
  • Model lockout: open-sse/services/accountFallback.ts.

So pre-screen should ask those three sources "is this target currently serveable?" — not record its own 429/503 cooldown. That removes the divergence risk entirely (one source of truth) and the setCooldown/isCooldownExpired map can largely go away. See docs/architecture/RESILIENCE_GUIDE.md for the 3-layer model.

2. Rebase + drop the stale providerStats.ts

The branch predates v3.8.10, so it re-creates src/lib/db/providerStats.ts (shipped in #3175). A rebase onto current release/v3.8.11 will drop that duplicate and surface the real conflict surface in combo.ts.

Tests

Keep combo-prescreen.test.ts, but please add one asserting the pre-screen honors an existing rateLimitedUntil set via markAccountUnavailable (i.e. it reads the canonical state, not its own cache). That's the behavior that proves #1.

Once it's rebased + reads the canonical cooldown state, ping me and I'll review/merge. Thanks for pushing on combo resilience! 🙏

@pizzav-xyz
pizzav-xyz force-pushed the fix/combo-cooldown-prescreen branch from 1447ea3 to 406fd5f Compare June 5, 2026 17:15
@pizzav-xyz

Copy link
Copy Markdown
Contributor Author

Thanks for the detailed guidance @diegosouzapw! Working on it now:

  1. Dropping the duplicate cooldownCache — rewiring preScreenTargets to read from the canonical 3-layer state (rateLimitedUntil via isAccountUnavailable, circuit breaker via getCircuitBreaker().getStatus(), model lockout via isModelLocked)
  2. Keeping parseRetryAfterMs — the Retry-After parsing is genuinely new
  3. Keeping the original E2E test — plus adding a test that proves pre-screen honors rateLimitedUntil
  4. Branch is rebased onto current main (was behind providerStats.ts which already shipped)

Will push the refactored branch shortly.

@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.11 to release/v3.8.12 June 5, 2026 23:45
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @pizzav-xyz — the rewire to canonical state is exactly what we discussed (dropped the parallel cooldownCache, reads rateLimitedUntil via isAccountUnavailable + getCircuitBreaker().getStatus(), removed the stale providerStats.ts). 🙌

I pushed one commit to your branch (you as author, me as co-author) with two mechanical fixes so it builds and runs:

  1. Syntax error (blocking): the three pre-screen skip branches you added inside the executeTarget arrow function used continue, which esbuild rejects (Cannot use "continue" here). Switched them to return null to match the existing skip branch (exhaustedProviders).
  2. Test setup: combo-provider-cooldown.test.ts was missing the const { ... } = harness destructure, so resetStorage/seedConnection/etc. were undefined. Restored it.

With those, all 8 of your tests pass on your branch's base. ✅

One blocker remains before this can merge — when your branch is rebased onto the current release/v3.8.12, it regresses the #3200 same-provider cascade guard (tests/unit/combo-same-provider-cascade.test.ts, added after your branch point):

✖ combo hits a failing provider only once before falling back across same-provider targets (#3200)
  expected the failing provider to be hit once then short-circuited by connection cooldown, got 3 calls
  3 !== 1

Root cause looks like the pre-screen reads the stale connection snapshot (target as any).connection.rateLimitedUntil captured at resolve time — so a connection that gets cooled mid-request (after the first same-provider target fails) isn't seen by the pre-screen, and the remaining same-provider targets still execute instead of being short-circuited.

To land it:

  1. git fetch && git rebase origin/release/v3.8.12 (picks up the [BUG] playground chat model list cannot fetch properly #3200 test).
  2. Make the pre-screen consult live cooldown state at dispatch (re-read the connection, or reuse the same live check the [BUG] playground chat model list cannot fetch properly #3200 short-circuit relies on) rather than the resolve-time snapshot.
  3. Confirm both combo-provider-cooldown.test.ts and combo-same-provider-cascade.test.ts are green.

Happy to pair on step 2 if useful. Leaving it open — you're close! 🚀

@pizzav-xyz
pizzav-xyz force-pushed the fix/combo-cooldown-prescreen branch from 8ef90f5 to 759f2e5 Compare June 6, 2026 14:38
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.12 to release/v3.8.13 June 6, 2026 15:04
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @pizzav-xyz for this, and apologies for the slow turnaround. I did a deep review and there are a few blockers that need your attention before it can land — leaving the PR open so you keep authorship:

1. Build break — broken import. open-sse/services/combo.ts now imports formatRetryAfterMs from accountFallback.ts, but that module only exports formatRetryAfter (no ...Ms variant). This makes the combo module throw has no exported member 'formatRetryAfterMs' on load — every combo request would crash. Either restore formatRetryAfter in the import list, or add the ...Ms export to accountFallback.ts first.

2. The connection-cooldown check is dead code as integrated. preScreenTargets reads (target as any).connection?.rateLimitedUntil, but the runtime target type ResolvedComboTarget (combo.ts:~337) has no .connection field — only connectionId: string | null. So in the real pipeline (resolveComboTargets → normalizeRuntimeStep) that branch is always undefined and never runs. To make it real you'd need to extend ResolvedComboTarget with an optional connection snapshot and populate it during target resolution. Otherwise please drop the rateLimitedUntil branch and rely on the circuit-breaker check (which works, since it uses target.provider).

3. Tests don't exercise the integrated path. combo-prescreen.test.ts passes targets with no real connection objects, so it only proves the existing isModelAvailable parallelization — the new cooldown branch is never hit. combo-provider-cooldown.test.ts tests preScreenTargets in isolation with hand-built .connection objects that don't exist in production. We'd need an integration test (seed a DB connection with rateLimitedUntil in the future, route a real combo, assert it's skipped) to prove the behavior end-to-end (Hard Rule #18).

Also: clampCooldownMs and parseRetryAfterMs are added but never wired into any production path, and parseRetryAfterMs overlaps the existing parsers in accountFallback.ts (parseDelayString/parseRetryAfterFromBody). Consider consolidating rather than adding a fourth parser.

Net: the genuinely new value here is the circuit-breaker short-circuit for the priority strategy — that part is worth keeping. If you fix the import, resolve the dead-code path (or remove it), and add one real integration test, I'll happily re-review and merge. Thank you!

@diegosouzapw

Copy link
Copy Markdown
Owner

Following up @pizzav-xyz — this still has the two blockers from my earlier review (the branch hasn't been updated since). Recapping so it's actionable:

  1. Build break: open-sse/services/combo.ts imports formatRetryAfterMs from accountFallback.ts, but that module only exports formatRetryAfter — the combo module throws on load. Restore formatRetryAfter or add the ...Ms export first.
  2. Dead code: preScreenTargets reads (target as any).connection?.rateLimitedUntil, but ResolvedComboTarget has no .connection field (only connectionId), so that branch never runs in the real pipeline. Either extend the target type + populate it during resolution, or drop the branch and rely on the circuit-breaker check.

It also currently conflicts with release/v3.8.13. Once the import is fixed, the dead-code path is resolved, and there's one integration test that exercises the real combo + cooldown interaction (Hard Rule #18), I'll re-review and merge. Thanks!

@pizzav-xyz
pizzav-xyz force-pushed the fix/combo-cooldown-prescreen branch from 759f2e5 to 777c60c Compare June 7, 2026 00:36
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.13 to release/v3.8.14 June 7, 2026 01:07
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.14 to release/v3.8.15 June 7, 2026 12:40
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.15 to release/v3.8.16 June 7, 2026 17:29
@pizzav-xyz
pizzav-xyz force-pushed the fix/combo-cooldown-prescreen branch from 777c60c to a6f4089 Compare June 7, 2026 19:58
- Add preScreenTargets() that checks provider profiles and model
  availability concurrently before dispatch
- Use circuit breaker status in pre-screen to skip OPEN providers
- Cache pre-screen results in executeTarget to avoid redundant checks
- Export PRE_SCREEN_CONCURRENCY (5) from comboConfig
- Add integration tests for pre-screening behavior
- Retains all existing test coverage

Addresses review feedback: removes dead code, broken imports, and
unused parsers from previous iteration. Rebased onto release/v3.8.13.
@pizzav-xyz
pizzav-xyz force-pushed the fix/combo-cooldown-prescreen branch from a6f4089 to 95114cc Compare June 8, 2026 14:42
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.16 to release/v3.8.17 June 8, 2026 21:02
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.17 to release/v3.8.18 June 9, 2026 11:32
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @pizzav-xyz! Verified the behavior end-to-end: preScreenTargets fans out provider-profile + availability checks at concurrency 5 for priority combos, and the attempt loop now fast-exits on an OPEN circuit breaker and reuses the pre-screened profile — your combo-prescreen (5) and combo-provider-cooldown (2) tests pass, including "skips the cooled provider target on the next request" and "pre-screen marks target unavailable when circuit breaker is OPEN". One note: there's no new cache here (the code correctly re-calls isModelAvailable because connection cooldowns can change mid-request), so I'm landing it under a clarified title — "parallel pre-screen + circuit-breaker fast-exit" — to match what it actually does. Merging into release/v3.8.18; ships next release. 🙌

@diegosouzapw
diegosouzapw merged commit c77215f into diegosouzapw:release/v3.8.18 Jun 9, 2026
1 of 2 checks passed
diegosouzapw added a commit that referenced this pull request Jun 9, 2026
@diegosouzapw diegosouzapw mentioned this pull request Jun 9, 2026
diegosouzapw added a commit that referenced this pull request Jun 9, 2026
- getPendingRequests() typed to real shape (was widened to object) → fixes
  unknown 'count' in the unified-requests view (#3401)
- streamChunks log payload cast to its declared type (callLogs.ts)
- preScreenTargets aligned to canonical IsModelAvailable signature (#3169),
  Promise.resolve-normalized so .catch never hits a bare boolean

All 5 gates green: lint(0 err) + typecheck:core + cycles + docs-all + unit + vitest(146).
diegosouzapw added a commit that referenced this pull request Jun 9, 2026
* chore(release): open v3.8.18 development cycle

* fix(catalog): stop Codex CLI model-catalog refresh from erroring (#3481)

Codex's model-catalog refresh (codex_models_manager) does
GET /v1/models?client_version=<v> and decodes a JSON object with a
TOP-LEVEL `models` array. OmniRoute answers in the OpenAI-standard
`{object,data}` shape, so codex fails with "missing field `models`"
and logs "failed to refresh available models" on every startup.

Detect codex clients via the `originator` / `user-agent` = `codex_*`
headers they send and add an EMPTY top-level `models: []` so the decode
succeeds. Non-codex OpenAI clients keep the byte-identical `{object,data}`
response.

The array is intentionally empty: codex replaces its built-in per-model
agent prompt (`base_instructions`, ~21k chars) with whatever a populated
entry carries for the selected model, so emitting our catalog would drop
the agent prompt to nothing and break codex's agent behaviour (verified
empirically against codex 0.137). An empty list keeps codex on its
built-in model info — same inference as before, minus the error.

Validated end-to-end with the real handler against codex 0.137:
"failed to refresh available models" → 0 occurrences, instructions
preserved (built-in Codex agent prompt, not empty).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: ignore quality reports and local prompt artifacts

Add generated quality gate reports, metrics files, and local setup prompt
artifacts to .gitignore to prevent committing environment-specific or
temporary files.

* fix(provider): detect Responses API format when body has `input` but … (#3490)

Integrated into release/v3.8.18

* fix(sse): normalize numeric provider ids to strings (#3451)

Integrated into release/v3.8.18

* feat(browserPool): resolve Playwright proxy from proxy_registry DB (#3492)

Integrated into release/v3.8.18

* fix(theoldllm): generate X-Request-Token server-side, drop Playwright (#3491)

Integrated into release/v3.8.18

* feat(plugins): add lifecycle hooks and theme-manager plugin (#3473)

Integrated into release/v3.8.18

* fix(combo): parallel pre-screen + circuit-breaker fast-exit for priority combos (#3169)

Integrated into release/v3.8.18

* feat(ui): unifi active and finished requests into single view #1422 (#3401)

Integrated into release/v3.8.18

* docs(changelog): record #3401, #3473, #3492, #3490, #3451, #3491, #3169 under v3.8.18

* feat(docs): add doc accuracy gate + refresh AGENTS.md counts (#3510)

Integrated into release/v3.8.18

* fix(sse): drop empty-choices chunks without usage instead of injecting retry text (#3513)

PR #3422 ('allow OpenAI usage-only empty choices chunks') reintroduced the
assistant-content injection '[OmniRoute] Upstream returned an empty response.
Please retry.' for empty `choices: []` chunks that carry no valid usage. Clients
(Goose/opencode) feed that text back as a turn and spin in a retry loop -- the
exact regression #3400 had fixed by dropping the chunk.

Restore the drop behavior for the no-usage case while preserving #3422's
standards-compliant forwarding of usage-only `include_usage` final chunks.
Realign the mislabeled stream-utils test (it asserted the injection) and add a
dedicated regression guard.

Reported-by: @mochizzan
Refs: #3502, #3388, #3400, #3422

* fix(authz): fall back to URL token when Authorization isn't a usable Bearer (#3504)

Integrated into release/v3.8.18

* fix(playground): authenticate via session, test key policy by id (#3503)

Integrated into release/v3.8.18

* docs(changelog): record #3510, #3504, #3503 under v3.8.18

* fix: llama base url normalization (#3519)

* docs(changelog): reconcile v3.8.18 — add #3519, #3513, #3435-repair, gitignore chore (full commit↔changelog coverage)

* fix(opencode-plugin): bound regex quantifiers in normaliseFreeLabel (polynomial-ReDoS)

CodeQL js/polynomial-redos: unbounded \s* before an anchored \s*$ allowed
O(n²) backtracking on attacker-influenced display names. Bounded to {0,8}/{1,8}
(ample for any real label spacing). Plugin builds + 254 tests green.

* fix(types): restore clean typecheck:core for v3.8.18 release gate

- getPendingRequests() typed to real shape (was widened to object) → fixes
  unknown 'count' in the unified-requests view (#3401)
- streamChunks log payload cast to its declared type (callLogs.ts)
- preScreenTargets aligned to canonical IsModelAvailable signature (#3169),
  Promise.resolve-normalized so .catch never hits a bare boolean

All 5 gates green: lint(0 err) + typecheck:core + cycles + docs-all + unit + vitest(146).

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Andrey Borodulin <borodulin@gmail.com>
Co-authored-by: Dmitrii Safronov <zimniy@cyberbrain.cc>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
@kilo-code-bot

kilo-code-bot Bot commented Jun 9, 2026

Copy link
Copy Markdown

Kilo Code Review could not run — your account is out of credits.

Add credits or switch to a free model to enable reviews on this change.

smartenok-ops added a commit to smartenok-ops/OmniRoute that referenced this pull request Jul 24, 2026
* fix(antigravity): alias agy gemini-3.1-pro -high/-low + stop masking upstream 4xx (#3229) (#3245)

agy's gemini-3.1-pro-high/-low had no alias, so resolveAntigravityModelId sent the
speculative -high/-low suffix verbatim to upstream, which rejects it (400) for
gemini-3.x. Worse, the non-stream executor branch fed the 4xx response into the SSE
collector, returning a synthetic empty {object:chat.completion} envelope that masked
the error. Alias both to gemini-3.1-pro, and surface real upstream errors via
buildErrorBody for non-ok non-stream responses. + unit tests.

* docs(changelog): combo on /v1/responses (#3227/#3233) + agy gemini 400 (#3229) (#3246)

* chore(release): finalize v3.8.11 changelog + repair release-gate test drift

Finalize the 3.8.11 cycle CHANGELOG and clear the failures the full test:unit
gate surfaced (release-branch drift — PR merges bypassed pre-push):

- CHANGELOG: date the [3.8.11] section (2026-06-05) + repo-housekeeping roll-up
- docs(env): document THEOLDLLM_NAV_TIMEOUT_MS in .env.example + ENVIRONMENT.md
  (env-doc-sync gate; #3217 added the var without docs)
- test(nvidia): exercise the #3226 bypass-fetch path via a local HTTP server and
  hoist the validator import (patching globalThis.fetch no longer intercepts the
  un-patched native fetch captured by proxyFetch)
- test(i18n): import the shipped normalizeComplianceEventTypes helper (#3185)
- test(model-caps): save synced metadata under the canonical gemini-3.1-pro key
  now that #3229 aliases gemini-3.1-pro-high/-low to it
- test(web-session): expect the grok-web "sso + sso-rw" credential hint (#3180)
- test(synced-models): isolate DATA_DIR so the #3199 hidden-override stops
  bleeding into the shared DB and breaking the re-run precondition

* fix(api,dashboard): validate /v1/images/edits JSON body + drop duplicate proxy handler

Clear the two release-gate failures the CI Lint+Build jobs surfaced (release-branch
drift — PR merges bypassed pre-push):

- fix(api): /v1/images/edits parsed request.json() without a Zod guard
  (route-validation t06 / hard rule #7). Add ImageEditJsonSchema.safeParse so a
  malformed body (non-object / wrong types) is rejected with 400 instead of
  silently parsed; valid JSON/data-URL bodies behave exactly as before. (#3214, #3215)
- fix(dashboard): remove a duplicate handleToggleProxyEnabled /
  handleTogglePerKeyProxyEnabled / handleDistributeProxies block in
  providers/[id]/page.tsx — a bad merge of the proxy PRs declared all three twice,
  breaking the webpack build ("Identifier already declared"). The removed copy was
  byte-identical to the kept one. (#3170, #3171, #3172)

* fix(security): clear CodeQL high alerts surfaced on the v3.8.11 release PR diff

CodeQL flagged 7 high alerts in code the cycle touched (the large release-PR diff
re-surfaces them). Resolved at the source — no dismissals:

- fix(images): resolveImageBaseUrl trimmed trailing slashes with `/\/+$/`, a
  polynomial-ReDoS pattern (js/polynomial-redos) on the configured node base URL.
  Replace it with a non-backtracking endsWith/slice loop.
- test(oauth): pin the Anthropic OAuth host with exact-equality asserts and a
  parsed-hostname negative check instead of substring `.includes()`
  (js/incomplete-url-substring-sanitization). The exact-equality assertions were
  already present, so coverage is unchanged.
- test(images): drop the redundant `!includes("generativelanguage.googleapis.com")`
  assert — the exact-equality assert on the resolved URL already guarantees it.

* chore(release): open v3.8.12 development cycle

Bump 3.8.11 → 3.8.12 across package.json, lockfile, electron/, open-sse/, and
docs/reference/openapi.yaml; add the [3.8.12] cycle placeholder to the root
CHANGELOG and the 41 i18n mirrors. Integration branch for the v3.8.12 cycle —
fixes/features land here via per-issue PRs and it merges to main at release time.

* docs(changelog): credit @wilsonicdev for the /v1/responses combo fix (#3242)

* fix(sse): strip every <omniModel> tag, not just the first (#454) (#3248)

Strip ALL <omniModel> tags before forwarding to provider (global regex variant).

Integrated into release/v3.8.12. Thanks @MikeTuev.

* fix(grok-web): add TLS fingerprint impersonation to bypass Cloudflare anti-bot (#3180) (#3249)

grok-web: TLS fingerprint impersonation to bypass Cloudflare anti-bot (#3180); sanitize executor error bodies (#12).

Integrated into release/v3.8.12. Thanks @wilsonicdev.

* feat(web-cookie): add tool-call translation to 8 executors via shared webTools helpers (#3259)

Add tool-call translation to 8 web-cookie executors via shared webTools helpers.

Integrated into release/v3.8.12. Thanks @oyi77.

* fix(providers): improve refresh validation and model catalog UI (#3261)

Provider refresh/validation, OpenRouter catalog and proxy UI fixes — incl. NVIDIA NIM /models-suffix path fix (real-VPS validated).

Integrated into release/v3.8.12. Thanks @strangersp.

* feat(provider): add Chipotle Pepper AI — free provider via reverse-engineered Amelia protocol (#3250)

Add Chipotle Pepper AI free provider (Amelia protocol); sanitize executor error body (#12).

Integrated into release/v3.8.12. Thanks @oyi77.

* fix(v1/responses): skip codex rewrite for combo names (#3233, #3227) (#3268)

Regression test for the /v1/responses combo-name codex-rewrite guard (#3233, #3227).

Integrated into release/v3.8.12. Thanks @wilsonicdev.

* fix(embeddings): block cross-dimension failover in embedding combos (#3256)

Block cross-dimension failover in embedding combos.

Integrated into release/v3.8.12.

* fix(ci): deploy-vps recreates PM2 via bin + gates on /login 200 (#3270)

Synchronized deploy-vps hardening (PM2 recreate via bin + /api/monitoring/health gate + fail-on-unhealthy). Supersedes #3262.

Integrated into release/v3.8.12.

* fix(db): detect SQLite driver-unavailable errors to avoid destructive rename (#3274)

Detect SQLite driver-unavailable errors to avoid destructive DB rename + optional FTS5 migration guard (split from #3073).

Integrated into release/v3.8.12. Thanks @zhiru.

* feat(dashboard): bulk activate/deactivate/retest for selected provider connections (#3271)

Bulk activate/deactivate/retest for selected provider connections.

Integrated into release/v3.8.12. Thanks @leninejunior.

* feat(free-tiers): per-model free-token budget + Monthly Budget dashboard card (#3263)

Free-token budget catalog + per-model budget + Monthly Budget dashboard card (joins #3257 + #3263 into one).

Integrated into release/v3.8.12.

* fix(sse): parse <tool_call name=...> wrapper from web-cookie providers (#3260) (#3275)

ds-web/deepseek-v4-pro emits tool calls wrapped as
<tool_call name="skill">{"name":"customize-opencode"}</tool_call> instead of the
canonical <tool>{json}</tool>. webTools.ts only matched <tool>...</tool>, so the block
was silently dropped (and when arguments were present, the surrounding tag leaked into
content). Add TOOL_CALL_TAG_RE to capture the JSON body — the real tool name comes from
the body, never the tag's name= attribute — and extend the early-exit + range stripping.

Regression test: tests/unit/web-tools-translation-3260.test.ts (RED before, GREEN after).
Existing web-tools suites stay green (26/26).

* fix(sse): strip reasoning_effort for non-reasoning Groq models (#3258) (#3277)

Regression of #764. Claude Code → Groq (llama-3.3-70b-versatile) returned HTTP 400
because the model was treated as reasoning-capable: supportsReasoning() defaulted to true,
so applyThinkingBudget did not strip reasoning params, and the claude→openai translator
forwarded reasoning_effort (and re-injected it from output_config.effort) — which Groq
rejects on non-reasoning models.

- providerRegistry: mark llama-3.3-70b-versatile + llama-4-scout supportsReasoning:false
  (gpt-oss / qwen3-32b keep reasoning — they accept reasoning_effort).
- stripThinkingConfig: also strip output_config.effort so the translator can't re-inject
  reasoning_effort downstream.

Regression test: tests/unit/thinking-budget-groq-3258.test.ts (RED before, GREEN after);
existing thinking-budget suites stay green (45/45).

* fix(api): build /v1/images/edits multipart as Buffer, not global FormData (#3273) (#3278)

A custom OpenAI-compatible image-edit provider received an empty `model`. In production
`globalThis.fetch` is patched with node_modules/undici's fetch, whose `FormData` class
differs from `globalThis.FormData`; passing a native FormData made undici serialize it as
the string "[object FormData]" (text/plain), dropping every field including `model`.

handleOpenAIImageEdit now assembles the multipart body as a Buffer with an explicit
boundary + Content-Type, which every fetch impl accepts verbatim.

Regression test: tests/unit/image-edits-multipart-3273.test.ts reproduces the exact prod
condition (routes through undici's fetch) — RED before (upstream got text/plain
[object FormData]), GREEN after. Existing image suites stay green (50/50).

* docs(changelog): complete v3.8.12 audit — all 14 merged PRs + contributors hall

Audited every commit since v3.8.11 one-by-one. Added the missing v3.8.12
entries (features #3250/#3259/#3263/#3271, fixes #3248/#3249/#3261/#3256/#3274,
maintenance #3270), repointed the combo-rewrite and web-tools entries to their
actually-merged PRs (#3268, #3275) instead of the closed #3242/issue links, and
added the v3.8.12 Contributors hall. Also co-credited @ibanunmangun on the
v3.8.11 #3203 OAuth fix (independent first diagnosis via #3193).

* fix(api): allow private webhook targets behind explicit opt-in (#3269) (#3279)

Webhooks hardcoded parseAndValidatePublicUrl, which blocks any RFC1918/loopback host —
breaking self-hosted setups that legitimately point webhooks at internal services
(n8n, Home Assistant, a LAN box). Provider URLs already had an opt-in
(OMNIROUTE_ALLOW_PRIVATE_PROVIDER_URLS); webhooks now reuse it.

- outboundUrlGuard: add parseAndValidateWebhookUrl — gates the private-host check on
  arePrivateProviderUrlsAllowed() (default OFF); protocol + embedded-credential checks
  stay unconditional.
- swap all webhook call sites (create/update/test/validate-url + dispatcher x2) to it.

Regression test: tests/unit/webhook-private-optin-3269.test.ts (RED before, GREEN after);
existing webhook SSRF/dispatcher suites stay green (33/33).

* fix(api): harden private webhook opt-in against cloud-metadata SSRF (#3269) (#3281)

Follow-up to the #3269 private-webhook opt-in. With the opt-in on, the private-host
check was bypassed entirely, leaving cloud-metadata endpoints (169.254.169.254,
metadata.google.internal, 100.100.100.200, link-local 169.254.0.0/16) reachable — the
classic SSRF -> IAM-credential pivot — and the webhook test endpoint returned the
upstream body, making it a content-exfiltration primitive against internal services.

- outboundUrlGuard: add isCloudMetadataHost(); parseAndValidateWebhookUrl blocks those
  hosts UNCONDITIONALLY, even when private targets are opted in.
- webhooks/[id]/test: redact responseBody for private targets (status + latency only).

Regression test: tests/unit/webhook-metadata-guard-3269.test.ts (RED before, GREEN after);
existing webhook SSRF/opt-in suites stay green (34/34).

* fix(sse): don't mark a valid Qoder PAT expired on a generic Cosy 500 (#3247) (#3283)

A working Qoder PAT was reported as "expired". The validator probes the Cosy endpoint
(api1.qoder.sh) — which IS the correct PAT path (the executor falls back to it after the
expected 401 from api.qoder.com). The bug was the verdict: isCosyAppError (added by #2860)
marked ANY Cosy 500 with "success":false as an auth failure, including a generic
{..."msgCode":500,"message":"Internal Server Error"} server fault — contradicting the
older #1391 "5xx = valid bypass" rule.

Narrow it: a Cosy 500 only marks the PAT invalid when the body carries an EXPLICIT auth
signal (unauthorized/forbidden/expired/token invalid/...); a generic Internal Server Error
falls back to valid-bypass. #2860's protection for genuine auth rejections is preserved.

Regression test: tests/unit/qoder-cli.test.ts — the two pre-existing generic-500 cases now
assert valid:true (they encoded the #3247 bug) + a new explicit-auth-signal case asserts
valid:false. 13/13 green.

* docs(free-tiers): richer budget-card image (28 models + first-month strip) + soften ToS framing to caution (#3284)

- Regenerate the README/dashboard mockup from the catalog: 28 pools in the grid
  (was 9), a balance-floored stacked bar (Mistral now ~40% of the bar, was ~90%),
  and a first-month signup-credit strip (~586M). Add the data-driven generator.
- FREE_TIERS.md: drop the alarming '🚫 Avoid / terms prohibit' framing — relabel
  those 19 providers as 'caution — worth checking', note their access is real and
  the OAuth/keyless ones aren't token-quantifiable (so out of the headline, not
  excluded as unusable).

* fix(quota): resolve poolUsage dead code, burn rate, saturation signals, webhooks, and embeddings enforcement (#3280)

Integrated into release/v3.8.12. Quota Sharing Engine fixes: poolUsageWithDimensions promoted to the QuotaStore interface, single-snapshot burn rate, zero-weight normalization, Anthropic saturation signals, quota.exceeded webhook on block, and embeddings enforcement. Validated: 10/10 PR tests + 34 quota/embedding regression files green, typecheck + lint clean. Dropped the committed .omo/ agent-tooling artifacts.

* docs(changelog): credit @wilsonicdev for the Qoder 500-bypass diagnosis (#3282/#3247)

The #3247 fix shipped via #3283 (parallel session) 46s after @wilsonicdev
filed the same fix in #3282, leaving his PR stranded with no credit — the
#3242 credit-theft pattern. Repoint the entry to the merged #3283, credit
@wilsonicdev as co-author for the independent diagnosis, and note #3283
refined it to keep rejecting on an explicit-auth-signal 500.

* fix(security): use crypto.randomInt/randomUUID in chipotle + URL parser in test (#3285)

Integrated into release/v3.8.12. CodeQL hardening on the Chipotle executor: Math.random → crypto.randomInt/randomUUID, and a strict URL hostname check in the test. Fixed the node:crypto import (crypto.randomInt is not on the Web Crypto global → would crash at WS-connect) and added a regression guard exercising both helpers.

* feat(models): add MiniMax M3 across all provider tiers (#3110) (#3287)

Integrated into release/v3.8.12. Registers MiniMax-M3 (1M context, Anthropic-compatible) across 8 provider tiers (minimax, minimax-cn, opencode, opencode-go, opencode-zen, trae, ollama-cloud, nvidia). Validated: 8/8 new registry tests + 25 registry/model-catalog regression files green, typecheck + lint clean. Complements the #3141 max_tokens spec already on release.

* docs(readme): consolidate community at top (Discord + Telegram + WhatsApp) + promote Free-Token Budget section (#3289)

- Add the official Telegram group (t.me/omnirouteOficial) and gather Discord,
  Telegram and both WhatsApp groups into one community card block at the top;
  remove the scattered WhatsApp links from the nav line and the Support section
  (now a pointer to the top).
- Move the Free-Token Budget section from the bottom (before License) up to a
  hero section near the top, retitled '💰 ~1.9B Free Tokens / Month'.

* fix(plugins): chain payload between emitHookBlocking handlers (#3286) (#3286)

Integrated into release/v3.8.12. Salvaged the emitHookBlocking payload-chaining fix from the now-closed plugins-v4 branch (#3221) and adapted it to the shipped release hooks.ts: each blocking handler now sees the body/metadata as mutated by previous handlers. TDD regression test included (RED before, GREEN after); existing plugins-hooks suites green (19+5), typecheck + lint clean.

* docs(changelog): add v3.8.12 entries for #3280/#3285/#3286/#3287

Quota Sharing Engine repair (#3280), MiniMax-M3 across 8 tiers (#3287),
emitHookBlocking payload chaining (#3286), Chipotle CodeQL hardening (#3285);
updated the contributors hall.

* chore(governance): raise coverage gate 40 -> 60

test:coverage now enforces 60/60/60/60 (statements/lines/functions/branches);
real coverage is ~75-82% so this tightens the floor without new test work.
Updates the c8 --check-coverage thresholds in package.json and the matching
references in CLAUDE.md (Quick Start, testing table, Copilot policy, Hard
Rule #9). Salvaged from the never-pushed chore/skills-governance-tdd-vps
branch; the i18n CLAUDE.md mirrors carry a separate pre-existing drift and
are not gated by check-docs-sync.

* chore(release): finalize v3.8.12 changelog + env-doc-sync + test drift

* fix(quota,sse): clear SonarCloud new-reliability findings on the v3.8.12 diff

- chipotle/grokTls: explicit null checks instead of a Promise in a boolean
  conditional (behavior-preserving; clears the 2 MAJOR reliability bugs)
- sqliteQuotaStore.poolUsage: drop the unreachable dimMap scan loops (dimMap
  was never populated) — the lightweight snapshot already returns no
  dimensions; poolUsageWithDimensions() is the plan-aware path
- BudgetTab: presentation role + keyboard handler on the checkbox wrapper

* Release v3.8.13 (#3327)

* chore(release): open v3.8.13 development cycle

Bump 3.8.12 → 3.8.13 across package.json, lockfile, electron/, open-sse/, and
docs/reference/openapi.yaml; add the [3.8.13] cycle placeholder to the root
CHANGELOG and the 41 i18n mirrors. Integration branch for the v3.8.13 cycle —
fixes/features land here via per-issue PRs and it merges to main at release time.

* fix(ci): skip auto-deploy when VPS host is unreachable from the runner (#3299)

Integrated into release/v3.8.13

* fix(dev): auto-rebuild better-sqlite3 on Node ABI mismatch at dev startup (#3301)

Integrated into release/v3.8.13

* feat(api): accept path-scoped API keys on client API routes (#3300)

Integrated into release/v3.8.13

* fix(sse): harden against empty responses causing Copilot Chat failures (#3297)

Integrated into release/v3.8.13

* fix(api): remove Completions.me rickroll provider (discussion #3293) (#3302)

Integrated into release/v3.8.13

* fix(opencode-provider): extract contextLength from live model catalog (#3298)

Integrated into release/v3.8.13

* feat(web-cookie): self-service login infrastructure + auto-refresh daemon (#3292)

Integrated into release/v3.8.13

* docs(changelog): record the v3.8.13 PRs merged this round (#3292/#3300/#3297/#3298/#3301/#3302/#3299)

* fix(auth): harden URL token extraction — drop query-string fallback, gate to client routes (security follow-up to #3300) (#3309)

Security follow-up to #3300 — integrated into release/v3.8.13

* docs: rename resolve-issues → review-issues skill references

* fix(dashboard): keep no-auth providers visible under 'Show configured only' (#3290) (#3312)

no-auth providers (opencode, duckduckgo-web, theoldllm, veoaifree-web) never
create a DB connection row so stats.total stays 0, which the configured-only
filter treated as 'unconfigured' and hid them — even though they are always
usable and appear unconditionally in /v1/models. filterConfiguredProviderEntries
now treats displayAuthType === 'no-auth' as configured.

Co-authored-by: uniQta <uniQta@users.noreply.github.com>

* fix(cli): resolve update paths relative to script + recursive backup (#3295) (#3313)

omniroute update always failed on a global install:
- getCurrentVersion() read package.json from process.cwd(), which on a global
  npm/brew install is the user's working dir, not the package root → null →
  'Could not determine current version'.
- createBackup() resolved bin/ from cwd too, and passed the 'cli' directory to
  copyFileSync → EISDIR, swallowed by the catch → 'Failed to create backup'.

Both now resolve package.json/bin relative to the script via import.meta.url,
and the backup uses cpSync({recursive:true}) so the cli/ directory is copied.

Co-authored-by: uniQta <uniQta@users.noreply.github.com>

* fix(theoldllm): read upstream body once to avoid [502] body-already-read (#3296) (#3314)

On the cached-token path the executor never enters the refresh branch, so the
same upstream Response was read with .text() twice (token-rejection check +
final body). A Response body is single-use, so the second read threw
'Body is unusable: Body has already been read', caught and surfaced as [502].

Read the body once into finalBody and only re-read after a token-rejection
refetch.

Co-authored-by: onizukashonan14-png <onizukashonan14-png@users.noreply.github.com>

* fix(sse): strip leaked internal tool envelopes from streaming output (#3311)

Integrated into release/v3.8.13

* fix(sse): expose Claude + Gemini budget tiers in the antigravity catalog (#3184) (#3303)

Integrated into release/v3.8.13 (#3184)

* fix(catalog): compute combo context_length from known targets only (#3304)

Integrated into release/v3.8.13 — live contextLength + known-targets combo context (#3298 follow-up)

* chore(i18n): add message keys for proxy UI + vscode/ollama endpoint (#3307)

Integrated into release/v3.8.13 — i18n message keys for proxy UI + vscode/ollama

* feat(dashboard): i18n the proxy settings UI (#3310)

Integrated into release/v3.8.13 — i18n the proxy settings UI

* feat(api): model catalog enrichment + MCP model-catalog tools (#3306)

Integrated into release/v3.8.13 — model catalog enrichment + MCP model-catalog tools, reconciled with #3309 URL-token hardening

* test(catalog): align Antigravity preview-alias test with #3303 budget tiers

#3303 added the Gemini `-high`/`-low` budget tiers to ANTIGRAVITY_PUBLIC_MODELS
(user-callable on the Antigravity OAuth backend, verified via #3184), but did
not update the catalog-route test that asserted `antigravity/gemini-3.1-pro-high`
must NOT be exposed. The assertion now reflects the intended behavior — the
client-visible budget alias IS surfaced — while keeping the legacy
`gemini-claude-*` alias keys unexposed. Caught running the full catalog suite
on the merged release HEAD (the #3303 round only ran the antigravity-aliases
and usage-hardening files).

* docs(changelog): record the 6 PRs merged this review round into v3.8.13

#3306/#3307/#3310 (New Features — VS Code split: catalog+MCP, i18n keys, proxy
UI i18n), #3311/#3303/#3304 (Bug Fixes — SSE envelope sanitizer, antigravity
budget tiers, combo known-targets context_length).

* chore(release): finalize v3.8.13 changelog and cleanup

Finalize the v3.8.13 changelog with release date, maintenance notes,
and contributor credits. Update MCP docs to reference the correct tool
inventory diagram, exclude nested .claude worktrees from ESLint scans,
and tighten a response sanitizer type guard.

* fix(dashboard): refresh connections after provider auth import (#3320)

Integrated into release/v3.8.13 — refresh connections after provider auth import

* fix(codex): strip client-only params on native /responses passthrough (#3317) (#3325)

A /v1/responses request against the built-in codex/ provider does an
openai-responses -> openai-responses passthrough (CodexExecutor.transformRequest
returns the body early for _nativeCodexPassthrough). It forwarded client-only
fields verbatim and the Codex upstream rejected them with 400 Unsupported
parameter: prompt_cache_retention / safety_identifier / user — breaking Factory
Droid (which injects all three). The chat-completions path already strips these
(base.ts #1884, openai-responses translator #2770) but the passthrough skips
translation. Strip the three fields in the shared block before the passthrough
return; user is removed unconditionally since Codex /responses always rejects it.

Co-authored-by: tycronk20 <tycronk20@users.noreply.github.com>

* fix(dashboard): normalize agent-bridge /state response to stop page crash (#3318) (#3326)

The Agent Bridge page seeded a well-shaped initialData default then replaced it
wholesale with the raw /api/tools/agent-bridge/state response. The route returns
{ server, agents } but the UI reads { serverState, agentStates, bypassPatterns,
mappings }, so serverState became undefined and AgentBridgeServerCard crashed on
serverState.running — surfaced as the full-page 'Internal Server Error' boundary
(client render error, not a real 5xx).

Add a shared normalizeAgentBridgeState() that maps the route shape into the page
contract (server.running/certExists -> serverState) and always returns safe
defaults (never undefined serverState). Wired into both the SSR loader (page.tsx)
and the polling hook. The legacy 'agents' entry shape differs from AgentStateEntry
so it is not coerced; full route<->page contract reconciliation (port, upstreamCa,
bypassPatterns, mappings, agentStates) is a follow-up.

Co-authored-by: tycronk20 <tycronk20@users.noreply.github.com>

* docs: VS Code/Ollama endpoints + env & i18n tooling (#3319)

Integrated into release/v3.8.13 — VS Code/Ollama docs + env & i18n tooling

* feat(provider): test-all endpoint, rate-limit overrides, visibility f… (#3267)

Integrated into release/v3.8.13 — provider test-all endpoint, rate-limit overrides, model visibility

* feat: auto-combo optimization, playground model dropdown, only-configured toggle (#3322)

Integrated into release/v3.8.13 — auto-combo candidate expansion + playground dropdown + only-configured toggle

* feat(api): VS Code Copilot Ollama-compatible BYOK endpoint (#3316)

Integrated into release/v3.8.13 — VS Code Copilot Ollama-compatible BYOK endpoint (reconciled with #3306/#3309 auth hardening)

* chore(release): document #3320 in the v3.8.13 changelog + contributor credits

---------

Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
Co-authored-by: Wilson <pedbookmed@gmail.com>
Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: uniQta <uniQta@users.noreply.github.com>
Co-authored-by: onizukashonan14-png <onizukashonan14-png@users.noreply.github.com>
Co-authored-by: tycronk20 <tycronk20@users.noreply.github.com>
Co-authored-by: Vinayrnani <vinayrnani@gmail.com>

* fix(electron): ship loginManager.js in the packaged app (#3292 regression) (#3334)

#3292 added electron/loginManager.js and a require("./loginManager") in
main.js but did not add it to electron-builder's build.files allowlist, so
the packaged app crashed at startup with "Cannot find module './loginManager'"
on the Linux/macOS smoke tests (v3.8.13 Electron release fragment).

Add loginManager.js to build.files, plus a regression test that asserts every
local require("./x") in the Electron entry points is shipped.

* fix(startup): correct autoRefreshDaemon import alias (@/ -> @omniroute/open-sse) (#3292) (#3335)

instrumentation-node.ts imported the #3292 cookie auto-refresh daemon via
"@/open-sse/services/autoRefreshDaemon". The @/ alias maps to src/, but the
daemon lives in the open-sse workspace, so the import resolved to the
non-existent src/open-sse/... and threw "Cannot find module" at runtime in the
built standalone. A try/catch made it non-fatal (the daemon silently never
ran), which kept typecheck and the dev server green, but the packaged Electron
app's strict startup-log smoke test failed on the "Cannot find module" line.

Use the correct @omniroute/open-sse alias, plus a regression test banning
@/open-sse/* imports across src/.

* fix(security): use trusted internal origin for provider auto-sync self-fetch (CodeQL #323 SSRF) (#3336)

POST /api/providers fires a credential-bearing self-fetch to the new
connection's /sync-models route (forwarding the management cookie + internal
sync auth headers). #3267 built that origin from new URL(request.url).origin —
the client-controlled Host header — so a (management-authenticated) caller
could redirect the internal request to an arbitrary host, exfiltrating the
internal sync auth token (CodeQL js/request-forgery, critical, alert #323).

Derive the origin from the trusted loopback/env-pinned base URL via a new
getModelSyncInternalBaseUrl() helper (same source the model-sync scheduler
already uses), never from the incoming request. Adds a regression test.

* fix(electron): swallow auto-updater check rejection to avoid unhandled rejection (#3339)

checkForUpdates() is fired unawaited from a setTimeout at startup. The
underlying autoUpdater.checkForUpdates() rejects on a 404 (release update
manifest not published yet), offline, or rate-limit — and the uncaught
rejection surfaced as an "Unhandled Rejection", which the packaged-app smoke
test treats as fatal (failed the macOS-intel v3.8.13 build; passed elsewhere
only by timing race). The autoUpdater "error" event still notifies the user;
wrap the await so the promise rejection never escapes. Adds a regression test.

* Release v3.8.14 (#3340)

* chore(release): open v3.8.14 development cycle

Version bump 3.8.13 -> 3.8.14 (root + electron + open-sse + openapi + lockfiles).
Seed the v3.8.14 changelog with the four post-tag hotfixes that shipped to
Docker/Electron in v3.8.13 but missed the immutable npm 3.8.13 (#3336 SSRF /
CodeQL #323, #3334/#3335/#3339 Electron packaging). i18n CHANGELOG mirrors get
the in-progress placeholder section.

* feat: add per-provider custom headers support for OpenAI/Anthropic-compatible nodes (#3338)

Integrated into release/v3.8.14

* fix: Kiro Builder ID token import fails with Bad credentials (#3333)

Integrated into release/v3.8.14 — adds Builder ID cached-creds + OIDC refresh path for Kiro token import, with regression tests (#3333).

* Improve code quality: auto-pr/docstrings-1780792063 (#3337)

Integrated into release/v3.8.14 — docstring for context analytics route re-export.

* fix(catalog): remove minimaxai/minimax-m3 from NVIDIA NIM tier (404 upstream) (#3329) (#3341)

NVIDIA NIM does not host minimaxai/minimax-m3 — every request returns
404 page not found, while sibling minimaxai/minimax-m2.7 on the same provider
works. Advertising a model that 404s is a catalog bug; remove it from the nvidia
tier (it remains on the tiers that actually serve MiniMax M3). Re-add only once
NVIDIA serves it.

Co-authored-by: mikmaneggahommie <mikmaneggahommie@users.noreply.github.com>

* fix(cli): write OpenCode config to ~/.config on all platforms incl. Windows (#3330) (#3343)

resolveOpencodeConfigDir used %APPDATA% on Windows, but OpenCode reads its
config from XDG ~/.config/opencode/ on every platform (on Windows:
%USERPROFILE%\.config\opencode\, NOT %APPDATA%). So a Windows user who
configured OpenCode via the dashboard had the file written where OpenCode never
looks — it silently had no effect.

Use the XDG path (XDG_CONFIG_HOME || ~/.config) unconditionally. Update the UI
note + route JSDoc, and flip the three tests that encoded the old %APPDATA%
behavior (t40 per-platform + card-note, cli-runtime-extended getCliConfigPaths).

Co-authored-by: abdulkadirozyurt <abdulkadirozyurt@users.noreply.github.com>

* fix(proxy): make auto-selection fallback opt-in (#3332) (#3344)

selectWorkingProxyFallback (Step 11 of resolveProxyForConnection) listed ALL
registry proxies, ignoring assignments and per-connection proxy_enabled, and
returned the first working one with level:'autoSelect'. So a single proxy added
to the registry silently became a global fallback for every connection's traffic.

Gate it behind a new PROXY_AUTO_SELECT_ENABLED feature flag (default off): the
fallback now no-ops unless the operator opts in. No registry proxy becomes a
silent global default anymore.

Co-authored-by: hertznsk <hertznsk@users.noreply.github.com>

* fix(sse): treat MiniMax M3 as multimodal so vision isn't stripped (#3328) (#3342)

MiniMax M3 via the opencode provider (oc/minimax-m3-free) appeared blind:
image inputs didn't reach the model, while the same model in Cline could
see them. Verified empirically that MiniMax M3 on the opencode upstream IS
multimodal -- a base64 image is described correctly (it returns 403 only
for remote image URLs, which it doesn't accept).

Root cause: OmniRoute treated MiniMax M3 as a non-vision model in two
places, so when compression was active the image was replaced with a text
placeholder before dispatch:
- compression's modelSupportsVision() heuristic (lite.ts) only matched
  gpt-4/4o/claude-3/gemini/vision -- minimax was absent -> replaceImageUrls
  stripped the image.
- the opencode minimax-m3-free catalog entry lacked supportsVision, so the
  combo vision-capability gate could also exclude/mishandle it.

Add 'minimax-m3' to the vision heuristic and supportsVision: true to the
opencode minimax-m3-free entry. TDD: a failing-then-passing test in
compression/lite.test.ts proves replaceImageUrls now keeps images for
minimax-m3 ids, plus a registry assertion mirroring the #2822 qwen test.

Reported-by: @mikmaneggahommie

* docs(i18n): translate 25 core documentation files to Indonesian (#3348)

Integrated into release/v3.8.14 — Indonesian i18n docs.

* fix(review): resolve /review-reviews battery findings (LEDGER-1..11) on v3.8.14 (#3350)

Integrated into release/v3.8.14 — /review-reviews battery hardening (LEDGER-1..11) for #3338 custom-headers + #3333 kiro, plus cycle-test drift fixes (#3329/#3330/#3332).

* fix(provider-proxy): honor per-account proxy toggles (#3349)

Integrated into release/v3.8.14 — honor per-account proxy toggles + auto-fallback opt-in via PROXY_AUTO_SELECT_ENABLED.

* fix(dashboard): remove duplicate Distribute Proxies button on provider page (#3352)

* fix(providers): reduce proxy label noise (#3346)

Integrated into release/v3.8.14 — reduce proxy label noise + a11y (aria-label/sr-only).

* fix(duckduckgo): restore bare Response contract and rebase onto release/v3.8.14 (#3323)

Integrated into release/v3.8.14 — browser-backed cookie providers (duckduckgo/claude-web) with restored executor contract + unit tests.

* fix(noauth): expose only usable model aliases (#3345)

Integrated into release/v3.8.14 — noauth usable-alias filtering + registry alias plumbing (veo-free).

* fix(dashboard): stop infinite config-load loop on Hermes Agent detail page (#3353)

* fix(electron): tree-kill the server on exit/update to release the omniroute.exe lock (#3347) (#3354)

* chore(release): finalize v3.8.14 changelog + clear release-gate drift

- CHANGELOG: finalize the v3.8.14 section (date, full New Features/Bug Fixes/
  Maintenance coverage of all 16 cycle commits, Contributors hall of 12).
- docs: document OMNIROUTE_BROWSER_POOL + WEB_COOKIE_USE_BROWSER (#3323) in
  .env.example + ENVIRONMENT.md; regenerate the id/llm.txt strict mirror (#3348
  had translated it; llm.txt mirrors must match root).
- test(proxy-fetch): #3323 made tlsClient.available a computed getter — stub it
  via Object.defineProperty instead of assignment (5 tests were red on the base).

* fix(translator): coerce Gemini functionDeclaration parameters to an OBJECT schema (#3357) (#3360)

* fix(gemini): resolve truncation/suppression of false positive textual tool call markers in backticks (#3358)

Integrated into release/v3.8.14 — Gemini/Antigravity textual tool-call marker normalization (no false-positive suppression + split-chunk buffering).

* docs(changelog): add #3358 Gemini textual tool-call normalization to v3.8.14

* fix(dashboard): surface real analytics error instead of generic placeholder (#3356) (#3361)

The Analytics page discarded the server's error body on a non-OK response and
rendered a generic "An error occurred", so users (and maintainers) could not see
why /api/usage/analytics 500'd after an upgrade. Now the route returns the real
reason via buildErrorBody (sanitized, Hard Rule #12) and the page surfaces it via
a new readFetchErrorMessage helper that handles both the OpenAI-style and legacy
error shapes.

Reported-by: @superti4r

---------

Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Someres <168349709+quanturbo@users.noreply.github.com>
Co-authored-by: Dong Mengzhe <154944819+Lang-Qiu@users.noreply.github.com>
Co-authored-by: mikmaneggahommie <mikmaneggahommie@users.noreply.github.com>
Co-authored-by: abdulkadirozyurt <abdulkadirozyurt@users.noreply.github.com>
Co-authored-by: hertznsk <hertznsk@users.noreply.github.com>
Co-authored-by: Krisna Santosa <54174372+KrisnaSantosa15@users.noreply.github.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: Wilson <pedbookmed@gmail.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Ardem2025 <ardemb22@gmail.com>

* docs(changelog): complete v3.8.14 — add #3356 + @nullbytef0x/@Ardem2025 to contributors (#3362)

The release PR #3340 was merged before these changelog lines landed: the #3356
Usage-Analytics-error bullet and the @nullbytef0x (#3357) / @Ardem2025 (#3358)
contributor rows. Code for all three was already in the squash; this only
completes the changelog/credits so the GitHub release notes are accurate.

* fix(ci): drop explicit any on executeWithUpstreamStartTimeout call (t11 any-budget) (#3364)

The v3.8.14 merge introduced `executeWithUpstreamStartTimeout<any>(...)` in
chatCore.ts, pushing the file's explicit-any count to 1 over its budget of 0
(check:any-budget:t11, a blocking CI lint-job gate). The generic T is already
inferable from the `execute` callback's return type, so drop the explicit
`<any>` and let inference do it — no behavior change, typecheck:core stays clean.

* test(translator): align gemini-2.5-flash maxOutputTokens cap to 65536 (#3358) (#3367)

#3358 added the gemini-2.5-flash model spec with its real 65536 max-output cap
(previously the model had no spec and fell to an 8192 default). The Claude→Gemini
clamp test still asserted 8192, so it failed deterministically — the single real
failure behind the v3.8.14 CI red (Unit Tests 3/8, Coverage Shard 3/8, Node
24/26 Compatibility 1/2 all hit this one test; E2E 5/6 was fail-fast collateral).

* Release v3.8.15 (#3373)

* chore(release): open v3.8.15 development cycle

Version bump 3.8.14 -> 3.8.15 (root + electron + open-sse + openapi + lockfiles)
and seed the v3.8.15 changelog placeholder (root + 41 i18n mirrors).

* fix(catalog): add getTokenLimit fallback for combo targets with unknown context (#3369)

Integrated into release/v3.8.15. Fixes applied on the contributor's branch: removed duplicate JSDoc opening in accountFallback.ts and dropped a test asserting unreachable catalog behavior (models with no registry/spec/synced source are filtered before the getTokenLimit fallback at catalog.ts:499).

* fix(combo): add 429 to PROVIDER_FAILURE_ERROR_CODES to prevent infinite retry loop (#3366)

Integrated into release/v3.8.15. Comment block reconciled on the contributor's branch to remove the contradictory 'intentionally excluded' text that remained from the original code.

* fix(auto-combo): include no-auth providers declaratively (#3365)

Integrated into release/v3.8.15. Cleanup applied on contributor's branch: removed duplicate migration 095 (already exists from PR #3338), reverted CHANGELOG.md and i18n changelogs to release versions (release process owns these), dropped package version-bump noise from stale fork base. Core feature — declarative no-auth via serviceKinds metadata, declarative VEO as 'video' provider, anonymousFallback flag for opencode-zen/opencode-go — integrated cleanly.

* fix(migrations): restore 095_provider_node_custom_headers migration

The squash merge of PR #3365 accidentally deleted this migration because
the cleanup commit on the contributor's branch included 'git rm' for the
file (which was a duplicate on their branch). The migration was merged
in v3.8.14 via PR #3338 and must be present in the release branch.

Restoring from git history.

* fix: update Command Code base URL from /alpha/ to /provider/v1/ (#3372)

Integrated into release/v3.8.15.

* feat(error-rules): provider-specific error classification with scope (#3370)

Integrated into release/v3.8.15. PR has genuine value beyond #3369: (1) getProviderErrorRuleMatch now accepts native Headers objects from fetch(); (2) checkFallbackError also uses the provider rule registry — the real end-to-end wiring in the combo fallback path; (3) S4 end-to-end test proving the wiring fires. Merge commit on contributor branch resolved the add/add conflict by taking the #3370 version throughout.

* fix(auto-combo): validate web-session credentials (#3371)

Integrated into release/v3.8.15. Core feature: provider-aware web-session credential validation — hasUsableWebSessionCredential() replaces the broad Object.keys check in virtualFactory.ts, ensuring only sessions with the required storageKeys are included in auto-combo. Cleanup: removed duplicate 095 migration, reverted CHANGELOG/i18n, dropped package bump noise.

* fix(migrations): restore 095_provider_node_custom_headers (deleted again by #3371 squash)

Same issue as after #3365: git rm in the contributor cleanup commit
was included in the squash, deleting this migration from release.
Permanent fix needed: use 'git checkout origin/release -- <file>'
instead of 'git rm' when cleaning up duplicate files in contributor branches.

* fix(kiro): probe Windows %APPDATA%\kiro\storage.db in auto-import (#3363) (#3375)

Integrated into release/v3.8.15. Test fix applied: kiro-windows-auto-import-3363.test.ts now sets DATA_DIR to a fresh temp dir before importing app modules, ensuring isAuthRequired() sees an empty settings DB (no password → auth not required). This fixed test 4 (synthetic SQLite) which was getting 401 due to settings DB state leakage.

* chore(release): finalize v3.8.15 changelog — 2026-06-07

---------

Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Muhammad Nabil Muyassar Rahman <65392758+TapZe@users.noreply.github.com>
Co-authored-by: kiro-agent[bot] <245459735+kiro-agent[bot]@users.noreply.github.com>

* chore(release): open v3.8.16 development cycle

* fix(ci): stop the E2E shard from being cancelled mid-run (timeout headroom) (#3387)

The heaviest E2E shard (5/6 — responsive viewport matrix + studio/smoke) overran
the job's 20m timeout-minutes because each shard re-runs `npm run build` (~5m)
before Playwright, then runs ~24 serial tests with retries:2. The job was killed
(CANCELLED mid-run, 'Terminate orphan process') instead of any test failing.

- Bump test-e2e timeout-minutes 20 -> 35 (cumulative build+tests headroom).
- Lower the Playwright per-test timeout 600s -> 180s so a genuine hang fails fast
  and visibly (a clear per-test timeout) instead of silently eating the job budget.

* fix(ci): give the heavy E2E shard headroom + stream live progress (#3392)

The 35m bump still wasn't enough — shard 5/6 (responsive viewport matrix +
studio/smoke, ~24 serial tests after a ~5m build) was still cancelled at 35m,
and the `github` Playwright reporter buffers output so the cancelled log showed
no per-test results (couldn't tell which test was slow).

- e2e timeout-minutes 35 -> 50 (the shard observably needs >35m; other shards
  finish in ~7m so they're unaffected).
- Playwright CI reporter github -> line so per-test progress + timing stream
  live to the job log, making any genuinely slow/hung test diagnosable.

* fix(env): correct casing of OMNIROUTE_TRACE in .env.example and related files (#3393)

Integrated into release/v3.8.16

* fix(featureFlags): update description for PRICING_SYNC_ENABLED to clarify environment variable requirement (#3394)

Integrated into release/v3.8.16

* fix(account-fallback): preserve provider cooldown dedupe state (#3381)

Integrated into release/v3.8.16

* ci(docker): also build & publish the -web image variant (#3389)

Integrated into release/v3.8.16

* fix(stream): solve false positive textual tool-call marker truncation using emitted content state (#3382)

Integrated into release/v3.8.16

* fix(stream): drop empty choices chunks instead of emitting retry text (#3400)

Integrated into release/v3.8.16

* feat: adaptive keepalive threshold for web-session providers (#3397)

Integrated into release/v3.8.16

* feat: add web-session pool observability (MCP tool + health-matrix) (#3395)

Integrated into release/v3.8.16

* fix(providers): refresh model list after provider sync (#3402)

Integrated into release/v3.8.16

* feat(vision-bridge): auto-route to fastest vision model (#3377)

Integrated into release/v3.8.16

* fix: server-side context cache pinning, stop proxy message leaks, persist context_cache_protection toggle (#3399)

Integrated into release/v3.8.16

* fix(sse): eliminate race window in usageTokenBuffer settings update (#3405)

Integrated into release/v3.8.16

* feat: add bulk web-session credential import endpoint (#3403)

Integrated into release/v3.8.16

* feat: add REST API for session pool health (dashboard interface) (#3404)

Integrated into release/v3.8.16

* docs: add Codex CLI configuration guide for OmniRoute

Add a comprehensive guide for configuring Codex CLI to use OmniRoute as an OpenAI-compatible backend.

Document ready-to-use config examples, Responses API routing behavior, context window settings, token limits, model profiles, and troubleshooting guidance to help users avoid direct-provider compatibility issues.

* fix(mitm): getMitmStatus stub returns graceful status in Docker (#3390) (#3408)

* fix(executor): strip trailing assistant text for Mistral (user-last required) (#3396) (#3409)

* fix(sanitizer+stream): tighten textual tool-call detection, flush partial buffer (#3355) (#3410)

* fix(docs+ui): add MDX frontmatter to Codex CLI guide, fix setState-in-effect lint

- docs/guides/CODEX-CLI-CONFIGURATION.md was missing the YAML frontmatter
  block required by fumadocs (title/version/lastUpdated), causing the
  production build to fail with "invalid frontmatter" MDX error.
- CodexCliGuideModal.tsx called setLoading/setError synchronously in a
  useEffect body, triggering the react-hooks/set-state-in-effect lint error.
  Refactored to an internal async function with an `cancelled` guard to
  prevent state updates on unmounted components.

* fix(tests): align test suite to post-#3355/#3366/#3399 behavior

- Remove 429 from PROVIDER_BREAKER_FAILURE_STATUSES; 429 belongs to
  connection cooldown, not whole-provider breaker (CLAUDE.md §resilience).
  PR #3366 correctly added 429 to PROVIDER_FAILURE_ERROR_CODES in
  accountFallback.ts (combo infinite-retry fix) but the parallel change
  to chat.ts was wrong — the integration test from v3.8.10 confirms this.

- Align stream-utils tests to PR #3399 (SYNTHETIC_CLAUDE_EMPTY_RESPONSE_TEXT
  → "", message.content → null) and PR #3355 (malformed tool-call buffer
  now emitted as plain text, not suppressed).

- Align services-branch-hardening test to PR #3399 (pinnedModel always
  null from applyComboAgentMiddleware; server-side session pinning replaced
  client-side <omniModel> tag extraction).

- Align combo-routing-engine context-cache tests to PR #3399 (no <omniModel>
  tag in output, no X-OmniRoute-Model header, priority routing unchanged).

* ci: speed up e2e shards with build and browser cache

Upload the Next.js build from the build job and reuse it across E2E
shards to avoid rebuilding in each shard. Increase Playwright sharding
from 6 to 9, cache Chromium browsers, and lower the E2E timeout to match
the faster expected runtime.

Add a Codex CLI configuration skill for OmniRoute setup and ignore local
credential-bearing setup prompts.

* fix(agentSkills): cast next-fetch opts to satisfy TypeScript overload check

* fix(ci+tests): fix E2E artifact (exclude 558MB standalone/node_modules, cp after download) and update skill count to 43

* fix(tests+ci): update 42→43 skill count, fix E2E artifact path

- Unit/integration tests: update hardcoded 42→43 in 7 test files
  (agentSkillTools-mcp, agentSkills-catalog, agentSkills-generator,
  agent-skills-content, agent-skills-discovery, listCapabilities-a2a)
  to match the 43rd skill (config-codex-cli) added in the previous commit.
- Include CONFIG_SKILL_IDS in integration content test ALL_IDS so
  skills/config-codex-cli/ is no longer "unexpected".
- listCapabilities.ts: change totalSkills from literal 42 to catalog.length
  so it adapts to catalog growth automatically.
- computeCoverage assertions: include config.have in totalSkills check.
- CI: switch E2E artifact from upload-artifact path (ambiguous stripping)
  to explicit tar archive. Fixes "Could not find a production build in
  ./.build/next" — the previous approach's download path was double-nested
  (.build/next/next/...) due to upload-artifact LCA computation. tar -czf
  stores .build/next/... relative to CWD; tar -xzf restores them verbatim.
- Also exclude .build/next/cache from the tar to keep archive lean.
- feat(translator): strip client_metadata in Responses→Chat translation
  (Mistral 422 extra_forbidden fix); add regression test.

* docs: update Codex CLI profile naming guidance

Update Codex CLI docs and configuration skill to use the v0.137+
profile file naming format: ~/.codex/<name>.config.toml instead of
the deprecated profile- prefix.

Clarify that missing profile files silently fall back to defaults, and
rename the setup workflow heading to match the config-codex-cli skill.

* ci(e2e): increase E2E shard timeout 30→45min for slow runners

* fix(e2e): wait for add-dialog close before clicking Edit (backdrop race)

* fix(e2e): dismiss import-models modal after adding connection (sync-models mock + close)

* fix(e2e): use .first() on Close button to avoid strict-mode violation (2 elements)

* chore(release): finalize v3.8.16 CHANGELOG — 2026-06-08

* chore(release): cover missing agentSkills TS-overload fix in CHANGELOG

* fix(docker): copy playwright from builder instead of npx fetch in runner-web

npx playwright falls back to a registry download when playwright is absent from
the slim runtime image's node_modules. On GitHub-hosted runners this download
fails with exit 127, breaking both amd64 and arm64 -web image builds.

Fix: COPY playwright and playwright-core from the builder stage and invoke
node node_modules/playwright/cli.js directly — no network access, same version,
and playwright remains available at runtime for web-session providers.

* chore(release): open v3.8.17 development cycle

* deps: bump electron from 42.3.2 to 42.3.3 in /electron (#3441)

Integrated into release/v3.8.17

* deps: bump the production group with 10 updates (#3444)

Integrated into release/v3.8.17

* deps: bump the development group with 4 updates (#3445)

Integrated into release/v3.8.17

* deps: bump electron-updater from 6.8.8 to 6.8.9 in /electron (#3442)

Integrated into release/v3.8.17

* deps: bump electron-builder from 26.14.0 to 26.15.2 in /electron (#3443)

Integrated into release/v3.8.17

* fix(sse): normalize provider ids to strings (#3427)

Integrated into release/v3.8.17

* fix(command-code): revert chat endpoint to /alpha/generate and fix model sync discovery (#3432)

Integrated into release/v3.8.17

* fix(analytics): scope SQL named params per query context (#3446) (#3447)

Integrated into release/v3.8.17

* fix claude-web and cleanup (#3449)

Integrated into release/v3.8.17

* feat: add ZenMux provider (Phase 2B of #3368) (#3429)

Integrated into release/v3.8.17

* feat: add LMArena provider (Phase 2A of #3368) (#3421)

Integrated into release/v3.8.17

* feat: add Gemini Business provider (Phase 2C of #3368) (#3436)

Integrated into release/v3.8.17

* fix: probe container bridge network IP in healthcheck (#3151) (#3434)

Integrated into release/v3.8.17

* fix(publish): remove onnxruntime CUDA binary from tarball to avoid 413 (#3437)

Integrated into release/v3.8.17

* docs(opencode-plugin): lead with the why — make plugin the recommended path over @omniroute/opencode-provider (#3418)

Integrated into release/v3.8.17

* docs: close critical documentation gaps (ACP, router strategies, APIs, compression) (#3438)

Integrated into release/v3.8.17

* fix(stream): allow OpenAI usage-only empty choices chunks (#3422)

Integrated into release/v3.8.17

* fix(catalog): make combos auto-compute context_length for any provider id form (#3417)

Integrated into release/v3.8.17

* fix(stream): resolve index mismatch in textual tool-call slicing and deduplicate containsTextualToolCallMarker (#3413)

Integrated into release/v3.8.17

* feat(opencode-plugin): per-prefix API format + debug logging + free-label normaliser (3 mrmm-fork backports) (#3420)

Integrated into release/v3.8.17

* fix(translator): use non-empty reasoning_content placeholder on cache miss instead of empty string (#3433)

Integrated into release/v3.8.17

* feat(plugin+api): auto combos + free model quota display + /api/combos/auto (#3435)

Integrated into release/v3.8.17

* test(auto-combo): cover same-provider connection identity (#3378)

Integrated into release/v3.8.17

* feat: add connection pagination, health filter, batch delete confirmation, and custom banned keywords (#3454)

Integrated into release/v3.8.17

* fix(translator): strip function_call.id for Vertex AI provider (#3440) (#3457)

Vertex AI's FunctionCall/FunctionResponse protos have no id field; emitting it made Vertex reject tool calls with 400 'Unknown name id'. The id is now stripped only when the routed provider is vertex/vertex-partner (threaded via credentials._provider), preserving it for the public Gemini API where Gemini 3+ uses it for signature matching.

Co-authored-by: nullbytef0x <nullbytef0x@users.noreply.github.com>

* fix(claude): respect client anthropic-beta instead of forcing thinking/effort betas (#3415) (#3458)

Claude Code -> claude-opus-4-8 turns intermittently died with 'tool call could not be parsed (retry also failed)'. OmniRoute's claude identity cloak rebuilt the anthropic-beta header from scratch and unconditionally forced interleaved-thinking-2025-05-14 (+ advanced-tool-use / effort for heavy agents), even when the client never negotiated them. The forced interleaved-thinking conflicts with tool_choice-forced turns, producing malformed opus tool_use streams (and sibling 400 'Thinking may not be enabled when tool_choice forces tool use').

selectBetaFlags now takes the client's inbound anthropic-beta: when present, thinking/effort betas are only emitted if the client requested them. Opaque clients (no header — the OAuth cloak path) keep the full set unchanged, so existing behavior and the #2454 model-tier gating are preserved.

Co-authored-by: Forcerecon <Forcerecon@users.noreply.github.com>

* fix(catalog): surface imported models on no-auth providers in /api/v1/models (#3200) (#3463)

The custom-models loop in getUnifiedModelsResponse gated every model through hasEligibleConnectionForModel(getConnectionsForProvider(...)). no-auth providers (theoldllm, etc.) never create DB connection rows, so that returned [] and the gate dropped every imported/custom model for them — the Playground dropdown showed nothing for imported models while built-in/custom models on auth providers worked. Built-in models survived because they go through providerSupportsModel(), which already has a no-auth bypass (#2798).

The custom-model gate now applies the same no-auth bypass, keeping the eligibility check (with parentProviderType) intact for auth providers.

Co-authored-by: tjengbudi <tjengbudi@users.noreply.github.com>
Co-authored-by: a2belugin <a2belugin@users.noreply.github.com>

* fix(browser): avoid bundling optional cloakbrowser import (#3460)

Integrated into release/v3.8.17

* fix(command-code): align CLI version header (#3462)

Integrated into release/v3.8.17

* Add Endpoint Token Saver visibility setting (#3461)

Integrated into release/v3.8.17

* docs(env): document COMMAND_CODE_VERSION override (#3462 follow-up)

#3462 added a process.env.COMMAND_CODE_VERSION read but did not document it,
tripping the env-doc-sync gate on the release branch (PR-merges bypass the
pre-commit check-docs-sync hook). Add the var to .env.example + ENVIRONMENT.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Add model catalog name feature flag (#3464)

Integrated into release/v3.8.17

* chore(release): v3.8.17 — 2026-06-09

CHANGELOG: 8 features, 15 bug fixes, 6 maintenance entries (29 bullets / 32 commits since v3.8.16).
i18n: sync [3.8.17] section to all 41 locale CHANGELOG files.

fix(translator): strip empty reasoning_content on non-tool-call kimi-k2 messages (#3433 regression)
fix(translator): update placeholder assertion for non-empty cache-miss behaviour (test alignment)
fix(executor): lmarena.ts return wrapper shape {response,url,headers,transformedBody} (#3421 regression)
test: align lmarena-provider + tool-request-sanitization to corrected executor contract

* fix(docs): add ACP.md frontmatter and flatten docs/meta.json pages format

fumadocs-mdx requires a YAML title in every .md file and does not support
nested object entries in meta.json pages arrays — both were introduced by
PR #3438 and broke the webpack build.

* fix(opencode-plugin): remove duplicated blocks + wire missing schema fields (#3435 merge corruption)

PR #3435's branch shipped a corrupted index.ts that never built — the npm
publish-opencode-plugin job failed on DTS errors. Root causes:

- Duplicate apiFormat block (ensureV1Suffix/DEFAULT_ANTHROPIC_PREFIXES/
  resolveApiBlock) — kept the canonical #3420 copy (anthropic url WITHOUT /v1),
  removed the duplicate that wrongly appended /v1 to the Anthropic SDK base.
- Duplicate debug-logging block (DebugLogEntry + debugLog* + createDebugLoggingFetch)
  with mid-file imports — kept the canonical copy using top-of-file imports.
- Local normaliseFreeLabel def superseded by the naming.ts extraction —
  removed it, routed the lone caller to the imported _normaliseFreeLabel.
- sdkBaseURL → resolvedBaseURL (undefined identifier in the auth loader).
- featuresSchema missing startupDebug + logLevel (referenced but never declared).
- shortProviderLabel dropped the prefix on long displayName + no alias; now
  keeps the long label, matching the test intent.

Plugin builds (DTS clean) and all 254 tests pass.

* fix(providerRegistry): update Claude model entries to latest versions (#3521)

* Release v3.8.18 (#3482)

* chore(release): open v3.8.18 development cycle

* fix(catalog): stop Codex CLI model-catalog refresh from erroring (#3481)

Codex's model-catalog refresh (codex_models_manager) does
GET /v1/models?client_version=<v> and decodes a JSON object with a
TOP-LEVEL `models` array. OmniRoute answers in the OpenAI-standard
`{object,data}` shape, so codex fails with "missing field `models`"
and logs "failed to refresh available models" on every startup.

Detect codex clients via the `originator` / `user-agent` = `codex_*`
headers they send and add an EMPTY top-level `models: []` so the decode
succeeds. Non-codex OpenAI clients keep the byte-identical `{object,data}`
response.

The array is intentionally empty: codex replaces its built-in per-model
agent prompt (`base_instructions`, ~21k chars) with whatever a populated
entry carries for the selected model, so emitting our catalog would drop
the agent prompt to nothing and break codex's agent behaviour (verified
empirically against codex 0.137). An empty list keeps codex on its
built-in model info — same inference as before, minus the error.

Validated end-to-end with the real handler against codex 0.137:
"failed to refresh available models" → 0 occurrences, instructions
preserved (built-in Codex agent prompt, not empty).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: ignore quality reports and local prompt artifacts

Add generated quality gate reports, metrics files, and local setup prompt
artifacts to .gitignore to prevent committing environment-specific or
temporary files.

* fix(provider): detect Responses API format when body has `input` but … (#3490)

Integrated into release/v3.8.18

* fix(sse): normalize numeric provider ids to strings (#3451)

Integrated into release/v3.8.18

* feat(browserPool): resolve Playwright proxy from proxy_registry DB (#3492)

Integrated into release/v3.8.18

* fix(theoldllm): generate X-Request-Token server-side, drop Playwright (#3491)

Integrated into release/v3.8.18

* feat(plugins): add lifecycle hooks and theme-manager plugin (#3473)

Integrated into release/v3.8.18

* fix(combo): parallel pre-screen + circuit-breaker fast-exit for priority combos (#3169)

Integrated into release/v3.8.18

* feat(ui): unifi active and finished requests into single view #1422 (#3401)

Integrated into release/v3.8.18

* docs(changelog): record #3401, #3473, #3492, #3490, #3451, #3491, #3169 under v3.8.18

* feat(docs): add doc accuracy gate + refresh AGENTS.md counts (#3510)

Integrated into release/v3.8.18

* fix(sse): drop empty-choices chunks without usage instead of injecting retry text (#3513)

PR #3422 ('allow OpenAI usage-only empty choices chunks') reintroduced the
assistant-content injection '[OmniRoute] Upstream returned an empty response.
Please retry.' for empty `choices: []` chunks that carry no valid usage. Clients
(Goose/opencode) feed that text back as a turn and spin in a retry loop -- the
exact regression #3400 had fixed by dropping the chunk.

Restore the drop behavior for the no-usage case while preserving #3422's
standards-compliant forwarding of usage-only `include_usage` final chunks.
Realign the mislabeled stream-utils test (it asserted the injection) and add a
dedicated regression guard.

Reported-by: @mochizzan
Refs: #3502, #3388, #3400, #3422

* fix(authz): fall back to URL token when Authorization isn't a usable Bearer (#3504)

Integrated into release/v3.8.18

* fix(playground): authenticate via session, test key policy by id (#3503)

Integrated into release/v3.8.18

* docs(changelog): record #3510, #3504, #3503 under v3.8.18

* fix: llama base url normalization (#3519)

* docs(changelog): reconcile v3.8.18 — add #3519, #3513, #3435-repair, gitignore chore (full commit↔changelog coverage)

* fix(opencode-plugin): bound regex quantifiers in normaliseFreeLabel (polynomial-ReDoS)

CodeQL js/polynomial-redos: unbounded \s* before an anchored \s*$ allowed
O(n²) backtracking on attacker-influenced display names. Bounded to {0,8}/{1,8}
(ample for any real label spacing). Plugin builds + 254 tests green.

* fix(types): restore clean typecheck:core for v3.8.18 release gate

- getPendingRequests() typed to real shape (was widened to object) → fixes
  unknown 'count' in the unified-requests view (#3401)
- streamChunks log payload cast to its declared type (callLogs.ts)
- preScreenTargets aligned to canonical IsModelAvailable signature (#3169),
  Promise.resolve-normalized so .catch never hits a bare boolean

All 5 gates green: lint(0 err) + typecheck:core + cycles + docs-all + unit + vitest(146).

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Andrey Borodulin <borodulin@gmail.com>
Co-authored-by: Dmitrii Safronov <zimniy@cyberbrain.cc>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>

* Release v3.8.19 (#3526)

* chore(release): open v3.8.19 development cycle

* chore(release): sync electron lockfile to 3.8.19

* feat(quality): quality-gate ratchet + anti-hallucination/rule-enforcement guardrails (Phases 0-6) (#3471)

* feat(quality): generic ratchet comparator (multi-metric, regression-only)

* chore(ci): Fase 0 quality-gate fixes — reconcile coverage gate (40->60), tier npm audit, wire orphaned contract gates, re-enable cheap husky pre-commit

* feat(quality): ratchet engine (collector + frozen baseline + CI job) and provider-consistency gate

- collect-metrics.mjs: emits quality-metrics.json (ESLint warnings + coverage when present)
- quality-baseline.json: frozen baseline (eslintWarnings=3482, regression-only)
- ci.yml: quality-gate job (ratchet + step summary + artifact) and check:provider-consistency in lint job
- check-provider-consistency.ts: every REGISTRY id must be a canonical provider (found krutrim half-registered → allowlisted as known pre-existing, blocks any NEW orphan)
- TDD: 9 tests (5 ratchet + 4 provider-consistency)

* feat(quality): Fase 2 anti-hallucination gates — fetch-targets, openapi-routes, deps allowlist

- check-fetch-targets: every dashboard fetch(/api/...) resolves to a real route.ts; found 7 pre-existing dashboard->route mismatches frozen as KNOWN_MISSING for triage
- check-openapi-routes: every openapi.yaml path resolves to a real route; found 1 stale spec entry (agent-bridge agents/{id}/state) frozen as KNOWN_STALE_SPEC
- check-deps: anti-slopsquatting allowlist (105 deps); new deps need explicit human-reviewed entry
- all wired into CI lint/docs jobs; TDD +12 tests (21 total across 5 gates)

* docs(quality): add quality-gates report + implementation plan to repo root

* feat(quality): Fase 3a — file-size ratchet (freeze 91 files >800 LOC, cap 800 for new)

- check-file-size.mjs: frozen files can only shrink; new files must be <= cap (kills the next 12k-line god-component)
- file-size-baseline.json: 91 files frozen at current LOC (largest 12883)
- wired into CI lint job; TDD 5 tests; --update ratchets the baseline down on shrink

* feat(quality): Fase 3b — duplicati…
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
* chore(release): open v3.8.18 development cycle

* fix(catalog): stop Codex CLI model-catalog refresh from erroring (diegosouzapw#3481)

Codex's model-catalog refresh (codex_models_manager) does
GET /v1/models?client_version=<v> and decodes a JSON object with a
TOP-LEVEL `models` array. OmniRoute answers in the OpenAI-standard
`{object,data}` shape, so codex fails with "missing field `models`"
and logs "failed to refresh available models" on every startup.

Detect codex clients via the `originator` / `user-agent` = `codex_*`
headers they send and add an EMPTY top-level `models: []` so the decode
succeeds. Non-codex OpenAI clients keep the byte-identical `{object,data}`
response.

The array is intentionally empty: codex replaces its built-in per-model
agent prompt (`base_instructions`, ~21k chars) with whatever a populated
entry carries for the selected model, so emitting our catalog would drop
the agent prompt to nothing and break codex's agent behaviour (verified
empirically against codex 0.137). An empty list keeps codex on its
built-in model info — same inference as before, minus the error.

Validated end-to-end with the real handler against codex 0.137:
"failed to refresh available models" → 0 occurrences, instructions
preserved (built-in Codex agent prompt, not empty).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: ignore quality reports and local prompt artifacts

Add generated quality gate reports, metrics files, and local setup prompt
artifacts to .gitignore to prevent committing environment-specific or
temporary files.

* fix(provider): detect Responses API format when body has `input` but … (diegosouzapw#3490)

Integrated into release/v3.8.18

* fix(sse): normalize numeric provider ids to strings (diegosouzapw#3451)

Integrated into release/v3.8.18

* feat(browserPool): resolve Playwright proxy from proxy_registry DB (diegosouzapw#3492)

Integrated into release/v3.8.18

* fix(theoldllm): generate X-Request-Token server-side, drop Playwright (diegosouzapw#3491)

Integrated into release/v3.8.18

* feat(plugins): add lifecycle hooks and theme-manager plugin (diegosouzapw#3473)

Integrated into release/v3.8.18

* fix(combo): parallel pre-screen + circuit-breaker fast-exit for priority combos (diegosouzapw#3169)

Integrated into release/v3.8.18

* feat(ui): unifi active and finished requests into single view diegosouzapw#1422 (diegosouzapw#3401)

Integrated into release/v3.8.18

* docs(changelog): record diegosouzapw#3401, diegosouzapw#3473, diegosouzapw#3492, diegosouzapw#3490, diegosouzapw#3451, diegosouzapw#3491, diegosouzapw#3169 under v3.8.18

* feat(docs): add doc accuracy gate + refresh AGENTS.md counts (diegosouzapw#3510)

Integrated into release/v3.8.18

* fix(sse): drop empty-choices chunks without usage instead of injecting retry text (diegosouzapw#3513)

PR diegosouzapw#3422 ('allow OpenAI usage-only empty choices chunks') reintroduced the
assistant-content injection '[OmniRoute] Upstream returned an empty response.
Please retry.' for empty `choices: []` chunks that carry no valid usage. Clients
(Goose/opencode) feed that text back as a turn and spin in a retry loop -- the
exact regression diegosouzapw#3400 had fixed by dropping the chunk.

Restore the drop behavior for the no-usage case while preserving diegosouzapw#3422's
standards-compliant forwarding of usage-only `include_usage` final chunks.
Realign the mislabeled stream-utils test (it asserted the injection) and add a
dedicated regression guard.

Reported-by: @mochizzan
Refs: diegosouzapw#3502, diegosouzapw#3388, diegosouzapw#3400, diegosouzapw#3422

* fix(authz): fall back to URL token when Authorization isn't a usable Bearer (diegosouzapw#3504)

Integrated into release/v3.8.18

* fix(playground): authenticate via session, test key policy by id (diegosouzapw#3503)

Integrated into release/v3.8.18

* docs(changelog): record diegosouzapw#3510, diegosouzapw#3504, diegosouzapw#3503 under v3.8.18

* fix: llama base url normalization (diegosouzapw#3519)

* docs(changelog): reconcile v3.8.18 — add diegosouzapw#3519, diegosouzapw#3513, diegosouzapw#3435-repair, gitignore chore (full commit↔changelog coverage)

* fix(opencode-plugin): bound regex quantifiers in normaliseFreeLabel (polynomial-ReDoS)

CodeQL js/polynomial-redos: unbounded \s* before an anchored \s*$ allowed
O(n²) backtracking on attacker-influenced display names. Bounded to {0,8}/{1,8}
(ample for any real label spacing). Plugin builds + 254 tests green.

* fix(types): restore clean typecheck:core for v3.8.18 release gate

- getPendingRequests() typed to real shape (was widened to object) → fixes
  unknown 'count' in the unified-requests view (diegosouzapw#3401)
- streamChunks log payload cast to its declared type (callLogs.ts)
- preScreenTargets aligned to canonical IsModelAvailable signature (diegosouzapw#3169),
  Promise.resolve-normalized so .catch never hits a bare boolean

All 5 gates green: lint(0 err) + typecheck:core + cycles + docs-all + unit + vitest(146).

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Andrey Borodulin <borodulin@gmail.com>
Co-authored-by: Dmitrii Safronov <zimniy@cyberbrain.cc>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
* chore(release): open v3.8.18 development cycle

* fix(catalog): stop Codex CLI model-catalog refresh from erroring (diegosouzapw#3481)

Codex's model-catalog refresh (codex_models_manager) does
GET /v1/models?client_version=<v> and decodes a JSON object with a
TOP-LEVEL `models` array. OmniRoute answers in the OpenAI-standard
`{object,data}` shape, so codex fails with "missing field `models`"
and logs "failed to refresh available models" on every startup.

Detect codex clients via the `originator` / `user-agent` = `codex_*`
headers they send and add an EMPTY top-level `models: []` so the decode
succeeds. Non-codex OpenAI clients keep the byte-identical `{object,data}`
response.

The array is intentionally empty: codex replaces its built-in per-model
agent prompt (`base_instructions`, ~21k chars) with whatever a populated
entry carries for the selected model, so emitting our catalog would drop
the agent prompt to nothing and break codex's agent behaviour (verified
empirically against codex 0.137). An empty list keeps codex on its
built-in model info — same inference as before, minus the error.

Validated end-to-end with the real handler against codex 0.137:
"failed to refresh available models" → 0 occurrences, instructions
preserved (built-in Codex agent prompt, not empty).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: ignore quality reports and local prompt artifacts

Add generated quality gate reports, metrics files, and local setup prompt
artifacts to .gitignore to prevent committing environment-specific or
temporary files.

* fix(provider): detect Responses API format when body has `input` but … (diegosouzapw#3490)

Integrated into release/v3.8.18

* fix(sse): normalize numeric provider ids to strings (diegosouzapw#3451)

Integrated into release/v3.8.18

* feat(browserPool): resolve Playwright proxy from proxy_registry DB (diegosouzapw#3492)

Integrated into release/v3.8.18

* fix(theoldllm): generate X-Request-Token server-side, drop Playwright (diegosouzapw#3491)

Integrated into release/v3.8.18

* feat(plugins): add lifecycle hooks and theme-manager plugin (diegosouzapw#3473)

Integrated into release/v3.8.18

* fix(combo): parallel pre-screen + circuit-breaker fast-exit for priority combos (diegosouzapw#3169)

Integrated into release/v3.8.18

* feat(ui): unifi active and finished requests into single view diegosouzapw#1422 (diegosouzapw#3401)

Integrated into release/v3.8.18

* docs(changelog): record diegosouzapw#3401, diegosouzapw#3473, diegosouzapw#3492, diegosouzapw#3490, diegosouzapw#3451, diegosouzapw#3491, diegosouzapw#3169 under v3.8.18

* feat(docs): add doc accuracy gate + refresh AGENTS.md counts (diegosouzapw#3510)

Integrated into release/v3.8.18

* fix(sse): drop empty-choices chunks without usage instead of injecting retry text (diegosouzapw#3513)

PR diegosouzapw#3422 ('allow OpenAI usage-only empty choices chunks') reintroduced the
assistant-content injection '[OmniRoute] Upstream returned an empty response.
Please retry.' for empty `choices: []` chunks that carry no valid usage. Clients
(Goose/opencode) feed that text back as a turn and spin in a retry loop -- the
exact regression diegosouzapw#3400 had fixed by dropping the chunk.

Restore the drop behavior for the no-usage case while preserving diegosouzapw#3422's
standards-compliant forwarding of usage-only `include_usage` final chunks.
Realign the mislabeled stream-utils test (it asserted the injection) and add a
dedicated regression guard.

Reported-by: @mochizzan
Refs: diegosouzapw#3502, diegosouzapw#3388, diegosouzapw#3400, diegosouzapw#3422

* fix(authz): fall back to URL token when Authorization isn't a usable Bearer (diegosouzapw#3504)

Integrated into release/v3.8.18

* fix(playground): authenticate via session, test key policy by id (diegosouzapw#3503)

Integrated into release/v3.8.18

* docs(changelog): record diegosouzapw#3510, diegosouzapw#3504, diegosouzapw#3503 under v3.8.18

* fix: llama base url normalization (diegosouzapw#3519)

* docs(changelog): reconcile v3.8.18 — add diegosouzapw#3519, diegosouzapw#3513, diegosouzapw#3435-repair, gitignore chore (full commit↔changelog coverage)

* fix(opencode-plugin): bound regex quantifiers in normaliseFreeLabel (polynomial-ReDoS)

CodeQL js/polynomial-redos: unbounded \s* before an anchored \s*$ allowed
O(n²) backtracking on attacker-influenced display names. Bounded to {0,8}/{1,8}
(ample for any real label spacing). Plugin builds + 254 tests green.

* fix(types): restore clean typecheck:core for v3.8.18 release gate

- getPendingRequests() typed to real shape (was widened to object) → fixes
  unknown 'count' in the unified-requests view (diegosouzapw#3401)
- streamChunks log payload cast to its declared type (callLogs.ts)
- preScreenTargets aligned to canonical IsModelAvailable signature (diegosouzapw#3169),
  Promise.resolve-normalized so .catch never hits a bare boolean

All 5 gates green: lint(0 err) + typecheck:core + cycles + docs-all + unit + vitest(146).

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Andrey Borodulin <borodulin@gmail.com>
Co-authored-by: Dmitrii Safronov <zimniy@cyberbrain.cc>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
- getPendingRequests() typed to real shape (was widened to object) → fixes
  unknown 'count' in the unified-requests view (diegosouzapw#3401)
- streamChunks log payload cast to its declared type (callLogs.ts)
- preScreenTargets aligned to canonical IsModelAvailable signature (diegosouzapw#3169),
  Promise.resolve-normalized so .catch never hits a bare boolean

All 5 gates green: lint(0 err) + typecheck:core + cycles + docs-all + unit + vitest(146).
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
* chore(release): open v3.8.18 development cycle

* fix(catalog): stop Codex CLI model-catalog refresh from erroring (diegosouzapw#3481)

Codex's model-catalog refresh (codex_models_manager) does
GET /v1/models?client_version=<v> and decodes a JSON object with a
TOP-LEVEL `models` array. OmniRoute answers in the OpenAI-standard
`{object,data}` shape, so codex fails with "missing field `models`"
and logs "failed to refresh available models" on every startup.

Detect codex clients via the `originator` / `user-agent` = `codex_*`
headers they send and add an EMPTY top-level `models: []` so the decode
succeeds. Non-codex OpenAI clients keep the byte-identical `{object,data}`
response.

The array is intentionally empty: codex replaces its built-in per-model
agent prompt (`base_instructions`, ~21k chars) with whatever a populated
entry carries for the selected model, so emitting our catalog would drop
the agent prompt to nothing and break codex's agent behaviour (verified
empirically against codex 0.137). An empty list keeps codex on its
built-in model info — same inference as before, minus the error.

Validated end-to-end with the real handler against codex 0.137:
"failed to refresh available models" → 0 occurrences, instructions
preserved (built-in Codex agent prompt, not empty).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: ignore quality reports and local prompt artifacts

Add generated quality gate reports, metrics files, and local setup prompt
artifacts to .gitignore to prevent committing environment-specific or
temporary files.

* fix(provider): detect Responses API format when body has `input` but … (diegosouzapw#3490)

Integrated into release/v3.8.18

* fix(sse): normalize numeric provider ids to strings (diegosouzapw#3451)

Integrated into release/v3.8.18

* feat(browserPool): resolve Playwright proxy from proxy_registry DB (diegosouzapw#3492)

Integrated into release/v3.8.18

* fix(theoldllm): generate X-Request-Token server-side, drop Playwright (diegosouzapw#3491)

Integrated into release/v3.8.18

* feat(plugins): add lifecycle hooks and theme-manager plugin (diegosouzapw#3473)

Integrated into release/v3.8.18

* fix(combo): parallel pre-screen + circuit-breaker fast-exit for priority combos (diegosouzapw#3169)

Integrated into release/v3.8.18

* feat(ui): unifi active and finished requests into single view diegosouzapw#1422 (diegosouzapw#3401)

Integrated into release/v3.8.18

* docs(changelog): record diegosouzapw#3401, diegosouzapw#3473, diegosouzapw#3492, diegosouzapw#3490, diegosouzapw#3451, diegosouzapw#3491, diegosouzapw#3169 under v3.8.18

* feat(docs): add doc accuracy gate + refresh AGENTS.md counts (diegosouzapw#3510)

Integrated into release/v3.8.18

* fix(sse): drop empty-choices chunks without usage instead of injecting retry text (diegosouzapw#3513)

PR diegosouzapw#3422 ('allow OpenAI usage-only empty choices chunks') reintroduced the
assistant-content injection '[OmniRoute] Upstream returned an empty response.
Please retry.' for empty `choices: []` chunks that carry no valid usage. Clients
(Goose/opencode) feed that text back as a turn and spin in a retry loop -- the
exact regression diegosouzapw#3400 had fixed by dropping the chunk.

Restore the drop behavior for the no-usage case while preserving diegosouzapw#3422's
standards-compliant forwarding of usage-only `include_usage` final chunks.
Realign the mislabeled stream-utils test (it asserted the injection) and add a
dedicated regression guard.

Reported-by: @mochizzan
Refs: diegosouzapw#3502, diegosouzapw#3388, diegosouzapw#3400, diegosouzapw#3422

* fix(authz): fall back to URL token when Authorization isn't a usable Bearer (diegosouzapw#3504)

Integrated into release/v3.8.18

* fix(playground): authenticate via session, test key policy by id (diegosouzapw#3503)

Integrated into release/v3.8.18

* docs(changelog): record diegosouzapw#3510, diegosouzapw#3504, diegosouzapw#3503 under v3.8.18

* fix: llama base url normalization (diegosouzapw#3519)

* docs(changelog): reconcile v3.8.18 — add diegosouzapw#3519, diegosouzapw#3513, diegosouzapw#3435-repair, gitignore chore (full commit↔changelog coverage)

* fix(opencode-plugin): bound regex quantifiers in normaliseFreeLabel (polynomial-ReDoS)

CodeQL js/polynomial-redos: unbounded \s* before an anchored \s*$ allowed
O(n²) backtracking on attacker-influenced display names. Bounded to {0,8}/{1,8}
(ample for any real label spacing). Plugin builds + 254 tests green.

* fix(types): restore clean typecheck:core for v3.8.18 release gate

- getPendingRequests() typed to real shape (was widened to object) → fixes
  unknown 'count' in the unified-requests view (diegosouzapw#3401)
- streamChunks log payload cast to its declared type (callLogs.ts)
- preScreenTargets aligned to canonical IsModelAvailable signature (diegosouzapw#3169),
  Promise.resolve-normalized so .catch never hits a bare boolean

All 5 gates green: lint(0 err) + typecheck:core + cycles + docs-all + unit + vitest(146).

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Andrey Borodulin <borodulin@gmail.com>
Co-authored-by: Dmitrii Safronov <zimniy@cyberbrain.cc>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants