feat(vertex): dynamic model discovery (surface image models + full catalog) - #3712
Conversation
There was a problem hiding this comment.
Code Review
This pull request extracts Gemini model parsing into a reusable utility (parseGeminiModelsList) and introduces dynamic model discovery for Vertex AI and Vertex Partner providers. It also adds unit tests for the new parser. A critical issue was identified in the Vertex AI discovery logic, where non-existent helper functions (isExpressApiKey and looksLikeServiceAccountJson) are imported, which would cause a runtime crash. It is recommended to determine the credential type by attempting to parse it as JSON instead.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| const { isExpressApiKey, looksLikeServiceAccountJson, parseSAFromApiKey, getAccessToken } = | ||
| await import("@omniroute/open-sse/executors/vertex.ts"); | ||
| if (accessToken) { | ||
| bearerToken = accessToken; | ||
| } else if (isExpressApiKey(credential)) { | ||
| queryKey = credential; | ||
| } else if (looksLikeServiceAccountJson(credential)) { | ||
| bearerToken = await getAccessToken(parseSAFromApiKey(credential)); | ||
| } |
There was a problem hiding this comment.
The helper functions isExpressApiKey and looksLikeServiceAccountJson are imported from @omniroute/open-sse/executors/vertex.ts, but they are not exported or defined in that file. This will cause a runtime TypeError and crash the model discovery process for Vertex AI.
Instead of importing these non-existent helpers, we can determine the credential type directly by checking if the credential is valid JSON (which indicates a Service Account JSON) or a plain string (which indicates an Express API key).
const { parseSAFromApiKey, getAccessToken } =
await import("@omniroute/open-sse/executors/vertex.ts");
if (accessToken) {
bearerToken = accessToken;
} else if (credential) {
let isSaJson = false;
try {
JSON.parse(credential);
isSaJson = true;
} catch {}
if (isSaJson) {
bearerToken = await getAccessToken(parseSAFromApiKey(credential));
} else {
queryKey = credential;
}
}
Code Review SummaryStatus: 2 Issues Found | Recommendation: Address before merge Overview
Issue Details (click to expand)WARNING
Other Observations (not in diff)Issues found in unchanged code that cannot receive inline comments:
Files Reviewed (3 files)
Fix these issues in Kilo Cloud Reviewed by nex-n2-pro:free · 1,980,315 tokens |
| const headers: Record<string, string> = { "Content-Type": "application/json" }; | ||
| if (bearerToken) headers["Authorization"] = `Bearer ${bearerToken}`; | ||
|
|
||
| const allModels: any[] = []; |
There was a problem hiding this comment.
WARNING: allModels is typed as any[] instead of GeminiDiscoveryModel[]
parseGeminiModelsList returns GeminiDiscoveryModel[], but the array collecting its results is typed as any[]. This loses type safety for all downstream operations on allModels (e.g., buildApiDiscoveryResponse at line 2142) and prevents the compiler from catching property mismatches if the parser's return type ever changes.
const allModels: GeminiDiscoveryModel[] = [];Reply with @kilocode-bot fix it to have Kilo Code address this issue.
There was a problem hiding this comment.
To use Kilo from GitHub you first need to link your GitHub account to Kilo. Link your Kilo account to continue. After linking, mention me again in this issue or pull request.
The Vertex provider returned only a small hardcoded registry list (10 chat models, no image models) from /v1/models, so image-generation models (imagen-*, gemini-*-image) and the full live catalog were never available. Wire Vertex into the model-discovery pipeline: a dedicated vertex/vertex-partner branch lists models from https://generativelanguage.googleapis.com/v1beta/models, reachable by both Express-mode API keys (?key=) and Service Account JSON (minted OAuth Bearer). Credential type is detected locally (JSON object => Service Account, opaque string => Express key) so the branch has no dependency on optional executor helpers. supportedGenerationMethods are mapped to endpoints via a shared, unit-tested parseGeminiModelsList helper (extracted from the gemini config). Falls back to cached/static catalog on error or missing credentials.
890c643 to
7921241
Compare
| }); | ||
| } | ||
|
|
||
| if (provider === "vertex" || provider === "vertex-partner") { |
There was a problem hiding this comment.
WARNING: vertex-partner dynamic discovery can replace the partner-only catalog
This branch treats vertex-partner the same as vertex and returns buildApiDiscoveryResponse(allModels) when the Generative Language endpoint returns any models. The live endpoint lists Google Gemini/Imagen models, not the partner models in the static vertex-partner registry (open-sse/config/providerRegistry.ts:3458, e.g. DeepSeek-V4-*, Qwen3.6-*, GLM-5.1-FP8, and Claude partner models). A successful discovery response can therefore drop those partner-only models instead of augmenting the local catalog.
Merge discovered models with the local catalog for vertex-partner, or exclude vertex-partner from this branch if it is not served by the Generative Language API.
Reply with @kilocode-bot fix it to have Kilo Code address this issue.
There was a problem hiding this comment.
To use Kilo from GitHub you first need to link your GitHub account to Kilo. Link your Kilo account to continue. After linking, mention me again in this issue or pull request.
|
Thanks @artickc — merged into This is a great follow-up to your Vertex Express-mode work (#3690): Vertex now flows through the standard model-discovery pipeline, so Imagen / Validated locally: |
* chore(release): open v3.8.23 development cycle * fix(anthropic): strip top_p when temperature is set to avoid 400 (#3691) Integrated into release/v3.8.23 * fix(vertex): support Vertex AI Express-mode API keys (#3690) Integrated into release/v3.8.23 * fix(stream): error on empty Claude SSE instead of synthetic success (#3689) Integrated into release/v3.8.23 * fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (#3692) Integrated into release/v3.8.23 * docs: add FUNDING.yml and Support section to README (#3698) Integrated into release/v3.8.23 * feat: gemini - handle known ratelimits (#3686) Integrated into release/v3.8.23 * fix: stream combo fails over on empty content-filtered response (#3685) (#3702) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (#3696) (#3703) Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(auto-combo): add auto-updating model intelligence scoring (#3660) Integrated into release/v3.8.23 * fix(gemini): context-mode fallback for signatureless tool calls (#3688) (#3704) * chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (#3705) * feat(vertex): dynamic model discovery via Generative Language models API (#3712) Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean. * fix(combo): gate reasoning token buffer (#3700) Integrated into release/v3.8.23. Makes the #3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean. * refactor(#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (#3717) Phase 1g-1j of #3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix. Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * refactor(#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (#3721) Phase 1k-1m of #3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * docs(changelog): restore #3590 bullet lost on the v3.8.20 release branch The fix itself reached main pre-tag via cherry-pick #3591, but its changelog bullet (commit e33fdd4) only ever existed on release/v3.8.20 after the squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md). * fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (#3722) Integrated into release/v3.8.23 * refactor(#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (#3725) Phase 1n-1s of #3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (#3629) Integrated into release/v3.8.23 * refactor(#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (#3727) Phase 1t of #3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (#3726) Integrated into release/v3.8.23 * feat(vertex): self-tracked USD spend since account added (#3724) Integrated into release/v3.8.23 * fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (#3288) (#3723) Integrated into release/v3.8.23 * fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import #3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed because typecheck:core does not cover src/sse and no test in the merge gates loaded chatHelpers via tsx; any consumer that did (chat-context-relay and chat-route-coverage suites, integration harnesses) failed at module load with 'await can only be used inside an async function'. safeLogEvents is fire-and-forget logging with an outer try/catch, so making it async (and 'void'-ing the single chat.ts call site) preserves behavior exactly. Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts went from failing-at-load to green (+14 tests destravados). * fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (#3699) Integrated into release/v3.8.23 * fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (#3728) Integrated into release/v3.8.23 * fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (#3729) Integrated into release/v3.8.23 * chore(deps): bump actions/upload-artifact from 4 to 7 (#3735) Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml). * chore(deps): bump actions/cache from 4 to 5 (#3734) Integrated into release/v3.8.23 — actions/cache v4→v5. * chore(deps): bump actions/download-artifact from 4 to 8 (#3733) Integrated into release/v3.8.23 — download-artifact v4→v8. * feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (#3741) Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes #3739, related #2879. Integrated into release/v3.8.23. * i18n: comprehensive zh-CN translation improvements (#3736) Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green. Integrated into release/v3.8.23. * chore(release): v3.8.23 — 2026-06-12 - CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits) - fix(webdav): resolve promise on writeStream finish, not req end — eliminates intermittent 500 on PUT update (writeStream may not have flushed at rename time) - test(autoCombo): stub DB calls from PR #3660 in tieredRotation.test.ts to prevent 5s timeout in vitest (getModelIntelligenceBySource DB init path) - chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE allowlist (introduced by PR #3726 setup-open-code.mjs, not OmniRoute config vars) - chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated) * fix(model-family): fallback lookup also tries bare model name with dots getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" → "gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The lookup always missed, returning null for any model whose dots are part of the name rather than a version separator. Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22). * feat: expose API key cost drilldown + quota % used (#3742) Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule #18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release. Integrated into release/v3.8.23. * feat: add provider display modes — All / Configured / Compact (#3743) Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23. Integrated into release/v3.8.23. * fix(cache): scope semantic-cache signature to API key (#3740) Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests. Integrated into release/v3.8.23. * fix(responses): apply OpenAI Responses API stream=false spec default (#3708) resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected. Integrated into release/v3.8.23. * chore(release): reconcile CI gates for v3.8.23 - file-size baseline: re-freeze 8 files grown by PRs #3742/#3743/#3740 (cost drilldown, provider display modes, cache key isolation) - ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift) - .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (#3741, env-doc-sync) - CHANGELOG: add formatted bullets for #3742, #3743, #3708, #3740, model-family-fallback fix; remove duplicate raw ### Fixed section * test: restore assert count to satisfy check:test-masking gate Three test files had net assertion removals after behavior-changing PRs: - chatcore-translation-paths: emergency fallback moved to routing layer (#3699) — add body error assertion + model-name guard - executor-vertex-extended: non-JSON is now Express API key (#3690) — add projects/-path guard to the express-key URL test - stream-utils: empty streams now emit error (#3685) — add code/message/ status/completePayload guards to both passthrough and translate variants All new assertions are meaningful (code enum value, 5xx range, non-empty message, onComplete must-not-fire contract). * fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it --------- Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com> Co-authored-by: NOXX - Commiter <artur1992123@mail.ru> Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com> Co-authored-by: Markus Hartung <mail@hartmark.se> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com> Co-authored-by: Randi <55005611+rdself@users.noreply.github.com> Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com> Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com> Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com> Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com> Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
* chore(release): open v3.8.23 development cycle * fix(anthropic): strip top_p when temperature is set to avoid 400 (diegosouzapw#3691) Integrated into release/v3.8.23 * fix(vertex): support Vertex AI Express-mode API keys (diegosouzapw#3690) Integrated into release/v3.8.23 * fix(stream): error on empty Claude SSE instead of synthetic success (diegosouzapw#3689) Integrated into release/v3.8.23 * fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (diegosouzapw#3692) Integrated into release/v3.8.23 * docs: add FUNDING.yml and Support section to README (diegosouzapw#3698) Integrated into release/v3.8.23 * feat: gemini - handle known ratelimits (diegosouzapw#3686) Integrated into release/v3.8.23 * fix: stream combo fails over on empty content-filtered response (diegosouzapw#3685) (diegosouzapw#3702) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (diegosouzapw#3696) (diegosouzapw#3703) Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(auto-combo): add auto-updating model intelligence scoring (diegosouzapw#3660) Integrated into release/v3.8.23 * fix(gemini): context-mode fallback for signatureless tool calls (diegosouzapw#3688) (diegosouzapw#3704) * chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (diegosouzapw#3705) * feat(vertex): dynamic model discovery via Generative Language models API (diegosouzapw#3712) Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean. * fix(combo): gate reasoning token buffer (diegosouzapw#3700) Integrated into release/v3.8.23. Makes the diegosouzapw#3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean. * refactor(diegosouzapw#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (diegosouzapw#3717) Phase 1g-1j of diegosouzapw#3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix. Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * refactor(diegosouzapw#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (diegosouzapw#3721) Phase 1k-1m of diegosouzapw#3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * docs(changelog): restore diegosouzapw#3590 bullet lost on the v3.8.20 release branch The fix itself reached main pre-tag via cherry-pick diegosouzapw#3591, but its changelog bullet (commit db04ef2) only ever existed on release/v3.8.20 after the squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md). * fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (diegosouzapw#3722) Integrated into release/v3.8.23 * refactor(diegosouzapw#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (diegosouzapw#3725) Phase 1n-1s of diegosouzapw#3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (diegosouzapw#3629) Integrated into release/v3.8.23 * refactor(diegosouzapw#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (diegosouzapw#3727) Phase 1t of diegosouzapw#3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (diegosouzapw#3726) Integrated into release/v3.8.23 * feat(vertex): self-tracked USD spend since account added (diegosouzapw#3724) Integrated into release/v3.8.23 * fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (diegosouzapw#3288) (diegosouzapw#3723) Integrated into release/v3.8.23 * fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import diegosouzapw#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed because typecheck:core does not cover src/sse and no test in the merge gates loaded chatHelpers via tsx; any consumer that did (chat-context-relay and chat-route-coverage suites, integration harnesses) failed at module load with 'await can only be used inside an async function'. safeLogEvents is fire-and-forget logging with an outer try/catch, so making it async (and 'void'-ing the single chat.ts call site) preserves behavior exactly. Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts went from failing-at-load to green (+14 tests destravados). * fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (diegosouzapw#3699) Integrated into release/v3.8.23 * fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (diegosouzapw#3728) Integrated into release/v3.8.23 * fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (diegosouzapw#3729) Integrated into release/v3.8.23 * chore(deps): bump actions/upload-artifact from 4 to 7 (diegosouzapw#3735) Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml). * chore(deps): bump actions/cache from 4 to 5 (diegosouzapw#3734) Integrated into release/v3.8.23 — actions/cache v4→v5. * chore(deps): bump actions/download-artifact from 4 to 8 (diegosouzapw#3733) Integrated into release/v3.8.23 — download-artifact v4→v8. * feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (diegosouzapw#3741) Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes diegosouzapw#3739, related diegosouzapw#2879. Integrated into release/v3.8.23. * i18n: comprehensive zh-CN translation improvements (diegosouzapw#3736) Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green. Integrated into release/v3.8.23. * chore(release): v3.8.23 — 2026-06-12 - CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits) - fix(webdav): resolve promise on writeStream finish, not req end — eliminates intermittent 500 on PUT update (writeStream may not have flushed at rename time) - test(autoCombo): stub DB calls from PR diegosouzapw#3660 in tieredRotation.test.ts to prevent 5s timeout in vitest (getModelIntelligenceBySource DB init path) - chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE allowlist (introduced by PR diegosouzapw#3726 setup-open-code.mjs, not OmniRoute config vars) - chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated) * fix(model-family): fallback lookup also tries bare model name with dots getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" → "gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The lookup always missed, returning null for any model whose dots are part of the name rather than a version separator. Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22). * feat: expose API key cost drilldown + quota % used (diegosouzapw#3742) Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule diegosouzapw#18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release. Integrated into release/v3.8.23. * feat: add provider display modes — All / Configured / Compact (diegosouzapw#3743) Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23. Integrated into release/v3.8.23. * fix(cache): scope semantic-cache signature to API key (diegosouzapw#3740) Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests. Integrated into release/v3.8.23. * fix(responses): apply OpenAI Responses API stream=false spec default (diegosouzapw#3708) resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected. Integrated into release/v3.8.23. * chore(release): reconcile CI gates for v3.8.23 - file-size baseline: re-freeze 8 files grown by PRs diegosouzapw#3742/diegosouzapw#3743/diegosouzapw#3740 (cost drilldown, provider display modes, cache key isolation) - ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift) - .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (diegosouzapw#3741, env-doc-sync) - CHANGELOG: add formatted bullets for diegosouzapw#3742, diegosouzapw#3743, diegosouzapw#3708, diegosouzapw#3740, model-family-fallback fix; remove duplicate raw ### Fixed section * test: restore assert count to satisfy check:test-masking gate Three test files had net assertion removals after behavior-changing PRs: - chatcore-translation-paths: emergency fallback moved to routing layer (diegosouzapw#3699) — add body error assertion + model-name guard - executor-vertex-extended: non-JSON is now Express API key (diegosouzapw#3690) — add projects/-path guard to the express-key URL test - stream-utils: empty streams now emit error (diegosouzapw#3685) — add code/message/ status/completePayload guards to both passthrough and translate variants All new assertions are meaningful (code enum value, 5xx range, non-empty message, onComplete must-not-fire contract). * fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it --------- Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com> Co-authored-by: NOXX - Commiter <artur1992123@mail.ru> Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com> Co-authored-by: Markus Hartung <mail@hartmark.se> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com> Co-authored-by: Randi <55005611+rdself@users.noreply.github.com> Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com> Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com> Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com> Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com> Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
* chore(release): open v3.8.23 development cycle * fix(anthropic): strip top_p when temperature is set to avoid 400 (diegosouzapw#3691) Integrated into release/v3.8.23 * fix(vertex): support Vertex AI Express-mode API keys (diegosouzapw#3690) Integrated into release/v3.8.23 * fix(stream): error on empty Claude SSE instead of synthetic success (diegosouzapw#3689) Integrated into release/v3.8.23 * fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (diegosouzapw#3692) Integrated into release/v3.8.23 * docs: add FUNDING.yml and Support section to README (diegosouzapw#3698) Integrated into release/v3.8.23 * feat: gemini - handle known ratelimits (diegosouzapw#3686) Integrated into release/v3.8.23 * fix: stream combo fails over on empty content-filtered response (diegosouzapw#3685) (diegosouzapw#3702) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (diegosouzapw#3696) (diegosouzapw#3703) Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(auto-combo): add auto-updating model intelligence scoring (diegosouzapw#3660) Integrated into release/v3.8.23 * fix(gemini): context-mode fallback for signatureless tool calls (diegosouzapw#3688) (diegosouzapw#3704) * chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (diegosouzapw#3705) * feat(vertex): dynamic model discovery via Generative Language models API (diegosouzapw#3712) Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean. * fix(combo): gate reasoning token buffer (diegosouzapw#3700) Integrated into release/v3.8.23. Makes the diegosouzapw#3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean. * refactor(diegosouzapw#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (diegosouzapw#3717) Phase 1g-1j of diegosouzapw#3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix. Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * refactor(diegosouzapw#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (diegosouzapw#3721) Phase 1k-1m of diegosouzapw#3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * docs(changelog): restore diegosouzapw#3590 bullet lost on the v3.8.20 release branch The fix itself reached main pre-tag via cherry-pick diegosouzapw#3591, but its changelog bullet (commit e33fdd4) only ever existed on release/v3.8.20 after the squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md). * fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (diegosouzapw#3722) Integrated into release/v3.8.23 * refactor(diegosouzapw#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (diegosouzapw#3725) Phase 1n-1s of diegosouzapw#3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (diegosouzapw#3629) Integrated into release/v3.8.23 * refactor(diegosouzapw#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (diegosouzapw#3727) Phase 1t of diegosouzapw#3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (diegosouzapw#3726) Integrated into release/v3.8.23 * feat(vertex): self-tracked USD spend since account added (diegosouzapw#3724) Integrated into release/v3.8.23 * fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (diegosouzapw#3288) (diegosouzapw#3723) Integrated into release/v3.8.23 * fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import diegosouzapw#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed because typecheck:core does not cover src/sse and no test in the merge gates loaded chatHelpers via tsx; any consumer that did (chat-context-relay and chat-route-coverage suites, integration harnesses) failed at module load with 'await can only be used inside an async function'. safeLogEvents is fire-and-forget logging with an outer try/catch, so making it async (and 'void'-ing the single chat.ts call site) preserves behavior exactly. Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts went from failing-at-load to green (+14 tests destravados). * fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (diegosouzapw#3699) Integrated into release/v3.8.23 * fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (diegosouzapw#3728) Integrated into release/v3.8.23 * fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (diegosouzapw#3729) Integrated into release/v3.8.23 * chore(deps): bump actions/upload-artifact from 4 to 7 (diegosouzapw#3735) Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml). * chore(deps): bump actions/cache from 4 to 5 (diegosouzapw#3734) Integrated into release/v3.8.23 — actions/cache v4→v5. * chore(deps): bump actions/download-artifact from 4 to 8 (diegosouzapw#3733) Integrated into release/v3.8.23 — download-artifact v4→v8. * feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (diegosouzapw#3741) Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes diegosouzapw#3739, related diegosouzapw#2879. Integrated into release/v3.8.23. * i18n: comprehensive zh-CN translation improvements (diegosouzapw#3736) Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green. Integrated into release/v3.8.23. * chore(release): v3.8.23 — 2026-06-12 - CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits) - fix(webdav): resolve promise on writeStream finish, not req end — eliminates intermittent 500 on PUT update (writeStream may not have flushed at rename time) - test(autoCombo): stub DB calls from PR diegosouzapw#3660 in tieredRotation.test.ts to prevent 5s timeout in vitest (getModelIntelligenceBySource DB init path) - chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE allowlist (introduced by PR diegosouzapw#3726 setup-open-code.mjs, not OmniRoute config vars) - chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated) * fix(model-family): fallback lookup also tries bare model name with dots getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" → "gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The lookup always missed, returning null for any model whose dots are part of the name rather than a version separator. Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22). * feat: expose API key cost drilldown + quota % used (diegosouzapw#3742) Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule diegosouzapw#18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release. Integrated into release/v3.8.23. * feat: add provider display modes — All / Configured / Compact (diegosouzapw#3743) Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23. Integrated into release/v3.8.23. * fix(cache): scope semantic-cache signature to API key (diegosouzapw#3740) Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests. Integrated into release/v3.8.23. * fix(responses): apply OpenAI Responses API stream=false spec default (diegosouzapw#3708) resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected. Integrated into release/v3.8.23. * chore(release): reconcile CI gates for v3.8.23 - file-size baseline: re-freeze 8 files grown by PRs diegosouzapw#3742/diegosouzapw#3743/diegosouzapw#3740 (cost drilldown, provider display modes, cache key isolation) - ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift) - .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (diegosouzapw#3741, env-doc-sync) - CHANGELOG: add formatted bullets for diegosouzapw#3742, diegosouzapw#3743, diegosouzapw#3708, diegosouzapw#3740, model-family-fallback fix; remove duplicate raw ### Fixed section * test: restore assert count to satisfy check:test-masking gate Three test files had net assertion removals after behavior-changing PRs: - chatcore-translation-paths: emergency fallback moved to routing layer (diegosouzapw#3699) — add body error assertion + model-name guard - executor-vertex-extended: non-JSON is now Express API key (diegosouzapw#3690) — add projects/-path guard to the express-key URL test - stream-utils: empty streams now emit error (diegosouzapw#3685) — add code/message/ status/completePayload guards to both passthrough and translate variants All new assertions are meaningful (code enum value, 5xx range, non-empty message, onComplete must-not-fire contract). * fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it --------- Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com> Co-authored-by: NOXX - Commiter <artur1992123@mail.ru> Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com> Co-authored-by: Markus Hartung <mail@hartmark.se> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com> Co-authored-by: Randi <55005611+rdself@users.noreply.github.com> Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com> Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com> Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com> Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com> Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
…API (diegosouzapw#3712) Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean.
* chore(release): open v3.8.23 development cycle * fix(anthropic): strip top_p when temperature is set to avoid 400 (diegosouzapw#3691) Integrated into release/v3.8.23 * fix(vertex): support Vertex AI Express-mode API keys (diegosouzapw#3690) Integrated into release/v3.8.23 * fix(stream): error on empty Claude SSE instead of synthetic success (diegosouzapw#3689) Integrated into release/v3.8.23 * fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (diegosouzapw#3692) Integrated into release/v3.8.23 * docs: add FUNDING.yml and Support section to README (diegosouzapw#3698) Integrated into release/v3.8.23 * feat: gemini - handle known ratelimits (diegosouzapw#3686) Integrated into release/v3.8.23 * fix: stream combo fails over on empty content-filtered response (diegosouzapw#3685) (diegosouzapw#3702) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (diegosouzapw#3696) (diegosouzapw#3703) Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(auto-combo): add auto-updating model intelligence scoring (diegosouzapw#3660) Integrated into release/v3.8.23 * fix(gemini): context-mode fallback for signatureless tool calls (diegosouzapw#3688) (diegosouzapw#3704) * chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (diegosouzapw#3705) * feat(vertex): dynamic model discovery via Generative Language models API (diegosouzapw#3712) Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean. * fix(combo): gate reasoning token buffer (diegosouzapw#3700) Integrated into release/v3.8.23. Makes the diegosouzapw#3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean. * refactor(diegosouzapw#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (diegosouzapw#3717) Phase 1g-1j of diegosouzapw#3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix. Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * refactor(diegosouzapw#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (diegosouzapw#3721) Phase 1k-1m of diegosouzapw#3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * docs(changelog): restore diegosouzapw#3590 bullet lost on the v3.8.20 release branch The fix itself reached main pre-tag via cherry-pick diegosouzapw#3591, but its changelog bullet (commit a6b99843f) only ever existed on release/v3.8.20 after the squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md). * fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (diegosouzapw#3722) Integrated into release/v3.8.23 * refactor(diegosouzapw#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (diegosouzapw#3725) Phase 1n-1s of diegosouzapw#3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (diegosouzapw#3629) Integrated into release/v3.8.23 * refactor(diegosouzapw#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (diegosouzapw#3727) Phase 1t of diegosouzapw#3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%). Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> * fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (diegosouzapw#3726) Integrated into release/v3.8.23 * feat(vertex): self-tracked USD spend since account added (diegosouzapw#3724) Integrated into release/v3.8.23 * fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (diegosouzapw#3288) (diegosouzapw#3723) Integrated into release/v3.8.23 * fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import diegosouzapw#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed because typecheck:core does not cover src/sse and no test in the merge gates loaded chatHelpers via tsx; any consumer that did (chat-context-relay and chat-route-coverage suites, integration harnesses) failed at module load with 'await can only be used inside an async function'. safeLogEvents is fire-and-forget logging with an outer try/catch, so making it async (and 'void'-ing the single chat.ts call site) preserves behavior exactly. Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts went from failing-at-load to green (+14 tests destravados). * fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (diegosouzapw#3699) Integrated into release/v3.8.23 * fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (diegosouzapw#3728) Integrated into release/v3.8.23 * fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (diegosouzapw#3729) Integrated into release/v3.8.23 * chore(deps): bump actions/upload-artifact from 4 to 7 (diegosouzapw#3735) Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml). * chore(deps): bump actions/cache from 4 to 5 (diegosouzapw#3734) Integrated into release/v3.8.23 — actions/cache v4→v5. * chore(deps): bump actions/download-artifact from 4 to 8 (diegosouzapw#3733) Integrated into release/v3.8.23 — download-artifact v4→v8. * feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (diegosouzapw#3741) Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes diegosouzapw#3739, related diegosouzapw#2879. Integrated into release/v3.8.23. * i18n: comprehensive zh-CN translation improvements (diegosouzapw#3736) Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green. Integrated into release/v3.8.23. * chore(release): v3.8.23 — 2026-06-12 - CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits) - fix(webdav): resolve promise on writeStream finish, not req end — eliminates intermittent 500 on PUT update (writeStream may not have flushed at rename time) - test(autoCombo): stub DB calls from PR diegosouzapw#3660 in tieredRotation.test.ts to prevent 5s timeout in vitest (getModelIntelligenceBySource DB init path) - chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE allowlist (introduced by PR diegosouzapw#3726 setup-open-code.mjs, not OmniRoute config vars) - chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated) * fix(model-family): fallback lookup also tries bare model name with dots getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" → "gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The lookup always missed, returning null for any model whose dots are part of the name rather than a version separator. Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22). * feat: expose API key cost drilldown + quota % used (diegosouzapw#3742) Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule diegosouzapw#18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release. Integrated into release/v3.8.23. * feat: add provider display modes — All / Configured / Compact (diegosouzapw#3743) Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23. Integrated into release/v3.8.23. * fix(cache): scope semantic-cache signature to API key (diegosouzapw#3740) Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests. Integrated into release/v3.8.23. * fix(responses): apply OpenAI Responses API stream=false spec default (diegosouzapw#3708) resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected. Integrated into release/v3.8.23. * chore(release): reconcile CI gates for v3.8.23 - file-size baseline: re-freeze 8 files grown by PRs diegosouzapw#3742/diegosouzapw#3743/diegosouzapw#3740 (cost drilldown, provider display modes, cache key isolation) - ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift) - .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (diegosouzapw#3741, env-doc-sync) - CHANGELOG: add formatted bullets for diegosouzapw#3742, diegosouzapw#3743, diegosouzapw#3708, diegosouzapw#3740, model-family-fallback fix; remove duplicate raw ### Fixed section * test: restore assert count to satisfy check:test-masking gate Three test files had net assertion removals after behavior-changing PRs: - chatcore-translation-paths: emergency fallback moved to routing layer (diegosouzapw#3699) — add body error assertion + model-name guard - executor-vertex-extended: non-JSON is now Express API key (diegosouzapw#3690) — add projects/-path guard to the express-key URL test - stream-utils: empty streams now emit error (diegosouzapw#3685) — add code/message/ status/completePayload guards to both passthrough and translate variants All new assertions are meaningful (code enum value, 5xx range, non-empty message, onComplete must-not-fire contract). * fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it --------- Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com> Co-authored-by: NOXX - Commiter <artur1992123@mail.ru> Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com> Co-authored-by: Markus Hartung <mail@hartmark.se> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com> Co-authored-by: Randi <55005611+rdself@users.noreply.github.com> Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com> Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com> Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com> Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com> Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com> Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
Problem
The
vertexprovider returned only its small hardcoded registry list from/v1/models(10 chat models, e.g.gemini-3-flash-preview,gemini-3.1-flash-lite, plus a few partner models) — and no image-generation models at all. Vertex AI accounts expose far more (Imagen,gemini-*-image, embeddings, etc.), but those were never discoverable becausevertexwas not wired into the model-discovery pipeline (PROVIDER_MODELS_CONFIG/ per-provider discovery branches). It fell through to the staticlocal_catalog.Fix
Add a dedicated
vertex/vertex-partnerbranch to the model-discovery route (GET /api/providers/[id]/models) that lists models from the Google Generative Language endpoint:Both Vertex auth modes are supported (reusing the executor's existing credential helpers):
?key=query param (verified: returns the full live catalog, 55 models incl.gemini-3-pro-image,gemini-2.5-flash-image,imagen-4.0-*).parseSAFromApiKey→getAccessToken) and sent asAuthorization: Bearer.Each model's
supportedGenerationMethodsis mapped to OmniRoute endpoints via a shared, unit-testedparseGeminiModelsListhelper (extracted from the existinggeminiconfig so both share one implementation):generateContent/generateAnswerchatpredict/predictLongRunningimages(Imagen)embedContentembeddingsbidiGenerateContentaudioDiscovered models flow through the standard
buildApiDiscoveryResponse→persistDiscoveredModelspath (pagination vianextPageToken), so the unified catalog auto-swaps the static list for the live one. On any error or missing/unusable credential it falls back to the cached then static catalog (no regression for existing Vertex connections).Tests
tests/unit/gemini-models-parser.test.ts(5 tests, all pass) covers the method→endpoint mapping, themodels/prefix stripping + display-name handling, image models (gemini-*-image→ chat,imagen-*→ images), embeddings/audio, and empty/missing-input tolerance.Notes
providerSpecificData.autoSync); this PR only makes Vertex capable of dynamic discovery.geminiprovider behavior unchanged (it now callsparseGeminiModelsList).