Skip to content

Release v3.8.23 - #3693

Merged
diegosouzapw merged 42 commits into
mainfrom
release/v3.8.23
Jun 13, 2026
Merged

diegosouzapw merged 42 commits into
mainfrom
release/v3.8.23

Conversation

@diegosouzapw

@diegosouzapw diegosouzapw commented Jun 12, 2026 •

Copy link
Copy Markdown
Owner

[3.8.23] — TBD

✨ New Features

  • Emergency budget fallback: opt-out env switch OMNIROUTE_EMERGENCY_FALLBACK (#3741 — thanks @zoispag): adds an OMNIROUTE_EMERGENCY_FALLBACK environment variable that disables the budget-exhaustion emergency reroute to nvidia/openai/gpt-oss-120b entirely when set to false or 0. Default behavior (enabled) is unchanged.

  • Auto-Combo: live model intelligence scoring via Arena ELO + models.dev (#3660 — thanks @pizzav-xyz): replaces the static fitness lookup with a 5-layer resolution chain (user override → Arena ELO → models.dev tiers → hardcoded map → neutral fallback). A sync pipeline auto-fetches Arena AI leaderboard ELO scores and derives intelligence tiers from models.dev capabilities; combo picks now update as leaderboard rankings change without any manual configuration.

  • Vertex AI: dynamic model discovery (#3712 — thanks @artickc): the vertex provider now queries the Generative Language models API at runtime to surface the full account catalog — including image-generation models (Imagen, gemini-*-image), embeddings, and partner models — instead of returning only the small hardcoded registry list.

  • Vertex AI: self-tracked USD spend on the Limits page (#3724 — thanks @artickc): since the Google Cloud Billing API is inaccessible via the proxy credential, Vertex connections now track their own cumulative USD spend locally (based on token-cost accounting) and display it on the Limits page as "$ used since account added."

  • Gemini: rate-limit metadata for known per-model RPM/RPD caps (#3686 — thanks @hartmark): injects known rate-limit headers (RPM/RPD) for Gemini models that carry per-model limits (e.g. Gemma 4's 15 RPM / generous RPD), so the cooldown engine applies them correctly instead of locking out the whole account on daily-limit hits.

  • Model Lockout: full settings UI with success-decay recovery (#3629 — thanks @Chewji9875): end-to-end wiring of the per-model lockout feature — settings UI (enable/disable, configure thresholds), backend integration, structured error classification, and a success-decay mechanism that gradually recovers a locked model's fitness as successful calls accumulate.

🔧 Bug Fixes

  • @omniroute/opencode-plugin bundled in the npm tarball + omniroute setup opencode CLI command (#3726 — thanks @herjarsa)
  • MiMoCode 403 "Illegal access" fixed (#3728 — thanks @felipesartori)
  • "Test all models" flow: i18n crash, status icons, auto-hide (#3729 — thanks @felipesartori)
  • OAuth token-refresh invalidation loop fixed (#3692 — thanks @diegosouzapw)
  • safeLogEvents async hotfix (thanks @diegosouzapw)
  • Kiro: quota tracking for IAM Identity Center accounts (#3722 — thanks @artickc)
  • Empty Claude SSE stream now surfaces a real error (#3689 — thanks @TechNickAI)
  • Vertex AI Express-mode API keys (#3690 — thanks @artickc)
  • Anthropic: strip top_p when temperature is set (#3691 — thanks @zhiru)
  • Combo reasoning token buffer: conservative application + feature flag (#3700 — thanks @rdself)
  • Emergency budget fallback: cross-provider credential leak fixed (#3699 — thanks @diegosouzapw)
  • /v1/messages/count_tokens now honors the connection's proxy assignment (#3699 — thanks @diegosouzapw)
  • Gemini: context-mode fallback for signatureless tool calls (#3688 — thanks @diegosouzapw)
  • Antigravity: preserve gemini-3.1-pro High/Low budget tiers (#3696 — thanks @diegosouzapw)
  • Stream combo: fail over on empty/content-filtered response (#3685 — thanks @diegosouzapw)
  • Qwen Web: migrated to v2 chat API (#3723 — thanks @diegosouzapw)
  • WebDAV PUT update: resolve promise on finish not end (thanks @diegosouzapw): eliminates intermittent 500 on second PUT — write stream might not have flushed at renameSync time.
  • Model family fallback lookup: also try literal dot-form key (thanks @diegosouzapw): getNextFamilyFallback normalized gemini-3.1-pro-high → gemini-3-1-pro-high but the map key uses dots. Pre-existing since v3.8.22.

♻️ Code Quality

🌍 Internationalization

  • zh-CN: comprehensive Simplified Chinese translation improvements (#3736 — thanks @sdfsdfw2)

📝 Maintenance

✅ Tests


Quality Gate

  • lint: ✅ 0 errors
  • typecheck:core: ✅ 0 errors
  • check:cycles: ✅ clean
  • check:docs-all: ✅ (warnings are pre-existing drift, non-strict)
  • check-env-doc-sync: ✅ 13/13
  • vitest (MCP + autoCombo): ✅ 171/171
  • unit tests (node:test): ✅ (pre-existing check-docs-symbols failure on research docs, unrelated to this cycle)

Coverage of commits since previous tag

  • Range: v3.8.22..HEAD
  • Commits inspected: 34
  • Changelog bullets: 31 (covers all 34 — 3 are rollup/maintenance commits)

⚠️ After merging: run Phase 2 (Local VPS homologation) before tagging.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@kilo-code-bot

kilo-code-bot Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

Code Review Summary

Status: 2 Issues Found | Recommendation: Address before merge

Overview

Severity Count
CRITICAL 0
WARNING 2
SUGGESTION 0
Issue Details (click to expand)

WARNING

File Line Issue
open-sse/services/emergencyFallback.ts 71 Emergency fallback env parser only disables for lowercase "false" or "0"; case/whitespace variants and accidental empty values remain enabled.
Other Observations (not in diff)

Issues found in unchanged code that cannot receive inline comments:

File Line Issue
tests/integration/combo-provider-exhaustion.test.ts 642 Existing active CodeQL warning: substring check for openai.com can match arbitrary URL segments, so the credential-leak guard may not reliably detect an OpenAI key sent to a non-OpenAI host.
Files Reviewed (1 file)
  • tests/unit/ui/rtl-logical-classes.test.tsx - no issues (file rename only, no code changes)

This PR contains only a file rename operation. All previous issues remain on unchanged files.

Fix these issues in Kilo Cloud


Reviewed by laguna-m.1-20260312:free · 186,978 tokens

@github-actions

github-actions Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

CI Coverage Report

  • Coverage job: skipped
  • PR test policy: success

Coverage artifact was not available for this run.

zhiru and others added 25 commits June 12, 2026 02:40
Integrated into release/v3.8.23
… (#3702)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… (#3703)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
…API (#3712)

Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean.
Integrated into release/v3.8.23. Makes the #3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean.
…54) (#3717)

Phase 1g-1j of #3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix.

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
…55) (#3721)

Phase 1k-1m of #3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
The fix itself reached main pre-tag via cherry-pick #3591, but its changelog
bullet (commit e33fdd4) only ever existed on release/v3.8.20 after the
squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch
leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md).
…177) (#3725)

Phase 1n-1s of #3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
…cation, and success-decay recovery (#3629)

Integrated into release/v3.8.23
…ARGET REACHED ✅) (#3727)

Phase 1t of #3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
…opencode' CLI command (#3726)

Integrated into release/v3.8.23
… every chatHelpers import

#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside
safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed
because typecheck:core does not cover src/sse and no test in the merge gates
loaded chatHelpers via tsx; any consumer that did (chat-context-relay and
chat-route-coverage suites, integration harnesses) failed at module load with
'await can only be used inside an async function'.

safeLogEvents is fire-and-forget logging with an outer try/catch, so making it
async (and 'void'-ing the single chat.ts call site) preserves behavior exactly.

Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts
went from failing-at-load to green (+14 tests destravados).
… + combo/proxy audit fixes (#3699)

Integrated into release/v3.8.23
Comment thread tests/integration/combo-provider-exhaustion.test.ts Fixed
felipesartori and others added 5 commits June 12, 2026 12:05
Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml).
Integrated into release/v3.8.23 — actions/cache v4→v5.
Integrated into release/v3.8.23 — download-artifact v4→v8.
Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes #3739, related #2879.

Integrated into release/v3.8.23.

export function isEmergencyFallbackEnvEnabled(): boolean {
const raw = process.env.OMNIROUTE_EMERGENCY_FALLBACK;
return raw !== "false" && raw !== "0";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

WARNING: Emergency fallback env parser treats most typo/falsey values as enabled

OMNIROUTE_EMERGENCY_FALLBACK is documented as a disable switch, but isEmergencyFallbackEnvEnabled() only rejects lowercase "false" and "0". Values such as "False", " FALSE ", or an accidental empty string keep fallback enabled, so an operator may believe fallback is disabled while the runtime still redirects budget failures.

Suggested change
return raw !== "false" && raw !== "0";
return raw?.trim() !== "" && raw?.trim().toLowerCase() !== "false" && raw?.trim().toLowerCase() !== "0";

Reply with @kilocode-bot fix it to have Kilo Code address this issue.

sdfsdfw2 and others added 10 commits June 12, 2026 20:53
Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green.

Integrated into release/v3.8.23.
- CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits)
- fix(webdav): resolve promise on writeStream finish, not req end — eliminates
  intermittent 500 on PUT update (writeStream may not have flushed at rename time)
- test(autoCombo): stub DB calls from PR #3660 in tieredRotation.test.ts to prevent
  5s timeout in vitest (getModelIntelligenceBySource DB init path)
- chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE
  allowlist (introduced by PR #3726 setup-open-code.mjs, not OmniRoute config vars)
- chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated)
getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" →
"gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The
lookup always missed, returning null for any model whose dots are part of the
name rather than a version separator.

Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both
naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22).
Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule #18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release.

Integrated into release/v3.8.23.
Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23.

Integrated into release/v3.8.23.
Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests.

Integrated into release/v3.8.23.
…3708)

resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected.

Integrated into release/v3.8.23.
- file-size baseline: re-freeze 8 files grown by PRs #3742/#3743/#3740
  (cost drilldown, provider display modes, cache key isolation)
- ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift)
- .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (#3741, env-doc-sync)
- CHANGELOG: add formatted bullets for #3742, #3743, #3708, #3740,
  model-family-fallback fix; remove duplicate raw ### Fixed section
Three test files had net assertion removals after behavior-changing PRs:
- chatcore-translation-paths: emergency fallback moved to routing layer
  (#3699) — add body error assertion + model-name guard
- executor-vertex-extended: non-JSON is now Express API key (#3690) —
  add projects/-path guard to the express-key URL test
- stream-utils: empty streams now emit error (#3685) — add code/message/
  status/completePayload guards to both passthrough and translate variants

All new assertions are meaningful (code enum value, 5xx range, non-empty
message, onComplete must-not-fire contract).
@diegosouzapw
diegosouzapw merged commit de60b4b into main Jun 13, 2026
10 of 11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.