Skip to content

fix: add reasoning token buffer for combo routing (fixes #3587) - #3588

Merged
diegosouzapw merged 3 commits into
diegosouzapw:mainfrom
herjarsa:fix/normalize-owned-by
Jun 11, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:mainfrom
herjarsa:fix/normalize-owned-by

Conversation

@herjarsa

Copy link
Copy Markdown
Contributor

Summary

This PR fixes issue #3587 where reasoning models (deepseek-v4-flash, nemotron, etc.) consume ALL max_tokens for reasoning_tokens, leaving content empty.

Problem

When using combos that include reasoning models, the models consume all allocated tokens for , resulting in empty field. This causes OpenCode to treat the response as failed and return empty responses to clients.

Solution

  1. Detect reasoning models via from
  2. Add token buffer (50% + 1000 tokens) to before sending requests to reasoning models
  3. This ensures reasoning models have enough tokens for both reasoning and content output

Changes

  • :
    • Import and
    • In : Added check for reasoning-only empty content (content empty but reasoning_content present with reasoning consuming 90%+ of tokens)
    • In main combo loop: Added max_tokens buffer before call
    • In round-robin handler: Added same buffer before call

Testing

  • TypeScript compilation passes (no errors in combo.ts)
  • All existing tests pass
  • The fix is backward compatible - non-reasoning models are unaffected

Related

@herjarsa
herjarsa requested a review from diegosouzapw as a code owner June 10, 2026 17:59

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces configuration variables for global provider cooldown tracking in .env.example and performs extensive formatting and whitespace normalization across numerous source, test, and documentation files. However, the reviewer noted a critical issue: the PR title indicates a fix for a reasoning token buffer in combo routing, but the actual implementation files (such as combo.ts) are completely missing from the changes. Additionally, the reviewer recommended correcting the variable names referenced in the .env.example comments for provider cooldown tracking to accurately match the newly defined environment variables (PROVIDER_COOLDOWN_MIN_MS and PROVIDER_COOLDOWN_MAX_MS).

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread .env.example
# Accepted values: true|1|on (force on), false|0|off (force off), unset (use Dashboard).
# RATE_LIMIT_AUTO_ENABLE=

# Provider cooldown tracking: minimum time (ms) before a failed provider/connection

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

This pull request is titled 'fix: add reasoning token buffer for combo routing (fixes #3587)', but the actual implementation changes to combo.ts (and any other backend files) are completely missing from the diff. Instead, the PR contains unrelated changes such as provider cooldown tracking configuration in .env.example, extensive formatting updates across multiple files, and unrelated changelog entries. Please verify your branch and ensure that only the relevant changes for #3587 are committed, and that the actual fix is included.

Comment thread .env.example
Comment on lines +1174 to +1178
# Provider cooldown tracking: minimum time (ms) before a failed provider/connection
# can be retried. Prevents subsequent requests from re-walking failing providers.
# Scaled exponentially: minCooldown * 2^(failures-1), capped at maxRetryCooldownMs.
# Used by: open-sse/services/providerCooldownTracker.ts
# PROVIDER_COOLDOWN_MIN_MS=5000

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The comment refers to minCooldown and maxRetryCooldownMs, but the actual environment variables defined here are PROVIDER_COOLDOWN_MIN_MS and PROVIDER_COOLDOWN_MAX_MS. Updating the comment to use the correct variable names will improve clarity and prevent confusion.

# Provider cooldown tracking: minimum time (ms) before a failed provider/connection
# can be retried. Prevents subsequent requests from re-walking failing providers.
# Scaled exponentially: PROVIDER_COOLDOWN_MIN_MS * 2^(failures-1), capped at PROVIDER_COOLDOWN_MAX_MS.
# Used by: open-sse/services/providerCooldownTracker.ts
# PROVIDER_COOLDOWN_MIN_MS=5000

…#3587)

Reasoning models (deepseek-v4-flash, nemotron, etc.) consume ALL max_tokens
for reasoning_tokens, leaving content empty. The fix:

1. Detects reasoning models via supportsReasoning() before sending requests
2. Adds 50% + 1000 token buffer to max_tokens for reasoning models
3. Detects reasoning-only empty content in validateResponseQuality()
4. Applies buffer in both main combo loop and round-robin handler
…buffer tests

- 6 unit tests for validateResponseQuality:
  - reasoning consumed 90%+ tokens → invalid
  - reasoning consumed < 90% tokens → valid
  - no usage data → valid
  - content present + reasoning + tokens exhausted → valid
  - completion_tokens_details.reasoning_tokens → invalid
  - completion_tokens=0 → safe (no division by zero)
- 2 integration tests for handleComboChat:
  - reasoning model gets max_tokens buffer (4096 → 6144)
  - non-reasoning model does not get buffer (4096 → 4096)
…ound-robin

handleRoundRobinCombo mutated the shared `body` to apply the diegosouzapw#3587 reasoning
buffer, so each reasoning model in the rotation (and each retry) re-read an
already-buffered max_tokens and multiplied it again (4096 -> 6144 -> 9216 -> ...),
eventually overshooting the model's real limit and triggering 400s. Apply the
buffer to a per-attempt copy instead, anchored on the original max_tokens.
handleComboChat was already safe (it clones attemptBody per attempt).

Adds a round-robin regression test asserting two reasoning models both buffer
from the original 4096 (6144), not [6144, 9216].

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
@diegosouzapw
diegosouzapw merged commit 4bfd9e2 into diegosouzapw:main Jun 11, 2026
2 checks passed
diegosouzapw pushed a commit that referenced this pull request Jun 11, 2026
@kilo-code-bot

kilo-code-bot Bot commented Jun 11, 2026 •

Copy link
Copy Markdown

Code Review Summary

Status: No Issues Found | Recommendation: Merge

Overview

The PR addresses Issue #3587 by handling reasoning models that consume all allocated tokens for reasoning, leaving content empty. The solution includes:

  1. Response validation (validateResponseQuality): Detects when reasoning consumed 90%+ of completion tokens and marks the response as invalid, triggering combo fallback
  2. Token buffer (handleComboChat & handleRoundRobinCombo): Adds a 50% buffer + 1000 floor to max_tokens for reasoning models
  3. No compounding in round-robin: The RR loop correctly reads from the original body object for buffer calculation, preventing exponential token growth across iterations

Code Quality Assessment

  • The supportsReasoning import is correctly added to combo.ts
  • Both priority and round-robin strategies apply the buffer consistently
  • Round-robin correctly creates a per-attempt copy (attemptBody = { ...body, max_tokens: buffered }) to avoid mutating shared state
  • Tests cover all key scenarios: token exhaustion detection, normal reasoning, missing usage data, and compounding prevention
  • The 90% threshold is well-chosen to detect true exhaustion while allowing some reasoning room
Files Reviewed (3 files)
  • open-sse/services/combo.ts — Core combo routing logic with reasoning token handling
  • tests/unit/combo-quality-validator-reasoning.test.ts — Response validation tests
  • tests/unit/combo-routing-engine.test.ts — Routing engine tests including buffer application

Reviewed by laguna-m.1-20260312:free · 1,301,832 tokens

@diegosouzapw

Copy link
Copy Markdown
Owner

Merged into release/v3.8.21 — thank you, @herjarsa! 🎉

This fixes #3587 (reasoning models like deepseek-v4-flash/nemotron consuming all of max_tokens on reasoning and returning empty content). Your two-part approach is solid: rejecting empty-content-but-reasoning-exhausted responses in validateResponseQuality so the combo loop retries/falls back, plus the max_tokens buffer for reasoning models. The TDD coverage was thorough.

One small adjustment was made on top (co-authored): in handleRoundRobinCombo the buffer mutated the shared body, so with 2+ reasoning models in a rotation (or retries) it compounded — 4096 → 6144 → 9216 → ... — eventually overshooting the model's real limit. Switched it to apply the buffer to a per-attempt copy anchored on the original max_tokens (matching what handleComboChat already did with attemptBody), and added a round-robin regression test. Ships in v3.8.21!

diegosouzapw added a commit that referenced this pull request Jun 11, 2026
…o reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@diegosouzapw diegosouzapw mentioned this pull request Jun 11, 2026
diegosouzapw added a commit that referenced this pull request Jun 11, 2026
* chore(release): open v3.8.21 development cycle

* fix: pass through valid max_tokens-truncated responses instead of fake 502 (#3572) (#3595)

* fix: /v1/completions returns legacy text-completion format, not chat (#3571) (#3596)

* fix: z.ai/GLM coding plan no longer shows Monthly 0% when no monthly cap (#3580) (#3597)

* docs: mark DISCOVERY_TOOL_DESIGN endpoints as Phase-2 not-yet-implemented (#3498) (#3599)

* fix(agent-bridge): add validate-only upstream-ca/test route (#3488) (#3600)

* fix(gamification): add level/badges/badges-earned profile routes (#3484)

* security(oauth): migrate 5 public client_ids to resolvePublicCred (#3493)

* fix(mcp): ship MCP server source closure in npm files + coverage gate (#3578)

* fix: add reasoning token buffer for combo routing (fixes #3587) (#3588)

Integrated into release/v3.8.21

* Refactor: Extract chatCore phases into modular files (#3598)

Integrated into release/v3.8.21 — chatCore phase modularization. Adjusted: re-derive idempotencyKey for the save path after the check moved into the module (co-authored). Thanks @oyi77!

* docs(changelog): credit #3598 (chatCore modularization) + #3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(api): implement GET /api/guardrails + POST /api/guardrails/test, drop shadow/guardrails doc-fiction (#3496) (#3602)

Integrated into release/v3.8.21 — implements GET /api/guardrails + POST /api/guardrails/test, removes shadow/guardrails doc-fiction. TDD-validated (5/5) + check-docs-symbols/typecheck/eslint green.

* fix(gemini): isolate textual reasoning wrappers (#3605)

Split-out PR C from #3584. Isolates textual reasoning wrappers (<think>/<thinking>/<thought>/<internal_thought>, including malformed/open tags) into reasoning_content across both the non-streaming sanitizer and the Gemini streaming translator, with split-chunk buffering. Additive to the existing textual tool-call pipeline; does not touch the #3569 native functionResponse path. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(antigravity): normalize Gemini 3.5 Flash tier IDs (#3603)

Split-out PR A from #3584. Normalizes the Antigravity/agy Gemini 3.5 Flash tier IDs to clean public names (gemini-3.5-flash-low/medium/high), maps them to the live upstream IDs at the executor boundary, and removes Antigravity from the global model resolver so the executor owns wire normalization. Maintainer follow-up: kept gemini-3.5-flash-preview as a hidden backward-compat alias routing to the High tier (so saved combos/configs keep working). Live-validated the tier set via the agy CLI catalog. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(agent-bridge): surface real MITM startup-failure cause, not always port 443 (#3606) (#3608)

Integrated into release/v3.8.21 (#3606)

* fix(oauth): surface real Kiro import-token failure cause, not a bare 500 (#3589) (#3609)

Integrated into release/v3.8.21 (#3589)

* docs(opencode-provider): soft-deprecate in favor of @omniroute/opencode-plugin (#3419) (#3613)

Integrated into release/v3.8.21 (#3419)

* fix(usage): normalize Antigravity and agy provider quotas (#3604)

Split-out PR B from #3584. Normalizes Antigravity/agy provider quotas: prefers retrieveUserQuota for live consumption, falls back to fetchAvailableModels and local usage_history, sanitizes cached Provider Limits so retired upstream IDs are not re-exposed, and schedules a deduplicated post-usage refresh. Maintainer follow-up: decoupled the post-usage refresh via a lightweight usageEvents bus (usageHistory no longer dynamic-imports providerLimits) so it does not pull the executors/translator graph into the typecheck-core surface — typecheck:core stays at 0. Integrated into release/v3.8.21. Thanks @dhaern!

* feat(cli): add autostart on/off/toggle shorthand for headless serve mode (#3331) (#3614)

Integrated into release/v3.8.21 (#3331)

* docs(changelog): credit #3603 (Flash tier IDs) + #3604 (provider quotas) + #3605 (reasoning wrappers)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(review): resolve findings from /review-reviews battery (v3.8.21 hardening) (#3618)

Pre-release hardening from the /review-reviews battery — 15 findings resolved (L1-L13,L15) + L14 live-verified WONTFIX, convergence re-review clean. lint/typecheck:core/test:vitest(146)/build green; zero new test:unit failures vs baseline 797de43.

* chore(release): v3.8.21 CHANGELOG + i18n + env-doc sync

---------

Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Raxxoor <manker_lol@hotmail.com>
diegosouzapw pushed a commit that referenced this pull request Jun 12, 2026
Integrated into release/v3.8.23. Makes the #3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean.
diegosouzapw added a commit to Chewji9875/OmniRoute that referenced this pull request Jun 12, 2026
…ockout surgical

Conflict-resolution repair against release/v3.8.23:
- combo.ts: rebuilt from the release base preserving diegosouzapw#3588/diegosouzapw#3700 reasoning-buffer
  and diegosouzapw#3556 providerCooldown; re-applied only the lockout hunks (pre-check skip,
  quality-failure record, success decay, first-transient record, final-failure
  record alongside provider cooldown).
- Restored to base: package.json (BUILD_SHA + undici), prepublish.ts, gemini
  parser/route/test, phase1f, ProviderDetailPageClient, ApiManagerPageClient,
  ExternalLinkModal, schemas.ts (diegosouzapw#3537 budget), ResilienceTab (ProviderCooldownCard
  + degradationThreshold) and settingsSchemas (Hard Rules diegosouzapw#15/diegosouzapw#17 spawn guard) —
  re-applying only the ModelLockoutCard wiring, modelLockout schema and i18n keys.
- Implemented the spec the new tests demanded: backoff cap (BACKOFF_CONFIG.max /
  options.maxCooldownMs), canonical provider aliases in lock keys (cx → codex),
  escalation window extended past the applied cooldown, and in-memory lockModel
  on the persistUnavailableState=false path (combo transient 429).
- Removed the 7 unregistered CLI commands (registry.mjs is generated from the
  OpenAPI spec; shipping them unwired would be dead code — they belong in their
  own PR via the generator).

Validation: 259/259 unit tests green (lockout + auth + full combo suites),
typecheck:core clean, eslint 0 errors.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
diegosouzapw added a commit to Chewji9875/OmniRoute that referenced this pull request Jun 12, 2026
…ockout surgical

Conflict-resolution repair against release/v3.8.23:
- combo.ts: rebuilt from the release base preserving diegosouzapw#3588/diegosouzapw#3700 reasoning-buffer
  and diegosouzapw#3556 providerCooldown; re-applied only the lockout hunks (pre-check skip,
  quality-failure record, success decay, first-transient record, final-failure
  record alongside provider cooldown).
- Restored to base: package.json (BUILD_SHA + undici), prepublish.ts, gemini
  parser/route/test, phase1f, ProviderDetailPageClient, ApiManagerPageClient,
  ExternalLinkModal, schemas.ts (diegosouzapw#3537 budget), ResilienceTab (ProviderCooldownCard
  + degradationThreshold) and settingsSchemas (Hard Rules diegosouzapw#15/diegosouzapw#17 spawn guard) —
  re-applying only the ModelLockoutCard wiring, modelLockout schema and i18n keys.
- Implemented the spec the new tests demanded: backoff cap (BACKOFF_CONFIG.max /
  options.maxCooldownMs), canonical provider aliases in lock keys (cx → codex),
  escalation window extended past the applied cooldown, and in-memory lockModel
  on the persistUnavailableState=false path (combo transient 429).
- Removed the 7 unregistered CLI commands (registry.mjs is generated from the
  OpenAPI spec; shipping them unwired would be dead code — they belong in their
  own PR via the generator).

Validation: 259/259 unit tests green (lockout + auth + full combo suites),
typecheck:core clean, eslint 0 errors.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
oyi77 pushed a commit to oyi77/OmniRoute that referenced this pull request Jun 12, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
oyi77 pushed a commit to oyi77/OmniRoute that referenced this pull request Jun 12, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
oyi77 pushed a commit to oyi77/OmniRoute that referenced this pull request Jun 12, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
oyi77 pushed a commit to oyi77/OmniRoute that referenced this pull request Jun 12, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
oyi77 pushed a commit to oyi77/OmniRoute that referenced this pull request Jun 12, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
oyi77 pushed a commit to oyi77/OmniRoute that referenced this pull request Jun 12, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
oyi77 pushed a commit to oyi77/OmniRoute that referenced this pull request Jun 12, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
diegosouzapw added a commit that referenced this pull request Jun 13, 2026
* chore(release): open v3.8.23 development cycle

* fix(anthropic): strip top_p when temperature is set to avoid 400 (#3691)

Integrated into release/v3.8.23

* fix(vertex): support Vertex AI Express-mode API keys (#3690)

Integrated into release/v3.8.23

* fix(stream): error on empty Claude SSE instead of synthetic success (#3689)

Integrated into release/v3.8.23

* fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (#3692)

Integrated into release/v3.8.23

* docs: add FUNDING.yml and Support section to README (#3698)

Integrated into release/v3.8.23

* feat: gemini - handle known ratelimits (#3686)

Integrated into release/v3.8.23

* fix: stream combo fails over on empty content-filtered response (#3685) (#3702)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (#3696) (#3703)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(auto-combo): add auto-updating model intelligence scoring (#3660)

Integrated into release/v3.8.23

* fix(gemini): context-mode fallback for signatureless tool calls (#3688) (#3704)

* chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (#3705)

* feat(vertex): dynamic model discovery via Generative Language models API (#3712)

Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean.

* fix(combo): gate reasoning token buffer (#3700)

Integrated into release/v3.8.23. Makes the #3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean.

* refactor(#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (#3717)

Phase 1g-1j of #3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix.

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* refactor(#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (#3721)

Phase 1k-1m of #3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* docs(changelog): restore #3590 bullet lost on the v3.8.20 release branch

The fix itself reached main pre-tag via cherry-pick #3591, but its changelog
bullet (commit e33fdd4) only ever existed on release/v3.8.20 after the
squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch
leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md).

* fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (#3722)

Integrated into release/v3.8.23

* refactor(#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (#3725)

Phase 1n-1s of #3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (#3629)

Integrated into release/v3.8.23

* refactor(#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (#3727)

Phase 1t of #3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (#3726)

Integrated into release/v3.8.23

* feat(vertex): self-tracked USD spend since account added (#3724)

Integrated into release/v3.8.23

* fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (#3288) (#3723)

Integrated into release/v3.8.23

* fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import

#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside
safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed
because typecheck:core does not cover src/sse and no test in the merge gates
loaded chatHelpers via tsx; any consumer that did (chat-context-relay and
chat-route-coverage suites, integration harnesses) failed at module load with
'await can only be used inside an async function'.

safeLogEvents is fire-and-forget logging with an outer try/catch, so making it
async (and 'void'-ing the single chat.ts call site) preserves behavior exactly.

Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts
went from failing-at-load to green (+14 tests destravados).

* fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (#3699)

Integrated into release/v3.8.23

* fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (#3728)

Integrated into release/v3.8.23

* fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (#3729)

Integrated into release/v3.8.23

* chore(deps): bump actions/upload-artifact from 4 to 7 (#3735)

Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml).

* chore(deps): bump actions/cache from 4 to 5 (#3734)

Integrated into release/v3.8.23 — actions/cache v4→v5.

* chore(deps): bump actions/download-artifact from 4 to 8 (#3733)

Integrated into release/v3.8.23 — download-artifact v4→v8.

* feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (#3741)

Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes #3739, related #2879.

Integrated into release/v3.8.23.

* i18n: comprehensive zh-CN translation improvements (#3736)

Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green.

Integrated into release/v3.8.23.

* chore(release): v3.8.23 — 2026-06-12

- CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits)
- fix(webdav): resolve promise on writeStream finish, not req end — eliminates
  intermittent 500 on PUT update (writeStream may not have flushed at rename time)
- test(autoCombo): stub DB calls from PR #3660 in tieredRotation.test.ts to prevent
  5s timeout in vitest (getModelIntelligenceBySource DB init path)
- chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE
  allowlist (introduced by PR #3726 setup-open-code.mjs, not OmniRoute config vars)
- chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated)

* fix(model-family): fallback lookup also tries bare model name with dots

getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" →
"gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The
lookup always missed, returning null for any model whose dots are part of the
name rather than a version separator.

Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both
naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22).

* feat: expose API key cost drilldown + quota % used (#3742)

Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule #18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release.

Integrated into release/v3.8.23.

* feat: add provider display modes — All / Configured / Compact (#3743)

Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23.

Integrated into release/v3.8.23.

* fix(cache): scope semantic-cache signature to API key (#3740)

Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests.

Integrated into release/v3.8.23.

* fix(responses): apply OpenAI Responses API stream=false spec default (#3708)

resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected.

Integrated into release/v3.8.23.

* chore(release): reconcile CI gates for v3.8.23

- file-size baseline: re-freeze 8 files grown by PRs #3742/#3743/#3740
  (cost drilldown, provider display modes, cache key isolation)
- ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift)
- .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (#3741, env-doc-sync)
- CHANGELOG: add formatted bullets for #3742, #3743, #3708, #3740,
  model-family-fallback fix; remove duplicate raw ### Fixed section

* test: restore assert count to satisfy check:test-masking gate

Three test files had net assertion removals after behavior-changing PRs:
- chatcore-translation-paths: emergency fallback moved to routing layer
  (#3699) — add body error assertion + model-name guard
- executor-vertex-extended: non-JSON is now Express API key (#3690) —
  add projects/-path guard to the express-key URL test
- stream-utils: empty streams now emit error (#3685) — add code/message/
  status/completePayload guards to both passthrough and translate variants

All new assertions are meaningful (code enum value, 5xx range, non-empty
message, onComplete must-not-fire contract).

* fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it

---------

Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
Co-authored-by: NOXX - Commiter <artur1992123@mail.ru>
Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com>
Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com>
Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com>
Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
diegosouzapw added a commit to Witroch4/OmniRoute that referenced this pull request Jun 13, 2026
…-size

Sync with release/v3.8.24 (branch was 29 commits behind v3.8.23). Two
review fixes pushed to the PR branch:

- Restore reasoningTokenBufferEnabled in comboRuntimeConfigSchema — the
  branch had dropped it (present on the v3.8.23 fork point and on the
  current release; diegosouzapw#3588/diegosouzapw#3700 reasoning-buffer combo config). Out of
  scope for the strict-mode feature; restored to avoid a regression.
- Re-baseline file-size for the feature growth (ApiManagerPageClient.tsx,
  apiKeys.ts, schemas.ts) plus the diegosouzapw#3780 base.ts carry-over.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
* chore(release): open v3.8.21 development cycle

* fix: pass through valid max_tokens-truncated responses instead of fake 502 (diegosouzapw#3572) (diegosouzapw#3595)

* fix: /v1/completions returns legacy text-completion format, not chat (diegosouzapw#3571) (diegosouzapw#3596)

* fix: z.ai/GLM coding plan no longer shows Monthly 0% when no monthly cap (diegosouzapw#3580) (diegosouzapw#3597)

* docs: mark DISCOVERY_TOOL_DESIGN endpoints as Phase-2 not-yet-implemented (diegosouzapw#3498) (diegosouzapw#3599)

* fix(agent-bridge): add validate-only upstream-ca/test route (diegosouzapw#3488) (diegosouzapw#3600)

* fix(gamification): add level/badges/badges-earned profile routes (diegosouzapw#3484)

* security(oauth): migrate 5 public client_ids to resolvePublicCred (diegosouzapw#3493)

* fix(mcp): ship MCP server source closure in npm files + coverage gate (diegosouzapw#3578)

* fix: add reasoning token buffer for combo routing (fixes diegosouzapw#3587) (diegosouzapw#3588)

Integrated into release/v3.8.21

* Refactor: Extract chatCore phases into modular files (diegosouzapw#3598)

Integrated into release/v3.8.21 — chatCore phase modularization. Adjusted: re-derive idempotencyKey for the save path after the check moved into the module (co-authored). Thanks @oyi77!

* docs(changelog): credit diegosouzapw#3598 (chatCore modularization) + diegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(api): implement GET /api/guardrails + POST /api/guardrails/test, drop shadow/guardrails doc-fiction (diegosouzapw#3496) (diegosouzapw#3602)

Integrated into release/v3.8.21 — implements GET /api/guardrails + POST /api/guardrails/test, removes shadow/guardrails doc-fiction. TDD-validated (5/5) + check-docs-symbols/typecheck/eslint green.

* fix(gemini): isolate textual reasoning wrappers (diegosouzapw#3605)

Split-out PR C from diegosouzapw#3584. Isolates textual reasoning wrappers (<think>/<thinking>/<thought>/<internal_thought>, including malformed/open tags) into reasoning_content across both the non-streaming sanitizer and the Gemini streaming translator, with split-chunk buffering. Additive to the existing textual tool-call pipeline; does not touch the diegosouzapw#3569 native functionResponse path. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(antigravity): normalize Gemini 3.5 Flash tier IDs (diegosouzapw#3603)

Split-out PR A from diegosouzapw#3584. Normalizes the Antigravity/agy Gemini 3.5 Flash tier IDs to clean public names (gemini-3.5-flash-low/medium/high), maps them to the live upstream IDs at the executor boundary, and removes Antigravity from the global model resolver so the executor owns wire normalization. Maintainer follow-up: kept gemini-3.5-flash-preview as a hidden backward-compat alias routing to the High tier (so saved combos/configs keep working). Live-validated the tier set via the agy CLI catalog. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(agent-bridge): surface real MITM startup-failure cause, not always port 443 (diegosouzapw#3606) (diegosouzapw#3608)

Integrated into release/v3.8.21 (diegosouzapw#3606)

* fix(oauth): surface real Kiro import-token failure cause, not a bare 500 (diegosouzapw#3589) (diegosouzapw#3609)

Integrated into release/v3.8.21 (diegosouzapw#3589)

* docs(opencode-provider): soft-deprecate in favor of @omniroute/opencode-plugin (diegosouzapw#3419) (diegosouzapw#3613)

Integrated into release/v3.8.21 (diegosouzapw#3419)

* fix(usage): normalize Antigravity and agy provider quotas (diegosouzapw#3604)

Split-out PR B from diegosouzapw#3584. Normalizes Antigravity/agy provider quotas: prefers retrieveUserQuota for live consumption, falls back to fetchAvailableModels and local usage_history, sanitizes cached Provider Limits so retired upstream IDs are not re-exposed, and schedules a deduplicated post-usage refresh. Maintainer follow-up: decoupled the post-usage refresh via a lightweight usageEvents bus (usageHistory no longer dynamic-imports providerLimits) so it does not pull the executors/translator graph into the typecheck-core surface — typecheck:core stays at 0. Integrated into release/v3.8.21. Thanks @dhaern!

* feat(cli): add autostart on/off/toggle shorthand for headless serve mode (diegosouzapw#3331) (diegosouzapw#3614)

Integrated into release/v3.8.21 (diegosouzapw#3331)

* docs(changelog): credit diegosouzapw#3603 (Flash tier IDs) + diegosouzapw#3604 (provider quotas) + diegosouzapw#3605 (reasoning wrappers)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(review): resolve findings from /review-reviews battery (v3.8.21 hardening) (diegosouzapw#3618)

Pre-release hardening from the /review-reviews battery — 15 findings resolved (L1-L13,L15) + L14 live-verified WONTFIX, convergence re-review clean. lint/typecheck:core/test:vitest(146)/build green; zero new test:unit failures vs baseline 6d24708.

* chore(release): v3.8.21 CHANGELOG + i18n + env-doc sync

---------

Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Raxxoor <manker_lol@hotmail.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
* chore(release): open v3.8.23 development cycle

* fix(anthropic): strip top_p when temperature is set to avoid 400 (diegosouzapw#3691)

Integrated into release/v3.8.23

* fix(vertex): support Vertex AI Express-mode API keys (diegosouzapw#3690)

Integrated into release/v3.8.23

* fix(stream): error on empty Claude SSE instead of synthetic success (diegosouzapw#3689)

Integrated into release/v3.8.23

* fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (diegosouzapw#3692)

Integrated into release/v3.8.23

* docs: add FUNDING.yml and Support section to README (diegosouzapw#3698)

Integrated into release/v3.8.23

* feat: gemini - handle known ratelimits (diegosouzapw#3686)

Integrated into release/v3.8.23

* fix: stream combo fails over on empty content-filtered response (diegosouzapw#3685) (diegosouzapw#3702)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (diegosouzapw#3696) (diegosouzapw#3703)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(auto-combo): add auto-updating model intelligence scoring (diegosouzapw#3660)

Integrated into release/v3.8.23

* fix(gemini): context-mode fallback for signatureless tool calls (diegosouzapw#3688) (diegosouzapw#3704)

* chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (diegosouzapw#3705)

* feat(vertex): dynamic model discovery via Generative Language models API (diegosouzapw#3712)

Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean.

* fix(combo): gate reasoning token buffer (diegosouzapw#3700)

Integrated into release/v3.8.23. Makes the diegosouzapw#3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean.

* refactor(diegosouzapw#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (diegosouzapw#3717)

Phase 1g-1j of diegosouzapw#3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix.

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* refactor(diegosouzapw#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (diegosouzapw#3721)

Phase 1k-1m of diegosouzapw#3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* docs(changelog): restore diegosouzapw#3590 bullet lost on the v3.8.20 release branch

The fix itself reached main pre-tag via cherry-pick diegosouzapw#3591, but its changelog
bullet (commit db04ef2) only ever existed on release/v3.8.20 after the
squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch
leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md).

* fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (diegosouzapw#3722)

Integrated into release/v3.8.23

* refactor(diegosouzapw#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (diegosouzapw#3725)

Phase 1n-1s of diegosouzapw#3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (diegosouzapw#3629)

Integrated into release/v3.8.23

* refactor(diegosouzapw#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (diegosouzapw#3727)

Phase 1t of diegosouzapw#3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (diegosouzapw#3726)

Integrated into release/v3.8.23

* feat(vertex): self-tracked USD spend since account added (diegosouzapw#3724)

Integrated into release/v3.8.23

* fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (diegosouzapw#3288) (diegosouzapw#3723)

Integrated into release/v3.8.23

* fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import

diegosouzapw#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside
safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed
because typecheck:core does not cover src/sse and no test in the merge gates
loaded chatHelpers via tsx; any consumer that did (chat-context-relay and
chat-route-coverage suites, integration harnesses) failed at module load with
'await can only be used inside an async function'.

safeLogEvents is fire-and-forget logging with an outer try/catch, so making it
async (and 'void'-ing the single chat.ts call site) preserves behavior exactly.

Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts
went from failing-at-load to green (+14 tests destravados).

* fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (diegosouzapw#3699)

Integrated into release/v3.8.23

* fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (diegosouzapw#3728)

Integrated into release/v3.8.23

* fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (diegosouzapw#3729)

Integrated into release/v3.8.23

* chore(deps): bump actions/upload-artifact from 4 to 7 (diegosouzapw#3735)

Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml).

* chore(deps): bump actions/cache from 4 to 5 (diegosouzapw#3734)

Integrated into release/v3.8.23 — actions/cache v4→v5.

* chore(deps): bump actions/download-artifact from 4 to 8 (diegosouzapw#3733)

Integrated into release/v3.8.23 — download-artifact v4→v8.

* feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (diegosouzapw#3741)

Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes diegosouzapw#3739, related diegosouzapw#2879.

Integrated into release/v3.8.23.

* i18n: comprehensive zh-CN translation improvements (diegosouzapw#3736)

Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green.

Integrated into release/v3.8.23.

* chore(release): v3.8.23 — 2026-06-12

- CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits)
- fix(webdav): resolve promise on writeStream finish, not req end — eliminates
  intermittent 500 on PUT update (writeStream may not have flushed at rename time)
- test(autoCombo): stub DB calls from PR diegosouzapw#3660 in tieredRotation.test.ts to prevent
  5s timeout in vitest (getModelIntelligenceBySource DB init path)
- chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE
  allowlist (introduced by PR diegosouzapw#3726 setup-open-code.mjs, not OmniRoute config vars)
- chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated)

* fix(model-family): fallback lookup also tries bare model name with dots

getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" →
"gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The
lookup always missed, returning null for any model whose dots are part of the
name rather than a version separator.

Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both
naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22).

* feat: expose API key cost drilldown + quota % used (diegosouzapw#3742)

Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule diegosouzapw#18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release.

Integrated into release/v3.8.23.

* feat: add provider display modes — All / Configured / Compact (diegosouzapw#3743)

Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23.

Integrated into release/v3.8.23.

* fix(cache): scope semantic-cache signature to API key (diegosouzapw#3740)

Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests.

Integrated into release/v3.8.23.

* fix(responses): apply OpenAI Responses API stream=false spec default (diegosouzapw#3708)

resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected.

Integrated into release/v3.8.23.

* chore(release): reconcile CI gates for v3.8.23

- file-size baseline: re-freeze 8 files grown by PRs diegosouzapw#3742/diegosouzapw#3743/diegosouzapw#3740
  (cost drilldown, provider display modes, cache key isolation)
- ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift)
- .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (diegosouzapw#3741, env-doc-sync)
- CHANGELOG: add formatted bullets for diegosouzapw#3742, diegosouzapw#3743, diegosouzapw#3708, diegosouzapw#3740,
  model-family-fallback fix; remove duplicate raw ### Fixed section

* test: restore assert count to satisfy check:test-masking gate

Three test files had net assertion removals after behavior-changing PRs:
- chatcore-translation-paths: emergency fallback moved to routing layer
  (diegosouzapw#3699) — add body error assertion + model-name guard
- executor-vertex-extended: non-JSON is now Express API key (diegosouzapw#3690) —
  add projects/-path guard to the express-key URL test
- stream-utils: empty streams now emit error (diegosouzapw#3685) — add code/message/
  status/completePayload guards to both passthrough and translate variants

All new assertions are meaningful (code enum value, 5xx range, non-empty
message, onComplete must-not-fire contract).

* fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it

---------

Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
Co-authored-by: NOXX - Commiter <artur1992123@mail.ru>
Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com>
Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com>
Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com>
Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
* chore(release): open v3.8.21 development cycle

* fix: pass through valid max_tokens-truncated responses instead of fake 502 (diegosouzapw#3572) (diegosouzapw#3595)

* fix: /v1/completions returns legacy text-completion format, not chat (diegosouzapw#3571) (diegosouzapw#3596)

* fix: z.ai/GLM coding plan no longer shows Monthly 0% when no monthly cap (diegosouzapw#3580) (diegosouzapw#3597)

* docs: mark DISCOVERY_TOOL_DESIGN endpoints as Phase-2 not-yet-implemented (diegosouzapw#3498) (diegosouzapw#3599)

* fix(agent-bridge): add validate-only upstream-ca/test route (diegosouzapw#3488) (diegosouzapw#3600)

* fix(gamification): add level/badges/badges-earned profile routes (diegosouzapw#3484)

* security(oauth): migrate 5 public client_ids to resolvePublicCred (diegosouzapw#3493)

* fix(mcp): ship MCP server source closure in npm files + coverage gate (diegosouzapw#3578)

* fix: add reasoning token buffer for combo routing (fixes diegosouzapw#3587) (diegosouzapw#3588)

Integrated into release/v3.8.21

* Refactor: Extract chatCore phases into modular files (diegosouzapw#3598)

Integrated into release/v3.8.21 — chatCore phase modularization. Adjusted: re-derive idempotencyKey for the save path after the check moved into the module (co-authored). Thanks @oyi77!

* docs(changelog): credit diegosouzapw#3598 (chatCore modularization) + diegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(api): implement GET /api/guardrails + POST /api/guardrails/test, drop shadow/guardrails doc-fiction (diegosouzapw#3496) (diegosouzapw#3602)

Integrated into release/v3.8.21 — implements GET /api/guardrails + POST /api/guardrails/test, removes shadow/guardrails doc-fiction. TDD-validated (5/5) + check-docs-symbols/typecheck/eslint green.

* fix(gemini): isolate textual reasoning wrappers (diegosouzapw#3605)

Split-out PR C from diegosouzapw#3584. Isolates textual reasoning wrappers (<think>/<thinking>/<thought>/<internal_thought>, including malformed/open tags) into reasoning_content across both the non-streaming sanitizer and the Gemini streaming translator, with split-chunk buffering. Additive to the existing textual tool-call pipeline; does not touch the diegosouzapw#3569 native functionResponse path. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(antigravity): normalize Gemini 3.5 Flash tier IDs (diegosouzapw#3603)

Split-out PR A from diegosouzapw#3584. Normalizes the Antigravity/agy Gemini 3.5 Flash tier IDs to clean public names (gemini-3.5-flash-low/medium/high), maps them to the live upstream IDs at the executor boundary, and removes Antigravity from the global model resolver so the executor owns wire normalization. Maintainer follow-up: kept gemini-3.5-flash-preview as a hidden backward-compat alias routing to the High tier (so saved combos/configs keep working). Live-validated the tier set via the agy CLI catalog. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(agent-bridge): surface real MITM startup-failure cause, not always port 443 (diegosouzapw#3606) (diegosouzapw#3608)

Integrated into release/v3.8.21 (diegosouzapw#3606)

* fix(oauth): surface real Kiro import-token failure cause, not a bare 500 (diegosouzapw#3589) (diegosouzapw#3609)

Integrated into release/v3.8.21 (diegosouzapw#3589)

* docs(opencode-provider): soft-deprecate in favor of @omniroute/opencode-plugin (diegosouzapw#3419) (diegosouzapw#3613)

Integrated into release/v3.8.21 (diegosouzapw#3419)

* fix(usage): normalize Antigravity and agy provider quotas (diegosouzapw#3604)

Split-out PR B from diegosouzapw#3584. Normalizes Antigravity/agy provider quotas: prefers retrieveUserQuota for live consumption, falls back to fetchAvailableModels and local usage_history, sanitizes cached Provider Limits so retired upstream IDs are not re-exposed, and schedules a deduplicated post-usage refresh. Maintainer follow-up: decoupled the post-usage refresh via a lightweight usageEvents bus (usageHistory no longer dynamic-imports providerLimits) so it does not pull the executors/translator graph into the typecheck-core surface — typecheck:core stays at 0. Integrated into release/v3.8.21. Thanks @dhaern!

* feat(cli): add autostart on/off/toggle shorthand for headless serve mode (diegosouzapw#3331) (diegosouzapw#3614)

Integrated into release/v3.8.21 (diegosouzapw#3331)

* docs(changelog): credit diegosouzapw#3603 (Flash tier IDs) + diegosouzapw#3604 (provider quotas) + diegosouzapw#3605 (reasoning wrappers)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(review): resolve findings from /review-reviews battery (v3.8.21 hardening) (diegosouzapw#3618)

Pre-release hardening from the /review-reviews battery — 15 findings resolved (L1-L13,L15) + L14 live-verified WONTFIX, convergence re-review clean. lint/typecheck:core/test:vitest(146)/build green; zero new test:unit failures vs baseline 797de43.

* chore(release): v3.8.21 CHANGELOG + i18n + env-doc sync

---------

Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Raxxoor <manker_lol@hotmail.com>
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
* chore(release): open v3.8.23 development cycle

* fix(anthropic): strip top_p when temperature is set to avoid 400 (diegosouzapw#3691)

Integrated into release/v3.8.23

* fix(vertex): support Vertex AI Express-mode API keys (diegosouzapw#3690)

Integrated into release/v3.8.23

* fix(stream): error on empty Claude SSE instead of synthetic success (diegosouzapw#3689)

Integrated into release/v3.8.23

* fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (diegosouzapw#3692)

Integrated into release/v3.8.23

* docs: add FUNDING.yml and Support section to README (diegosouzapw#3698)

Integrated into release/v3.8.23

* feat: gemini - handle known ratelimits (diegosouzapw#3686)

Integrated into release/v3.8.23

* fix: stream combo fails over on empty content-filtered response (diegosouzapw#3685) (diegosouzapw#3702)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (diegosouzapw#3696) (diegosouzapw#3703)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(auto-combo): add auto-updating model intelligence scoring (diegosouzapw#3660)

Integrated into release/v3.8.23

* fix(gemini): context-mode fallback for signatureless tool calls (diegosouzapw#3688) (diegosouzapw#3704)

* chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (diegosouzapw#3705)

* feat(vertex): dynamic model discovery via Generative Language models API (diegosouzapw#3712)

Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean.

* fix(combo): gate reasoning token buffer (diegosouzapw#3700)

Integrated into release/v3.8.23. Makes the diegosouzapw#3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean.

* refactor(diegosouzapw#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (diegosouzapw#3717)

Phase 1g-1j of diegosouzapw#3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix.

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* refactor(diegosouzapw#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (diegosouzapw#3721)

Phase 1k-1m of diegosouzapw#3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* docs(changelog): restore diegosouzapw#3590 bullet lost on the v3.8.20 release branch

The fix itself reached main pre-tag via cherry-pick diegosouzapw#3591, but its changelog
bullet (commit e33fdd4) only ever existed on release/v3.8.20 after the
squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch
leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md).

* fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (diegosouzapw#3722)

Integrated into release/v3.8.23

* refactor(diegosouzapw#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (diegosouzapw#3725)

Phase 1n-1s of diegosouzapw#3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (diegosouzapw#3629)

Integrated into release/v3.8.23

* refactor(diegosouzapw#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (diegosouzapw#3727)

Phase 1t of diegosouzapw#3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (diegosouzapw#3726)

Integrated into release/v3.8.23

* feat(vertex): self-tracked USD spend since account added (diegosouzapw#3724)

Integrated into release/v3.8.23

* fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (diegosouzapw#3288) (diegosouzapw#3723)

Integrated into release/v3.8.23

* fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import

diegosouzapw#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside
safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed
because typecheck:core does not cover src/sse and no test in the merge gates
loaded chatHelpers via tsx; any consumer that did (chat-context-relay and
chat-route-coverage suites, integration harnesses) failed at module load with
'await can only be used inside an async function'.

safeLogEvents is fire-and-forget logging with an outer try/catch, so making it
async (and 'void'-ing the single chat.ts call site) preserves behavior exactly.

Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts
went from failing-at-load to green (+14 tests destravados).

* fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (diegosouzapw#3699)

Integrated into release/v3.8.23

* fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (diegosouzapw#3728)

Integrated into release/v3.8.23

* fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (diegosouzapw#3729)

Integrated into release/v3.8.23

* chore(deps): bump actions/upload-artifact from 4 to 7 (diegosouzapw#3735)

Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml).

* chore(deps): bump actions/cache from 4 to 5 (diegosouzapw#3734)

Integrated into release/v3.8.23 — actions/cache v4→v5.

* chore(deps): bump actions/download-artifact from 4 to 8 (diegosouzapw#3733)

Integrated into release/v3.8.23 — download-artifact v4→v8.

* feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (diegosouzapw#3741)

Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes diegosouzapw#3739, related diegosouzapw#2879.

Integrated into release/v3.8.23.

* i18n: comprehensive zh-CN translation improvements (diegosouzapw#3736)

Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green.

Integrated into release/v3.8.23.

* chore(release): v3.8.23 — 2026-06-12

- CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits)
- fix(webdav): resolve promise on writeStream finish, not req end — eliminates
  intermittent 500 on PUT update (writeStream may not have flushed at rename time)
- test(autoCombo): stub DB calls from PR diegosouzapw#3660 in tieredRotation.test.ts to prevent
  5s timeout in vitest (getModelIntelligenceBySource DB init path)
- chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE
  allowlist (introduced by PR diegosouzapw#3726 setup-open-code.mjs, not OmniRoute config vars)
- chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated)

* fix(model-family): fallback lookup also tries bare model name with dots

getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" →
"gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The
lookup always missed, returning null for any model whose dots are part of the
name rather than a version separator.

Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both
naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22).

* feat: expose API key cost drilldown + quota % used (diegosouzapw#3742)

Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule diegosouzapw#18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release.

Integrated into release/v3.8.23.

* feat: add provider display modes — All / Configured / Compact (diegosouzapw#3743)

Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23.

Integrated into release/v3.8.23.

* fix(cache): scope semantic-cache signature to API key (diegosouzapw#3740)

Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests.

Integrated into release/v3.8.23.

* fix(responses): apply OpenAI Responses API stream=false spec default (diegosouzapw#3708)

resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected.

Integrated into release/v3.8.23.

* chore(release): reconcile CI gates for v3.8.23

- file-size baseline: re-freeze 8 files grown by PRs diegosouzapw#3742/diegosouzapw#3743/diegosouzapw#3740
  (cost drilldown, provider display modes, cache key isolation)
- ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift)
- .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (diegosouzapw#3741, env-doc-sync)
- CHANGELOG: add formatted bullets for diegosouzapw#3742, diegosouzapw#3743, diegosouzapw#3708, diegosouzapw#3740,
  model-family-fallback fix; remove duplicate raw ### Fixed section

* test: restore assert count to satisfy check:test-masking gate

Three test files had net assertion removals after behavior-changing PRs:
- chatcore-translation-paths: emergency fallback moved to routing layer
  (diegosouzapw#3699) — add body error assertion + model-name guard
- executor-vertex-extended: non-JSON is now Express API key (diegosouzapw#3690) —
  add projects/-path guard to the express-key URL test
- stream-utils: empty streams now emit error (diegosouzapw#3685) — add code/message/
  status/completePayload guards to both passthrough and translate variants

All new assertions are meaningful (code enum value, 5xx range, non-empty
message, onComplete must-not-fire contract).

* fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it

---------

Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
Co-authored-by: NOXX - Commiter <artur1992123@mail.ru>
Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com>
Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com>
Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com>
Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
Integrated into release/v3.8.23. Makes the diegosouzapw#3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean.
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…iegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
* chore(release): open v3.8.21 development cycle

* fix: pass through valid max_tokens-truncated responses instead of fake 502 (diegosouzapw#3572) (diegosouzapw#3595)

* fix: /v1/completions returns legacy text-completion format, not chat (diegosouzapw#3571) (diegosouzapw#3596)

* fix: z.ai/GLM coding plan no longer shows Monthly 0% when no monthly cap (diegosouzapw#3580) (diegosouzapw#3597)

* docs: mark DISCOVERY_TOOL_DESIGN endpoints as Phase-2 not-yet-implemented (diegosouzapw#3498) (diegosouzapw#3599)

* fix(agent-bridge): add validate-only upstream-ca/test route (diegosouzapw#3488) (diegosouzapw#3600)

* fix(gamification): add level/badges/badges-earned profile routes (diegosouzapw#3484)

* security(oauth): migrate 5 public client_ids to resolvePublicCred (diegosouzapw#3493)

* fix(mcp): ship MCP server source closure in npm files + coverage gate (diegosouzapw#3578)

* fix: add reasoning token buffer for combo routing (fixes diegosouzapw#3587) (diegosouzapw#3588)

Integrated into release/v3.8.21

* Refactor: Extract chatCore phases into modular files (diegosouzapw#3598)

Integrated into release/v3.8.21 — chatCore phase modularization. Adjusted: re-derive idempotencyKey for the save path after the check moved into the module (co-authored). Thanks @oyi77!

* docs(changelog): credit diegosouzapw#3598 (chatCore modularization) + diegosouzapw#3588 (combo reasoning buffer)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(api): implement GET /api/guardrails + POST /api/guardrails/test, drop shadow/guardrails doc-fiction (diegosouzapw#3496) (diegosouzapw#3602)

Integrated into release/v3.8.21 — implements GET /api/guardrails + POST /api/guardrails/test, removes shadow/guardrails doc-fiction. TDD-validated (5/5) + check-docs-symbols/typecheck/eslint green.

* fix(gemini): isolate textual reasoning wrappers (diegosouzapw#3605)

Split-out PR C from diegosouzapw#3584. Isolates textual reasoning wrappers (<think>/<thinking>/<thought>/<internal_thought>, including malformed/open tags) into reasoning_content across both the non-streaming sanitizer and the Gemini streaming translator, with split-chunk buffering. Additive to the existing textual tool-call pipeline; does not touch the diegosouzapw#3569 native functionResponse path. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(antigravity): normalize Gemini 3.5 Flash tier IDs (diegosouzapw#3603)

Split-out PR A from diegosouzapw#3584. Normalizes the Antigravity/agy Gemini 3.5 Flash tier IDs to clean public names (gemini-3.5-flash-low/medium/high), maps them to the live upstream IDs at the executor boundary, and removes Antigravity from the global model resolver so the executor owns wire normalization. Maintainer follow-up: kept gemini-3.5-flash-preview as a hidden backward-compat alias routing to the High tier (so saved combos/configs keep working). Live-validated the tier set via the agy CLI catalog. Integrated into release/v3.8.21. Thanks @dhaern!

* fix(agent-bridge): surface real MITM startup-failure cause, not always port 443 (diegosouzapw#3606) (diegosouzapw#3608)

Integrated into release/v3.8.21 (diegosouzapw#3606)

* fix(oauth): surface real Kiro import-token failure cause, not a bare 500 (diegosouzapw#3589) (diegosouzapw#3609)

Integrated into release/v3.8.21 (diegosouzapw#3589)

* docs(opencode-provider): soft-deprecate in favor of @omniroute/opencode-plugin (diegosouzapw#3419) (diegosouzapw#3613)

Integrated into release/v3.8.21 (diegosouzapw#3419)

* fix(usage): normalize Antigravity and agy provider quotas (diegosouzapw#3604)

Split-out PR B from diegosouzapw#3584. Normalizes Antigravity/agy provider quotas: prefers retrieveUserQuota for live consumption, falls back to fetchAvailableModels and local usage_history, sanitizes cached Provider Limits so retired upstream IDs are not re-exposed, and schedules a deduplicated post-usage refresh. Maintainer follow-up: decoupled the post-usage refresh via a lightweight usageEvents bus (usageHistory no longer dynamic-imports providerLimits) so it does not pull the executors/translator graph into the typecheck-core surface — typecheck:core stays at 0. Integrated into release/v3.8.21. Thanks @dhaern!

* feat(cli): add autostart on/off/toggle shorthand for headless serve mode (diegosouzapw#3331) (diegosouzapw#3614)

Integrated into release/v3.8.21 (diegosouzapw#3331)

* docs(changelog): credit diegosouzapw#3603 (Flash tier IDs) + diegosouzapw#3604 (provider quotas) + diegosouzapw#3605 (reasoning wrappers)

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>

* fix(review): resolve findings from /review-reviews battery (v3.8.21 hardening) (diegosouzapw#3618)

Pre-release hardening from the /review-reviews battery — 15 findings resolved (L1-L13,L15) + L14 live-verified WONTFIX, convergence re-review clean. lint/typecheck:core/test:vitest(146)/build green; zero new test:unit failures vs baseline 408d91a2c.

* chore(release): v3.8.21 CHANGELOG + i18n + env-doc sync

---------

Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Paijo <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Raxxoor <manker_lol@hotmail.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
* chore(release): open v3.8.23 development cycle

* fix(anthropic): strip top_p when temperature is set to avoid 400 (diegosouzapw#3691)

Integrated into release/v3.8.23

* fix(vertex): support Vertex AI Express-mode API keys (diegosouzapw#3690)

Integrated into release/v3.8.23

* fix(stream): error on empty Claude SSE instead of synthetic success (diegosouzapw#3689)

Integrated into release/v3.8.23

* fix(oauth): stop token-refresh invalidation loop + harden proxy resolution (diegosouzapw#3692)

Integrated into release/v3.8.23

* docs: add FUNDING.yml and Support section to README (diegosouzapw#3698)

Integrated into release/v3.8.23

* feat: gemini - handle known ratelimits (diegosouzapw#3686)

Integrated into release/v3.8.23

* fix: stream combo fails over on empty content-filtered response (diegosouzapw#3685) (diegosouzapw#3702)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(antigravity): preserve gemini-3.1-pro high/low budget tiers (diegosouzapw#3696) (diegosouzapw#3703)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(auto-combo): add auto-updating model intelligence scoring (diegosouzapw#3660)

Integrated into release/v3.8.23

* fix(gemini): context-mode fallback for signatureless tool calls (diegosouzapw#3688) (diegosouzapw#3704)

* chore(quality-gate): reconcile file-size baseline (27 files + providerLimits.ts) (diegosouzapw#3705)

* feat(vertex): dynamic model discovery via Generative Language models API (diegosouzapw#3712)

Integrated into release/v3.8.23. Vertex dynamic model discovery — surfaces image models (imagen-*, gemini-*-image), embeddings and audio from the live Generative Language catalog, with cached→static fallback and the shared parseGeminiModelsList helper. Validated: parser test 5/5, typecheck:core clean.

* fix(combo): gate reasoning token buffer (diegosouzapw#3700)

Integrated into release/v3.8.23. Makes the diegosouzapw#3588 reasoning token buffer safe and configurable: only inflates max_tokens when the model has a known, non-default output cap and the buffered value fits inside it; otherwise preserves/clamps the client limit. Adds the reasoningTokenBufferEnabled kill switch (default ON). Validated: combo-routing-engine 81/81, combo-config 25/25, combo-quality-validator-reasoning 12/12, phase1f 10/10, typecheck:core clean.

* refactor(diegosouzapw#3501): god-component Phase 1g-1j — client 4062→3408 LOC (-654) (diegosouzapw#3717)

Phase 1g-1j of diegosouzapw#3501: client 4062→3408 LOC. Pure extraction (ProviderPlaygroundPanel, useCommandCodeAuth, useExternalLinkFlow+ExternalLinkModal, useAuthFileHandlers) + loadConnProxies ReferenceError fix + phase1f test path fix.

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* refactor(diegosouzapw#3501): god-component Phase 1k-1m — client 3408→2553 LOC (-855) (diegosouzapw#3721)

Phase 1k-1m of diegosouzapw#3501: client 3408→2553 LOC. Pure extraction (useModelImportHandlers+ImportProgressModal, useModelVisibilityHandlers, ProviderModelsSection).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* docs(changelog): restore diegosouzapw#3590 bullet lost on the v3.8.20 release branch

The fix itself reached main pre-tag via cherry-pick diegosouzapw#3591, but its changelog
bullet (commit a6b99843f) only ever existed on release/v3.8.20 after the
squash-merge. Restored under [3.8.20] per the 2026-06-12 release-branch
leftover audit (_tasks/release-audit/release-leftovers-audit-2026-06-12.md).

* fix(kiro): resolve quota for IAM Identity Center accounts missing a profileArn (diegosouzapw#3722)

Integrated into release/v3.8.23

* refactor(diegosouzapw#3501): god-component Phase 1n-1s — client 2553→1376 LOC (-1177) (diegosouzapw#3725)

Phase 1n-1s of diegosouzapw#3501: client 2553→1376 LOC. Pure extraction (ConnectionsListPanel, ConnectionsHeaderToolbar, ZedImportCard, BatchTestResultsModal, AdaptaTutorialModal, useApiKeySave + helpers).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* feat(model-lockout): settings UI, backend integration, error classification, and success-decay recovery (diegosouzapw#3629)

Integrated into release/v3.8.23

* refactor(diegosouzapw#3501): god-component Phase 1t — client 1376→781 LOC (≤800 TARGET REACHED ✅) (diegosouzapw#3727)

Phase 1t of diegosouzapw#3501: client 1376→781 LOC (≤800 reached). Original god-component 12,882→781 (−94%).

Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>

* fix: bundle @omniroute/opencode-plugin inside omniroute + add 'setup opencode' CLI command (diegosouzapw#3726)

Integrated into release/v3.8.23

* feat(vertex): self-tracked USD spend since account added (diegosouzapw#3724)

Integrated into release/v3.8.23

* fix(qwen-web): migrate to v2 chat API with full cookie-jar replay (diegosouzapw#3288) (diegosouzapw#3723)

Integrated into release/v3.8.23

* fix(sse): make safeLogEvents async — 'await' in a sync function broke every chatHelpers import

diegosouzapw#3692 added a lazy 'await import(proxyEgress)' for egress-IP visibility inside
safeLogEvents, which is a sync function — an ES syntax error. It went unnoticed
because typecheck:core does not cover src/sse and no test in the merge gates
loaded chatHelpers via tsx; any consumer that did (chat-context-relay and
chat-route-coverage suites, integration harnesses) failed at module load with
'await can only be used inside an async function'.

safeLogEvents is fire-and-forget logging with an outer try/catch, so making it
async (and 'void'-ing the single chat.ts call site) preserves behavior exactly.

Validation: tests/unit/chat-context-relay.test.ts + chat-route-coverage.test.ts
went from failing-at-load to green (+14 tests destravados).

* fix(sse): remove cross-provider credential leak in emergency fallback + combo/proxy audit fixes (diegosouzapw#3699)

Integrated into release/v3.8.23

* fix(executors): inject MiMoCode anti-abuse marker so free endpoint stops 403ing (diegosouzapw#3728)

Integrated into release/v3.8.23

* fix(dashboard): repair "Test all models" — toast crash, status icons, auto-hide (diegosouzapw#3729)

Integrated into release/v3.8.23

* chore(deps): bump actions/upload-artifact from 4 to 7 (diegosouzapw#3735)

Integrated into release/v3.8.23 — aligns upload-artifact to v7 (already used across ci.yml).

* chore(deps): bump actions/cache from 4 to 5 (diegosouzapw#3734)

Integrated into release/v3.8.23 — actions/cache v4→v5.

* chore(deps): bump actions/download-artifact from 4 to 8 (diegosouzapw#3733)

Integrated into release/v3.8.23 — download-artifact v4→v8.

* feat(fallback): add OMNIROUTE_EMERGENCY_FALLBACK env switch (diegosouzapw#3741)

Adds an OMNIROUTE_EMERGENCY_FALLBACK env switch to disable the emergency budget-exhaustion fallback (reroute to free nvidia/gpt-oss-120b). Default unchanged (enabled). Closes diegosouzapw#3739, related diegosouzapw#2879.

Integrated into release/v3.8.23.

* i18n: comprehensive zh-CN translation improvements (diegosouzapw#3736)

Aligns zh-CN to en (hundreds of entries), translates batch-action labels + settings sidebar menu, adds categoryConfig/endpointTokenSaver keys, resolves __MISSING__ stubs. Sidebar/SidebarTab hardcoded strings replaced with t(). en.json purely additive (8 new sidebar.* keys, 0 removed); cli-i18n gate green.

Integrated into release/v3.8.23.

* chore(release): v3.8.23 — 2026-06-12

- CHANGELOG: complete v3.8.23 section (28 bullets, 27 commits)
- fix(webdav): resolve promise on writeStream finish, not req end — eliminates
  intermittent 500 on PUT update (writeStream may not have flushed at rename time)
- test(autoCombo): stub DB calls from PR diegosouzapw#3660 in tieredRotation.test.ts to prevent
  5s timeout in vitest (getModelIntelligenceBySource DB init path)
- chore(env-sync): add XDG_DATA_HOME + OMNIROUTE_OPENCODE_PLUGIN_DIR to IGNORE_FROM_CODE
  allowlist (introduced by PR diegosouzapw#3726 setup-open-code.mjs, not OmniRoute config vars)
- chore(cli): regenerated bin/cli/api-commands/*.mjs (7 new, 27 updated)

* fix(model-family): fallback lookup also tries bare model name with dots

getNextFamilyFallback normalized dots-to-hyphens ("gemini-3.1-pro-high" →
"gemini-3-1-pro-high") but MODEL_FAMILIES keys use the literal dot form. The
lookup always missed, returning null for any model whose dots are part of the
name rather than a version separator.

Fallback: try MODEL_FAMILIES[lookupKey] ?? MODEL_FAMILIES[bareModel] so both
naming conventions are covered. Fixes T30 test (pre-existing since v3.8.22).

* feat: expose API key cost drilldown + quota % used (diegosouzapw#3742)

Adds all-time USD cost per API key in the API Key Manager, a per-key deep-link into the Cost Explorer (filtered + grouped by model), URL-param hydration of range/groupBy/apiKeyIds, and a '% used' quota display. Review adjustments: extracted URL-param parsers to a tested module (Rule diegosouzapw#18), i18n'd the new strings (en + zh-CN), dropped the redundant webdav-handler entry already on release.

Integrated into release/v3.8.23.

* feat: add provider display modes — All / Configured / Compact (diegosouzapw#3743)

Replaces the Providers page configured-only toggle with All/Configured/Compact display modes (Compact = flat deduped grid, no-auth last). Persists the preference and migrates the legacy localStorage key. Rebased onto release/v3.8.23.

Integrated into release/v3.8.23.

* fix(cache): scope semantic-cache signature to API key (diegosouzapw#3740)

Adds the api_key_id dimension to generateSignature's SHA-256 hash so two callers with different API keys never receive each other's cached responses. Threads apiKeyId through checkSemanticCache + both write sites; migration 098 clears pre-existing key-less entries; unauthenticated requests stay isolated from keyed ones. 3 TDD tests.

Integrated into release/v3.8.23.

* fix(responses): apply OpenAI Responses API stream=false spec default (diegosouzapw#3708)

resolveStreamFlag now applies the stream=false-when-omitted default for sourceFormat=openai-responses (same as the existing claude path), so spec-compliant /v1/responses upstreams that return JSON no longer fall through to the wildcard-Accept heuristic and trigger STREAM_EARLY_EOF / 502. Codex CLI (stream:true) and explicit text/event-stream clients unaffected.

Integrated into release/v3.8.23.

* chore(release): reconcile CI gates for v3.8.23

- file-size baseline: re-freeze 8 files grown by PRs diegosouzapw#3742/diegosouzapw#3743/diegosouzapw#3740
  (cost drilldown, provider display modes, cache key isolation)
- ARCHITECTURE.md: update executor count 55→60 (check:docs-counts drift)
- .env.example: add OMNIROUTE_EMERGENCY_FALLBACK (diegosouzapw#3741, env-doc-sync)
- CHANGELOG: add formatted bullets for diegosouzapw#3742, diegosouzapw#3743, diegosouzapw#3708, diegosouzapw#3740,
  model-family-fallback fix; remove duplicate raw ### Fixed section

* test: restore assert count to satisfy check:test-masking gate

Three test files had net assertion removals after behavior-changing PRs:
- chatcore-translation-paths: emergency fallback moved to routing layer
  (diegosouzapw#3699) — add body error assertion + model-name guard
- executor-vertex-extended: non-JSON is now Express API key (diegosouzapw#3690) —
  add projects/-path guard to the express-key URL test
- stream-utils: empty streams now emit error (diegosouzapw#3685) — add code/message/
  status/completePayload guards to both passthrough and translate variants

All new assertions are meaningful (code enum value, 5xx range, non-empty
message, onComplete must-not-fire contract).

* fix(ci): move rtl-logical-classes test to ui/ so vitest:ui runner collects it

---------

Co-authored-by: Felipe Almeman <4226997+zhiru@users.noreply.github.com>
Co-authored-by: NOXX - Commiter <artur1992123@mail.ru>
Co-authored-by: Nick Sullivan <142708+TechNickAI@users.noreply.github.com>
Co-authored-by: Markus Hartung <mail@hartmark.se>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: PizzaV <103120356+pizzav-xyz@users.noreply.github.com>
Co-authored-by: Randi <55005611+rdself@users.noreply.github.com>
Co-authored-by: oyi77 <14921983+oyi77@users.noreply.github.com>
Co-authored-by: Chewji <126886556+Chewji9875@users.noreply.github.com>
Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
Co-authored-by: Felipe Sartori <felipesartori.ti@gmail.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Zois Pagoulatos <zpagoulatos@hotmail.com>
Co-authored-by: sdfsdfw2 <167810361+sdfsdfw2@users.noreply.github.com>
Co-authored-by: Witroch4 <witalo_rocha@hotmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] Reasoning models in combos consume all tokens for reasoning_content, leaving content empty

2 participants