Skip to content

Release v3.8.32 - #4418

Merged
diegosouzapw merged 57 commits into
mainfrom
release/v3.8.32
Jun 21, 2026
Merged

diegosouzapw merged 57 commits into
mainfrom
release/v3.8.32

Conversation

@diegosouzapw

@diegosouzapw diegosouzapw commented Jun 20, 2026 •

Copy link
Copy Markdown
Owner

Release v3.8.32 — release/v3.8.32 → main. Finalized release notes — 39 commits since v3.8.31.

[3.8.32] — 2026-06-20

✨ New Features

  • feat(oauth): import accounts from CLIProxyAPI — Settings → CLIProxyAPI now has an "Import accounts" button that reads the OAuth accounts CLIProxyAPI already saved in ~/.cli-proxy-api/ and imports them as OmniRoute connections, so you don't have to log into every account individually. CLIProxyAPI's unified auth-file format is parsed by type discriminator and the supported account types (Gemini, Codex, Claude/Anthropic, Antigravity, Qwen, Kimi) are upserted; unknown types are skipped. The preview never exposes tokens to the client. (thanks @powellnorma)
  • feat(routing): opt-in setting to echo the requested alias/combo name in the response model field — Settings → Routing now has an "Echo requested model name in responses" toggle (default off). When enabled, the response model field (non-streaming and every streamed SSE chunk) reports the alias or combo name the client requested instead of the upstream model name, so strict clients such as Claude Desktop — which reject a response whose model does not match the request with a 401 — work with aliases and combos. (thanks @thaiphuong1202)
  • feat(providers): expand the openai and gemini direct registries with first-class variants already known elsewhere — the openai provider entry now exposes gpt-4.1-mini, gpt-4.1-nano, o3-mini, and o4-mini (the latter two carry REASONING_UNSUPPORTED like o3), and the gemini entry now exposes gemini-2.0-flash-lite and gemini-3-flash-lite-preview. These models were already first-class throughout sibling subsystems (cost estimator, task fitness, free-model catalog, multiple aggregator registries) but happened to be missing from the direct openai/gemini namespaces. Embedding/TTS/image-gen models stay in their dedicated registries (embeddingRegistry.ts, audioRegistry.ts, imageRegistry.ts); legacy ids OmniRoute curated out (o1, gpt-4-turbo, …) are not restored. (thanks @East-rayyy)
  • feat(translator): OpenAI SSE → Gemini SSE conversion for /v1beta/models/{model}:streamGenerateContent — the @google/genai SDK (Gemini CLI) always calls :streamGenerateContent?alt=sse for chat and expects Gemini SSE chunks (no [DONE] sentinel — the stream just closes). The v1beta route was forwarding OpenAI SSE from handleChat unchanged, so the SDK crashed on the OpenAI [DONE] line with SyntaxError: Unexpected token 'D', "[DONE]" is not valid JSON. A new transformOpenAISSEToGeminiSSE() (in open-sse/translator/response/openai-to-gemini-sse.ts) rewrites each OpenAI delta into candidates[].content.parts[], maps finish_reason → finishReason (STOP / MAX_TOKENS / SAFETY), attaches usageMetadata + modelVersion on the final chunk, and surfaces reasoning_content as { thought: true } parts for thinking models. The non-streaming :generateContent action gets a sibling convertOpenAIResponseToGemini() for the JSON path. Streaming intent is now keyed off the URL action suffix (canonical Gemini convention) rather than the non-standard generationConfig.stream body field. (thanks @SteelMorgan)
  • feat(compression): unified compression configuration panel (Phase 1) — /dashboard/context/settings is now the single source of truth for compression: a master toggle plus per-engine on/off and level controls, with the dispatch pipeline derived from a stored engines map on CompressionConfig. A gate (enginesExplicit) ensures the new map only drives dispatch when an engines row was actually saved from the panel, so legacy/backfilled installs (the seeded default combo from migrations 042/043) keep their existing defaultMode behavior unchanged. The default-combo and per-engine routes are shimmed (410). (#4432 — thanks @diegosouzapw)
  • feat(mcp): register the web-session pool observability tools — the poolTools MCP tool set (web-session pool stats/health) was defined but never wired into createMcpServer(), so it was dead. It is now registered in server.ts with withScopeEnforcement against the typed read:health / write:resilience scopes (no enum inflation), giving MCP clients visibility into the pooled web-session lifecycle. (#4399, #3368 — thanks @diegosouzapw)
  • feat(providers): stronger no-auth and web-cookie provider validation (AUTH_007) — provider connection validation now handles no-auth and web-cookie providers explicitly: instead of returning a generic "Provider validation not supported", these providers report a precise AUTH_007 status so the dashboard surfaces actionable validation feedback for cookie/no-auth flows. (#4023 — thanks @oyi77)

🐛 Fixed

  • fix(embeddings): forward output dimensions to Gemini for consistent embedding dims. (thanks @nguyenha935)
  • fix(translator): sanitize Read tool args from non-Anthropic models to prevent retry loops. (thanks @GodrezJr2)
  • fix(usage): reuse Gemini CLI project ID for quota checks (avoid re-discovery). (thanks @Delcado19)
  • fix(dashboard): surface manual config CTA when Claude CLI detection fails (remote deployments). (thanks @anuragg-saxenaa)
  • fix(executors): granular reasoning_effort handling for Claude models on GitHub Copilot. (thanks @baslr)
  • fix(translator): strip Claude output_config before MiniMax (rejected upstream). (thanks @hiepau1231)
  • fix(translator): OpenAI audio input now reaches Gemini/Antigravity instead of being silently dropped — input_audio/audio content parts on the OpenAI→Gemini path matched no handler in convertOpenAIContentToParts and were discarded with no error. They are now mapped to a Gemini inlineData part with an audio/<format> mime type (wav, mp3, …). (thanks @mugnimaestra)
  • fix(combo): model lockout now honors a long upstream quota reset instead of retrying within minutes — when a combo target returned a quota error carrying an explicit long reset (e.g. Antigravity Resets in 160h27m24s, a Retry-After header), the per-model lockout capped at the short base cooldown (~minutes) and discarded the parsed reset, so the exhausted model kept being retried far too early. The lockout now applies the parsed reset when it exceeds the base cooldown, and the Antigravity error-message parser also matches the plural Resets in … phrasing. (thanks @Ansh7473)
  • fix(antigravity): Claude models no longer 400 with Unknown name "output_config" — Anthropic/Claude-Code-only fields (output_config, legacy output_format) leaked into the Google Cloud Code request envelope via its top-level field passthrough, and Google rejects unknown envelope fields with 400 Invalid JSON payload received. Unknown name "output_config" — breaking every Claude model served through Antigravity in IDEs. Those fields are now dropped before the envelope is built. (thanks @Duongkhanhtool)
  • fix(combo): round-robin members fail over faster under concurrency saturation via a configurable queue depth — when a round-robin combo member was saturated, requests sat in the per-model semaphore's unbounded queue and only failed over to the next member after the full queueTimeoutMs (default 30s) elapsed — so a burst of agentic requests deep-queued one hot member instead of spilling to healthy ones. The per-model semaphore now accepts a bounded queue depth and emits SEMAPHORE_QUEUE_FULL once it is full (the round-robin loop already cascades on that code), so a configured low depth fails over immediately. A new queueDepth combo-config knob (global default / provider override / per-combo, default 20 for backward compatibility; 0 = never queue → fail over now) is exposed in Settings → Combo Defaults. (#3872 — thanks @KooshaPari)
  • fix(pricing): align Claude Code (cc) pricing with current Anthropic per-MTok rates — the cc provider block in the default pricing table had stale numbers across every Claude 4.x family entry — most visibly, claude-opus-4-5-20251101 was billed at the deprecated Opus 4.1 rate (input $15 / output $75), and claude-haiku-4-5-20251001 was at half the current Haiku 4.5 rate. The cached (cache hit) and cache_creation (5-minute cache write) multipliers were also off across Opus 4.6/4.7/4.8, Sonnet 4.5/4.6, Haiku 4.5, and Fable 5. All eight entries now match the rates Anthropic publishes (input, 5m cache write at 1.25x input, cache hit at 0.1x input, output; reasoning billed at the output rate), so cost accounting on the dashboard and per-request usage events stop under- or over-reporting Claude Code spend. (thanks @chulanpro5)
  • fix(executors): sanitize Anthropic-shape content parts before GitHub Copilot /chat/completions — Claude models on GitHub Copilot driven from clients like Cursor IDE (e.g. gh/claude-sonnet-4.6) failed with Provider returned error: type has to be either 'image_url' or 'text' (reset after 30s) because the client passed through Anthropic-shape content parts (tool_use, tool_result, thinking) untouched, and the Copilot chat-completions endpoint only accepts text/image_url. GithubExecutor.transformRequest now serializes any unsupported part type as text (preserving the model's context), drops empty parts, and collapses to null when an assistant message's only content was tool_calls — tool_calls ride alongside untouched. Codex-family models still route through /responses unchanged. (thanks @cngznNN)
  • fix(sse): refactor stall detection to reduce false positives on slow but progressing streams. (thanks @zakirkun)
  • fix(executors): synthesize x-opencode-request for custom-named OpenCode providers — the OpenCode CLI only emits the x-opencode-* header set when the provider id starts with opencode; a custom-named provider (e.g. omniroute) instead sends x-session-affinity / x-session-id (mapped to x-opencode-session since fix(opencode-executor): x-opencode-* header forwarding is dead code #4022) but no request-correlation id, so x-opencode-request was silently dropped. OpencodeExecutor now synthesizes a fresh x-opencode-request on that session-affinity fallback path so custom-named providers are not disadvantaged on the opencode.ai upstream. x-opencode-client / x-opencode-project are intentionally not fabricated (no valid client source — an invented value risks upstream rejection) and remain forward-only; DefaultExecutor is untouched. (#4465 — thanks @pizzav-xyz)
  • fix(compression): RTK now compresses Anthropic-shape tool_result blocks — applyRtkCompression only compressed OpenAI-shape tool results (role:"tool"); Anthropic-shape tool results (tool_result content blocks inside a role:"user" message) were skipped, so coding agents speaking the Anthropic Messages format got zero RTK savings even though RTK's command-aware filters (e.g. git-status) would have compressed the output. RTK now treats a message containing a tool_result block as eligible (gated by applyToToolResults), captures Anthropic tool_use blocks for command resolution, and compresses each block's inner text (string or nested text-block array) while preserving type + tool_use_id exactly — matching what caveman/aggressive already did. (#4468 — thanks @diegosouzapw)
  • fix(dashboard): request-log auto-refresh no longer dies from a "ghost" load-more on first page load — the request-log viewer's infinite-scroll IntersectionObserver uses a 200px rootMargin, so its sentinel was already intersecting on mount whenever the first page didn't fill the scroll container. That fired a loadMore() with no user interaction, growing the window past PAGE_SIZE — and auto-refresh only polls while on the first page (limit <= pageSize), so it stayed permanently paused (only a manual filter change re-armed it). The observer now grows the window only after a genuine user scroll (new pure shouldTriggerInfiniteScroll guard), and a filter change re-arms the guard, so the default first-page view resumes its ~10s auto-refresh. (#4269 — thanks @tjengbudi)
  • fix(sse): large /v1/chat/completions requests no longer crash the server with a Node heap OOM — the chat request body was parsed multiple times along the route (route guard, injection guard, handler), buffering very large payloads several times and pushing concurrent agentic traffic into an out-of-memory crash. The body is now parsed once at the route guard and threaded through, so each request is buffered a single time. (#4380 — thanks @NakHalal)
  • fix(guardrails): tighten the system_prompt_leak heuristic to stop false positives on agent traffic — the leak detector flagged normal agent/tool conversations as prompt-leak attempts; it now requires an additional qualifier before flagging, so legitimate agent traffic is no longer blocked. (#4041 — thanks @KooshaPari)
  • fix(translator): drop orphan tool results on the Claude→OpenAI request path — a tool_result with no preceding matching tool_use (orphan) produced upstream 500/502 errors for Command Code / Custom OpenAI clients on ≥3.8.26. Orphan tool results are now filtered before the request is sent. (#4385 — thanks @adityapnusantara)
  • fix(providers): register API-key validators for Firecrawl and Jina Reader — both providers returned "Provider validation not supported" when validating their API key; they now have proper validators registered in SEARCH_VALIDATOR_CONFIGS. (#4401 — thanks @ponkcore)
  • fix(providers): generic web-cookie validator must not shadow per-provider validators — a follow-up to the AUTH_007 validation work (feat(providers): improve no-auth and web-cookie provider validation #4023): the generic web-cookie validator was matching before more specific per-provider validators, so provider-specific validation was skipped. Validator resolution now prefers the per-provider validator. (#4467 — thanks @diegosouzapw)
  • fix(translator): inject a placeholder message when the Responses API input[] is empty — a POST /v1/responses with input: [] translated to messages: [], which every upstream Chat-Completions provider rejects (surfaced as a confusing 406); a single placeholder user message is now injected, mirroring the existing empty-string handling. (#4393 — thanks @diegosouzapw)
  • fix(providers): serve the api.airforce live /models catalog instead of the stale seed — the api.airforce provider listed a stale hard-coded seed; it now serves the upstream live /models catalog. (#4395 — thanks @diegosouzapw)
  • fix(cli): non-interactive-safe prompts + context alias — the CLI's confirm()/prompt helpers no longer hang in non-interactive (piped/CI) contexts, and a singular context alias is accepted alongside contexts; the contexts workflow is documented. (#4439, #4397 — thanks @diegosouzapw)

🔒 Security

  • fix(sse): use a crypto-secure RNG for combo/deck load-balancing selection — random combo/deck member selection used a non-cryptographic PRNG, flagged by CodeQL (#665); it now uses a crypto-secure RNG. (#4457 — thanks @diegosouzapw)

📝 Maintenance

🙌 Contributors

Thanks to everyone whose work shipped in v3.8.32:

Contributor
@powellnorma community contributor
@thaiphuong1202 community contributor
@East-rayyy community contributor
@SteelMorgan community contributor
@nguyenha935 community contributor
@GodrezJr2 community contributor
@Delcado19 community contributor
@anuragg-saxenaa community contributor
@baslr community contributor
@hiepau1231 community contributor
@mugnimaestra community contributor
@Ansh7473 community contributor
@Duongkhanhtool community contributor
@KooshaPari community contributor
@chulanpro5 community contributor
@cngznNN community contributor
@zakirkun community contributor
@pizzav-xyz community contributor
@NakHalal community contributor
@adityapnusantara community contributor
@ponkcore community contributor
@tjengbudi community contributor
@oyi77 community contributor
@diegosouzapw maintainer

Quality Gate (local, pre-PR)

  • lint / typecheck:core / check:cycles / check:docs-all: pass
  • test:unit: 16611 pass (5 timer/timeout tests flaked under --test-concurrency=20 CPU starvation; all pass isolated on a quiet machine — not regressions)
  • test:vitest: 193 pass (17 files)
  • quality ratchet: eslintWarnings rebaselined 3839→3863 (documented cycle drift; release commit touches only CHANGELOG/README/baseline → 0 lint delta); openapiCoverage 37.9 within eps; i18n 78.4 == baseline

Full sharded CI runs on this PR — authoritative gate.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

…4397)

* chore(release): open v3.8.31 development cycle

* fix(mitm): exact host membership in MITM hosts test (CodeQL false positive) (#4386)

getMitmToolHosts returns string[], so .includes(host) is Array.prototype.includes
(exact membership). CodeQL's js/incomplete-url-substring-sanitization heuristic
misreads it as a String.includes() URL-substring sanitization check and raises a
HIGH alert. Switch to .some(h => h === host) — identical semantics, explicit intent,
no flagged pattern. Surfaced post-v3.8.30 (#4325) once the test landed on main.

Test-only change (no runtime behavior); the suite still passes and the CodeQL
re-scan on merge clears the alert.

* fix(codex): request reasoning summaries (#4359)

Adds reasoning.summary=auto + reasoning.encrypted_content include for Codex. Thanks @xz-dev.

* fix(embeddings): inject NVIDIA NIM input_type for asymmetric embed models (#4341)

NVIDIA NIM asymmetric embedding models (e.g. nvidia/nv-embedqa-e5-v5) reject
requests without an `input_type` ("query" | "passage") with 400 "'input_type'
parameter is required". The embedding registry now carries a model-level
default param for the asymmetric NVIDIA model, and the embeddings handler
injects a model's default params into the upstream body only when the client
omitted them, leaving a client-supplied value untouched.

Reported-by: hydraromania (decolua/9router#1378)

Co-authored-by: hydraromania <252583922+hydraromania@users.noreply.github.com>

* fix(api): migrate deprecated Codex [features].codex_hooks to [features].hooks (#4342)

Codex renamed the `codex_hooks` feature flag to `hooks`; recent Codex CLI
versions ignore the old key and warn. When OmniRoute rewrites an existing
config.toml (configure/reset Codex provider) it now renames
[features].codex_hooks -> [features].hooks, preserving the value and never
clobbering an already-present `hooks`, then drops the deprecated key. The
migration is a no-op when the flag is absent and runs on both the POST and
DELETE config paths.

Reported-by: Bian-Sh (decolua/9router#1327)

Co-authored-by: Bian-Sh <24520547+Bian-Sh@users.noreply.github.com>

* fix(translator): drop the null flush on the same-format response path (#4344)

The streaming response translator's same-format fast path returned
`[chunk]` unconditionally, so the end-of-stream null/flush signal
(chunk === null) propagated as a literal `[null]`. Downstream this surfaced
as an empty `data: null` SSE event between chunks and crashed strict clients
(e.g. Factory Droid BYOK on /v1/responses). The fast path now returns `[]`
for the null flush while still passing real chunks through unchanged.

Reported-by: thaitryhand (decolua/9router#1052)

Co-authored-by: thaitryhand <248103256+thaitryhand@users.noreply.github.com>

* fix(translator): strip assistant echo fields on the OpenAI target path (Mistral 422) (#4350)

Strict OpenAI-compatible upstreams (e.g. mistral/codestral-latest) reject
client-only assistant echo fields sent back as input with 422
extra_forbidden (the report hit messages[].assistant.reasoning_content via
Codex /responses). Only reasoning_content was stripped on the OpenAI target
path; the sibling fields reasoning / refusal / annotations / cache_control
leaked through. They are now all dropped on the non-reasoner OpenAI target
path. `audio` is intentionally preserved (OpenAI audio models reference a
prior assistant audio response by id; Mistral never emits audio).

Reported-by: xxy9468615 (decolua/9router#1649)

Co-authored-by: xxy9468615 <63351664+xxy9468615@users.noreply.github.com>

* fix(cli): honor TAILSCALE_AUTHKEY for non-interactive tailscale login (#4343)

* fix(cli): honor TAILSCALE_AUTHKEY for non-interactive tailscale login (port from 9router#1263)

startTailscaleLogin built `tailscale up` without ever reading
process.env.TAILSCALE_AUTHKEY, so a pre-authenticated / headless daemon
waited for an interactive auth URL and timed out (~15s). When
TAILSCALE_AUTHKEY is set it is now passed via `--auth-key=` (an argv element
to spawn(binary, args) — no shell interpolation, Hard Rule #13); when unset,
behavior is unchanged. The arg builder is extracted into a pure exported
`tailscaleUpArgs()` for testing.

Reported-by: ipeterpetrus (decolua/9router#1263)
Co-authored-by: ipeterpetrus <93033698+ipeterpetrus@users.noreply.github.com>

* chore(quality): rebaseline tailscaleTunnel.ts file-size to 1202 (#1263 +13)

---------

Co-authored-by: ipeterpetrus <93033698+ipeterpetrus@users.noreply.github.com>

* fix(dashboard): OAuth modal surfaces the real error on a non-JSON response (#4351)

* fix(dashboard): OAuth modal surfaces real error on non-JSON responses (port from 9router#1318)

The OAuth connect/reauth modal called `await res.json()` unconditionally, so
a non-JSON error response (e.g. a plain-text 500 page from a build/OAuth
endpoint) threw `Unexpected token 'I'...` and hid the real failure. New
shared helpers parseResponseBody / getErrorMessage (src/shared/utils/api.ts)
read the body safely (JSON when JSON, raw text otherwise) and produce a clean
message either way; every modal fetch site now uses them.

Reported-by: DNNYF (decolua/9router#1318)
Co-authored-by: DNNYF <74033321+DNNYF@users.noreply.github.com>

* fix(dashboard): type OAuth modal response body as Record<string, unknown> (t11 any-budget)

Switch the parseResponseBody casts from Record<string, any> to
Record<string, unknown> so OAuthModal.tsx stays within its t11 explicit-any
budget. getErrorMessage already takes unknown; the success paths typecheck
clean under strict:false. No runtime change.

---------

Co-authored-by: DNNYF <74033321+DNNYF@users.noreply.github.com>

* fix(translator): accept AI SDK-style { type: image, image: data-URL } content parts (#4345)

* fix(translator): accept AI SDK-style { type: image, image: "data:..." } parts (port from 9router#1330)

Several OpenAI-input translators only recognized images shaped as
`image_url.url` (or an object with `.source`/`.url`), so an AI SDK-style
content part where `image` is a bare data-URL STRING was silently dropped
before reaching a vision provider (OpenCode is one affected client; the gap
is generic). The OpenAI->Claude, OpenAI->Kiro and OpenAI->Gemini/Antigravity
translators now parse a string `image` data URL into each provider's native
image shape (Claude base64 source, Kiro images[].source.bytes, Gemini
inlineData).

Reported-by: mugnimaestra (decolua/9router#1330)
Co-authored-by: mugnimaestra <13349159+mugnimaestra@users.noreply.github.com>

* chore(quality): freeze openai-to-kiro.ts file-size at 807 (#1330 +9, over 800 cap)

---------

Co-authored-by: mugnimaestra <13349159+mugnimaestra@users.noreply.github.com>

* fix(dashboard): show a disabled connection's last error in the row (#4352)

* fix(dashboard): show a disabled connection's last error in the row (port from 9router#1447)

The provider card's error badge counts a disabled connection (isActive ===
false) that has an error — its effective status is still
error/expired/unavailable — but the connection row hid the lastError text for
disabled rows, so the operator saw the count without the cause. The row's
error-visibility decision is extracted into shouldShowConnectionLastError()
and now shows the error whenever there is one, regardless of the active
toggle.

Reported-by: ntdung6868 (decolua/9router#1447)
Co-authored-by: ntdung6868 <103993527+ntdung6868@users.noreply.github.com>

* chore(quality): rebaseline ConnectionRow.tsx file-size to 942 (#1447 +1 import)

---------

Co-authored-by: ntdung6868 <103993527+ntdung6868@users.noreply.github.com>

* fix(providers): bound the OAuth connection-test probe with a timeout (#4347)

* fix(providers): bound OAuth connection-test probe with a timeout (port from 9router#1449)

The OAuth path of "Test Connection One-by-One" called bare fetch() with no
AbortController/signal, so a provider probe that accepted the socket but never
responded wedged the test queue forever. Both the initial probe and the
post-refresh retry are now bounded with AbortSignal.timeout(30s) — matching the
API-key path's existing budget — and a timed-out probe resolves as a failure
with a clear "Test timed out after 30s" message in the route's normal error shape.

Reported-by: ntdung6868 (decolua/9router#1449)
Co-authored-by: ntdung6868 <103993527+ntdung6868@users.noreply.github.com>

* chore(quality): rebaseline providers test route file-size to 887 (#1449 + sibling #1444 growth)

---------

Co-authored-by: ntdung6868 <103993527+ntdung6868@users.noreply.github.com>

* fix(providers): label a deactivated account distinctly from a revoked token (#4353)

A Codex connection whose OAuth refresh is fully healthy but whose ChatGPT
account has been deactivated by the provider gets a 401 from the upstream
API. The connection test labeled that the same as a bad credential
("Token invalid or revoked" -> upstream_auth_error), so an operator could not
tell a deactivated account from a revoked token. The test now reads the
401/403 body and, when it indicates account deactivation, classifies it as
account_deactivated (which the dashboard already renders as "Account
Deactivated"); a plain auth 401 is unchanged.

Reported-by: ntdung6868 (decolua/9router#1444)

Co-authored-by: ntdung6868 <103993527+ntdung6868@users.noreply.github.com>

* fix(db): cascade-delete orphaned model aliases when a provider is removed (#4348)

* fix(db): cascade-delete orphaned model aliases when a provider is removed (port from 9router#1409)

Deleting a custom provider removed its connections and node but left the
imported model-alias rows (key=<alias>, value="<providerId>/<model>") behind,
so re-importing the same provider was blocked by stale "already exists" aliases.
Add a deleteModelAliasesForProvider(providerId) DB helper that drops every alias
whose stored value begins with "<providerId>/", and call it from the provider-node
DELETE handler so a fresh import is unblocked.

Reported-by: nguyenvanhuy0612 (decolua/9router#1409)
Co-authored-by: nguyenvanhuy0612 <57367674+nguyenvanhuy0612@users.noreply.github.com>

* chore(quality): rebaseline models.ts file-size to 1221 (#1409 + sibling #1294 growth)

---------

Co-authored-by: nguyenvanhuy0612 <57367674+nguyenvanhuy0612@users.noreply.github.com>

* fix(api): persist max_input_tokens/max_output_tokens when adding a custom model (#4349)

The POST /api/provider-models handler read the rest of the body but never
the two token-limit fields, and addCustomModel() had no parameter for them,
so the form values were dropped on write while the DB layer and /v1/models
catalog already round-trip inputTokenLimit/outputTokenLimit. Accept the two
optional limits in the schema, forward them through the handler, and persist
them in addCustomModel(). TDD: failing-then-passing unit test.

Reported-by: codename-zen (decolua/9router#1294)

Co-authored-by: codename-zen <263238141+codename-zen@users.noreply.github.com>

* docs: feature-documentation catch-up (v3.8.20 → v3.8.30) (#4391)

One-time reconciliation of the docs with every user-facing feature shipped since
v3.8.20 (we had never done a dedicated pass, so debt had accumulated):

- README: new '✨ What's New' section (curated v3.8.20→v3.8.30 highlights).
- New guides: CLI-INTEGRATIONS (all setup-*/launch commands), MITM-TPROXY-DECRYPT
  (transparent-decrypt epic), CONTEXT_EDITING (delegated Anthropic clear_tool_uses).
- Refreshed: AUTO-COMBO (auto/<category>:<tier> + Arena-ELO), API_REFERENCE
  (x-omniroute-no-memory), MEMORY (int8 quantization + off-by-default), RESILIENCE
  (model-lockout success-decay), RTK, AGENTBRIDGE, TRAFFIC_INSPECTOR, GUARDRAILS,
  CLOUD_AGENT, ENVIRONMENT, SETUP_GUIDE, CLI-TOOLS, MCP-SERVER.
- Regenerated PROVIDER_REFERENCE (231 providers); synced the count in README/CLAUDE/AGENTS.
- Allowlisted external-tool env vars (OPENAI_API_BASE, PROMPTFOO_PROVIDER_KEY) and the
  STREAM_RECOVERY config-object name in the docs-accuracy gates.

All claims source-verified; check:docs-all (sync/counts/env/links/fabricated) passes.
Going forward this runs every release via generate-release step 6b.

* fix(executors): don't inject thinking when tool_choice forces a tool (native Claude) (#4389)

Forced tool_choice now strips the adaptive thinking injection to avoid Anthropic 400. Thanks @NomenAK.

* fix(translator): Gemini accepts HTTP/HTTPS image URLs (port from 9router#344) (#4373)

OpenAI-style `image_url` parts with an `http://` or `https://` URL reached
`convertOpenAIContentToParts` and were dropped with only a `console.warn`,
because Gemini's `inlineData` requires base64 (the helper is synchronous and
cannot fetch+encode upstream assets). Gemini's `Part` schema, however, natively
accepts `fileData: { fileUri }` for remote URIs — the model fetches the asset
itself.

The helper now emits a `fileData` part (`mimeType: "image/*"`, inferred upstream
on fetch) for HTTP/HTTPS URLs instead of silently dropping them. Vision requests
that pass a URL — not a data: URI — now reach Gemini intact.

No behavioral change for:
- `data:` URIs → still emitted as `inlineData` with the parsed media type.
- Unsupported schemes (e.g. `ftp:`) → still skipped (Gemini would reject them).

The openai-to-claude side already passed HTTP/HTTPS URLs through as
`source: { type: "url", url }` (lines 573–578) — the upstream PR's Claude-side
change was already covered.

Regression test: tests/unit/gemini-helper-http-image-url-port344.test.ts
(4 cases: https URL, http URL, data: URI no-regression, unsupported-scheme guard).


Inspired-by: decolua/9router#344

Co-authored-by: Ibrahim Ryan <ryan@nuevanext.com>

* fix(executors): strip stream_options for qwen non-streaming / thinking Claude Code requests (port from 9router#663) (#4374)

Claude-Code-compatible providers force the executor-level `stream` flag on
via `upstreamStream = stream || isClaudeCodeCompatible`
(open-sse/handlers/chatCore.ts), but the outgoing body keeps the caller's
original `stream: false`. The shared `stream && targetFormat === "openai"`
branch in DefaultExecutor.transformRequest then injected
`stream_options: { include_usage: true }` onto a body that still said
`stream: false`, and qwen upstream rejected the request with
`400 "'stream_options' only set this when you set stream: true"`. The same
rejection surfaced when the body carried `thinking` / `enable_thinking`.

The qwen branch now skips the injection (and strips any client-sent
`stream_options`) when the body explicitly says `stream: false` or
requests thinking, leaving regular qwen streaming requests with the
include_usage injection intact. Other providers are unaffected.

Adds a TDD regression with 4 cases covering both opt-out paths and the
normal-streaming positive control.


Inspired-by: decolua/9router#663

Co-authored-by: anuragg-saxenaa <anuragg.saxenaa@gmail.com>

* fix(security): scope OAuth callback postMessage to a trusted-origin allowlist (port from 9router#998) (#4372)

The OAuth callback at `/callback` previously fell back to
`window.opener.postMessage({ code, state, ... }, "*")` whenever the opener
was cross-origin. The fallback was intended to support remote-OmniRoute +
local-loopback callbacks (where opener and callback live on different
origins), but the same code path also delivers the OAuth code/state to any
hostile opener that pops the well-known callback URL — letting that
attacker complete the OAuth flow as the user.

Replace the wildcard fallback with iteration over a fixed allowlist:
`window.location.origin` (same-origin parent — the popup-mode dashboard)
plus Codex's fixed loopback helper (`http://localhost:1455` and the IPv4
literal `http://127.0.0.1:1455`). The browser drops `postMessage` to any
opener whose actual origin is not in `targetOrigin`, so the message reaches
only known parents and is silently dropped for any other. The same-origin
fallback path is unchanged — methods 2 (`BroadcastChannel`) and 3
(`localStorage` storage event) still cover same-origin openers that COOP
severed.

The `openerSameOrigin` probe stays in place to drive the auto-close vs
manual-copy UI decision (no behavior change for the success path).

Adds a regression test (`tests/unit/ui/oauth-callback-postmessage-scope.test.tsx`)
that mounts the page with a stubbed cross-origin opener and asserts no
`postMessage` call ever uses `"*"` and every call lands on an allowlisted
origin. The test failed against the pre-fix code (red), passes after the
fix (green) — TDD per CLAUDE.md hard rule #18.

Partial port of upstream decolua/9router#998: the upstream PR also
re-enabled TLS verification on a DNS-bypass fetch in `open-sse/utils/proxyFetch.js`;
that part is N/A here because OmniRoute's `proxyFetch.ts` never disabled
TLS verification (no `rejectUnauthorized: false` anywhere in the file).


Inspired-by: decolua/9router#998

Co-authored-by: aeonframework <aeon@aeonframework.dev>

* fix(sse): default combo per-target timeout to 120s for fast failover (#4365)

Combo per-target timeout inherited the full FETCH_TIMEOUT_MS (600s) when a
combo did not set its own targetTimeoutMs, so a single hung/slow target stalled
the whole combo for up to 10 minutes before falling through to the next model.

Introduce DEFAULT_COMBO_TARGET_TIMEOUT_MS (120s) as the unset-default in
resolveComboTargetTimeoutMs (new 3rd arg) and wire it in phaseComboSetup. The
upstream ceiling (600s) and per-combo opt-out (targetTimeoutMs, up to the
ceiling) are preserved; single non-combo requests are unchanged. For streaming
requests this only bounds time-to-first-headers, so token generation is not cut
short.

TDD: failing-then-passing unit test in tests/unit/combo-config.test.ts.

* refactor(combo): de-dup exhausted-target skip predicate across both dispatchers (#4362)

Primeiro incremento da de-dup dos 2 dispatchers de combo (handleComboChat +
handleRoundRobinCombo). O bloco de pre-check #1731/#1731v2 (skip de target já
exhausted no provider/connection) era BYTE-IDÊNTICO nos dois (mesmas condições,
mesmas mensagens), diferindo só na tag de log e no control-flow.

- comboPredicates.ts: getExhaustedTargetSkipReason(target, exhaustedProviders,
  exhaustedConnections) — predicate PURO que retorna a mensagem de skip (ou null);
  cada dispatcher mantém seu próprio log-tag + control-flow (return null / continue)
  + fallbackCount. No mutate do stryker (cobertura de mutação).
- combo.ts: −20 linhas (os 2 blocos viram 1 chamada cada).
- 7 testes de caracterização travam condições + strings exatas.

Comportamento preservado: 376/376 testes combo (357 caracterização + 7 novos),
integração sse-correctness 5/5, typecheck 0, complexity neutro (1895), file-size
encolhe. Próximo incremento: de-dup do error-handling/exhausted-tracking (handleTargetError).

* refactor(combo): de-dup upstream-error exhaustion classification across both dispatchers (#4366)

Segundo incremento da de-dup dos 2 dispatchers (handleTargetError). Após cada erro
de target, ambos rodavam um bloco quase-idêntico que marca o provider exhausted
(#1731), a conexão connection-errored (#1731v2) ou o provider transiently rate-limited.

- combo/targetExhaustion.ts: applyComboTargetExhaustion(target, opts) — atualiza os
  3 Sets de exhaustion e retorna providerExhausted. As MUTAÇÕES de Set (que dirigem o
  skip de targets, lidas por getExhaustedTargetSkipReason) são BYTE-IDÊNTICAS nos dois;
  as diferenças reais viram parâmetros: tag, allAccountsRateLimited (termo extra do RR,
  false no handleComboChat), exhaustedLogLevel (info no handleComboChat, debug no RR).
  Connection-level extraído p/ markConnectionLevelExhaustion (privado, <15 complexity).
- combo.ts: −73 linhas; 4 imports órfãos removidos.
- 7 testes de caracterização travam as mutações + o return.

ÚNICA mudança de comportamento: o WORDING das mensagens de log do RR ganha o sufixo
'on remaining targets' (cosmético; mesmo #code, mesmas mutações, mesmos níveis de log).
376/376 combo (caracterização preservada), integração sse 5/5, typecheck 0, complexity
neutro (1895), file-size encolhe.

* refactor(chatCore): extract checkHeapPressureGuard leaf (god-file decomposition start) (#4371)

Primeiro incremento da decomposição do chatCore.ts (5127 LOC, hot-path mais quente).
O guard de memória do topo do handleChatCore (rejeita 503 quando o heap V8 passa o
threshold de shed) vira um leaf testável, co-locado com o threshold em heapPressure.ts.

- heapPressure.ts: checkHeapPressureGuard(heapUsedMb, thresholdMb) — retorna o result
  503 pronto ou null. Byte-idêntico ao guard inline (mesmo check, mesma 503, mesmo warn).
  A figura de heap fica em telemetria INTERNA, nunca no response do cliente (Hard Rule #12).
- chatCore.ts: o bloco inline (~22 ln) vira 3 linhas; import órfão de HEAP_PRESSURE_THRESHOLD_MB
  trocado por checkHeapPressureGuard.
- 3 testes novos (incl. assert Rule #12: o MB medido não vaza no payload).

complexity-baseline 1895->1896: drift de base pós-#4338 (medido com minhas mudanças
stashed = 1896); esta mudança é complexity-NEUTRA (helper complexity 2, handleChatCore só
perde código). 190/190 chatcore tests, typecheck 0, file-size encolhe.

* Localize CLI and stabilize fetch, memory, and coverage handling (#4383)

en-only i18n, fetch-start-timeout hardening, EngineConfigPage icon fix, CI build-artifact-reuse overhaul. Memory production hunk dropped as a no-op (tests kept). Thanks @JxnLexn.

* test(combo): reset circuit breakers between stream-readiness cases (restore green) (#4396)

The combo-dispatch cases in combo-stream-readiness-fallback.test.ts deliberately
fail `glm` (zombie streams / repeated 504s), which legitimately trips the
per-provider circuit breaker. That OPEN state is a module-level singleton, so it
leaked into the next test and combo.ts then SKIPPED `glm/*` targets entirely
("Skipping … circuit breaker OPEN"). That made "combo does not retry stream
readiness timeouts on the same model" never attempt glm/zombie — expected
['glm/zombie','openai/gpt-5.4-mini'] but got ['openai/gpt-5.4-mini'].

This was a pre-existing red on release/v3.8.31 (present at the cycle-open tip),
order-dependent: the test passes in isolation, fails after the preceding cases.
Add a test.beforeEach(resetAllCircuitBreakers) so each scenario starts from a
clean breaker slate. Test-isolation only — the breaker behavior is correct and
no production code or assertion changes. Full combo suite: 390/390 green.

* fix(cli): decline confirm() cleanly on non-interactive stdin + document contexts workflow

The `contexts remove` command already has `--yes` to skip confirmation, but when
run without it under a non-interactive stdin (pipe, CI, EOF) the [y/N] prompt could
never be answered — the readline question stayed pending and Node warned about an
"unsettled top-level await" at exit. confirm() now detects `!process.stdin.isTTY`
and declines cleanly (returns false), pointing at `--yes` for non-interactive use.
Exported confirm() for a regression test.

Docs: REMOTE-MODE.md gains a full "Managing contexts" section (list/current/use to
switch between remote and local, add/show/rename, remove with --yes, export/import),
fixes a `context current` -> `contexts current` typo, and the README remote-mode
snippet now shows switching back to local. Verified against the live CLI: command
signatures, --yes, and the non-TTY decline path all behave as documented.

Tests: cli-contexts.test.ts asserts confirm() declines on non-TTY stdin (RED before,
GREEN after). All docs gates (fabricated/links/symbols) pass.

* ci(t11): bump any-budget for executors/base.ts (2 false-positive "any" strings)

Unblocks the Fast Quality Gates on release/v3.8.31: `check:any-budget:t11` was red on
`open-sse/executors/base.ts` for ALL PRs (pre-existing base drift, unrelated to this
branch). The checker counts `\bany\b` after stripping comments but NOT strings, and the
native-Claude tool_choice logic uses the API value `"any"` in two string literals
(`tb.tool_choice === "any"`, `.type === "any"`). There are zero actual TypeScript
`any` types in the file — budget set to the matched count, mirroring the existing
cursor.ts false-positive entry right below it.

* perf: combos UI split + next config + 1-click redis + bifrost sidecar (#3932) (#4381)

Combos UI split + next.config perf + 1-click local Redis launcher + bifrost relay. Review fixes (co-author): --rm/--restart conflict, error sanitization, UI/route names, dead-guard re-doc, bifrost Zod, IPv6 + CLI test fixes. Thanks @KooshaPari.

* fix(plugin): prefix OC static-catalog combo+raw keys with providerId (#4384)

OC parses model ids on '/'; combo keys now carry 'omniroute/' (was 'combo/'). Live-validated against the VPS via OpenCode. Thanks @herjarsa.

* docs(env): document TAILSCALE_AUTHKEY (env/docs contract drift on .31)

Second pre-existing base-drift fix needed to get Fast Quality Gates green: the
`repository contract is in sync` test (check:env-doc-sync) was red on ALL .31 PRs.
PR #4343 added a code reference to `process.env.TAILSCALE_AUTHKEY`
(src/lib/tailscaleTunnel.ts) for non-interactive `tailscale up`, but never added the
var to .env.example / docs/reference/ENVIRONMENT.md — the contract requires code vars
to appear in both. Add the (commented) entry to .env.example next to TAILSCALE_BIN and
a row to ENVIRONMENT.md. Verified: `node scripts/check/check-env-doc-sync.mjs` → in sync.

Unrelated to this branch's confirm()/docs change; surfaced because touching a quality
script triggers the TIA fail-safe full unit suite.

* chore(release): v3.8.31 — 2026-06-20

Finalize the v3.8.31 release: reconcile the CHANGELOG (full commit-to-bullet
coverage, 26 bullets across Features/Fixed/Security/Maintenance), refresh the
README What's New section, back-fill the .30/.31 sections into the 41 i18n
CHANGELOG mirrors, and add TAILSCALE_AUTHKEY to the env contract
(.env.example + ENVIRONMENT.md).

* chore(release): align mitm-hosts test comment with main to clear merge conflict

The same CodeQL false-positive fix landed twice — #4386 on release/v3.8.31 and
#4387 directly on main — with only the explanatory comment differing (the
assertion is byte-identical). Adopt main's comment wording on the release branch
so the release PR merges without a comment-only conflict.

* test(translator): align stale openai->gemini remote-URL tests with #4373

Third pre-existing base-drift fix to get Fast Quality Gates green on .31. PR #4373
('Gemini accepts HTTP/HTTPS image URLs', port of 9router#344) intentionally changed
convertOpenAIContentToParts so remote http(s) image URLs pass through as a native
`fileData: { fileUri }` part instead of the old #2807 drop+warn — and added its own
test (gemini-helper-http-image-url-port344.test.ts) for the new behavior. But it left
three tests in translator-openai-to-gemini.test.ts asserting the OLD drop+warn
contract, so they fail deterministically (verified: the file fails in isolation on
clean .31). These only surface under the TIA fail-safe FULL suite, which a quality-
script touch triggers.

Align the three stale tests to the real, intended behavior (captured by running the
function): remote URLs -> fileData.fileUri (mimeType image/*), still never inlineData
(the sync path cannot fetch+encode). This is test-vs-code alignment to a deliberate,
separately-tested change — not a weakened assertion. 40/40 in the file; 94/94 across
the translator + env-doc combo that previously failed.

* chore(quality): reconcile complexity baseline 1896->1900 (/review-prs v3.8.31 batch) (#4410)

* fix(release): reconcile full-CI drift for v3.8.31 (gemini tests #4373, any-budget #4389, masking allowlist #4384)

The release PR's full CI surfaced cumulative cycle drift the per-PR fast gates
skip:
- tests/unit/translator-openai-to-gemini.test.ts: realign 3 cases to #4373's
  HTTP/HTTPS-URL fileData pass-through (they asserted the old warn-and-drop).
- scripts/check/check-t11-any-budget.mjs: base.ts budget 0→2 — #4389 compares
  tool_choice against the string literal "any" (not a TS any type).
- config/quality/test-masking-allowlist.json: allowlist #4384's opencode combos
  net-assert reduction (obsolete combo/ namespace removed).
No production behavior change.

* chore(release): re-trigger full CI for v3.8.31 finalization

Force a fresh pull_request CI run on the head carrying the cycle-drift fixes
(gemini #4373 tests, any-budget #4389, masking allowlist #4384) — the prior
synchronize event did not spawn a ci.yml run.

* chore: re-trigger CI (no Actions runs registered for prior push)

---------

Co-authored-by: Xiangzhe <32761048+xz-dev@users.noreply.github.com>
Co-authored-by: hydraromania <252583922+hydraromania@users.noreply.github.com>
Co-authored-by: Bian-Sh <24520547+Bian-Sh@users.noreply.github.com>
Co-authored-by: thaitryhand <248103256+thaitryhand@users.noreply.github.com>
Co-authored-by: xxy9468615 <63351664+xxy9468615@users.noreply.github.com>
Co-authored-by: ipeterpetrus <93033698+ipeterpetrus@users.noreply.github.com>
Co-authored-by: DNNYF <74033321+DNNYF@users.noreply.github.com>
Co-authored-by: mugnimaestra <13349159+mugnimaestra@users.noreply.github.com>
Co-authored-by: ntdung6868 <103993527+ntdung6868@users.noreply.github.com>
Co-authored-by: nguyenvanhuy0612 <57367674+nguyenvanhuy0612@users.noreply.github.com>
Co-authored-by: codename-zen <263238141+codename-zen@users.noreply.github.com>
Co-authored-by: Anton <39598727+NomenAK@users.noreply.github.com>
Co-authored-by: Ibrahim Ryan <ryan@nuevanext.com>
Co-authored-by: anuragg-saxenaa <anuragg.saxenaa@gmail.com>
Co-authored-by: aeonframework <aeon@aeonframework.dev>
Co-authored-by: Jan Leon <Jan.gaschler@gmail.com>
Co-authored-by: KooshaPari <42529354+KooshaPari@users.noreply.github.com>
Co-authored-by: Hernan Javier Ardila Sanchez <hjasgr@gmail.com>
@github-actions

github-actions Bot commented Jun 20, 2026 •

Copy link
Copy Markdown
Contributor

CI Coverage Report

  • Coverage job: skipped
  • PR test policy: failure

Coverage artifact was not available for this run.

PR Test Policy

This PR changes production code in src/, open-sse/, electron/, or bin/ without accompanying automated tests.

diegosouzapw and others added 15 commits June 20, 2026 16:48
…ver (#3872) (#4390)

Round-robin combo members deep-queued under concurrency saturation: the
per-model rate-limit semaphore had an unbounded queue and only emitted
SEMAPHORE_TIMEOUT after the full queueTimeoutMs (default 30s), so failover to
the next combo member happened far too late (or the client died first).

The per-model semaphore now accepts a bounded queue depth and rejects with
SEMAPHORE_QUEUE_FULL once the queue is full — the round-robin loop already
cascades to the next member on that code, so a low depth fails over immediately.
A new `queueDepth` combo-config knob (global default / provider override /
per-combo; default 20 for backward compatibility, 0 = never queue) is plumbed
via a resolveComboQueueDepth helper and surfaced in Settings → Combo Defaults.

TDD: rateLimitSemaphore.test.ts (bounded queue + SEMAPHORE_QUEUE_FULL, RED
before the maxQueueSize cap) and combo-config.test.ts (queueDepth cascade,
helper clamps, schema range).

Co-authored-by: KooshaPari <KooshaPari@users.noreply.github.com>
The pool tools (omniroute_pool_status/sessions/reset/warm/health) shipped in
open-sse/mcp-server/tools/poolTools.ts but were never imported or registered in
server.ts, so they were defined-but-dead — the PR4 observability step of the
#3368 web-session roadmap was one wiring step from live.

Wire them through the standard registration loop (import + tool-count tally +
RESERVED_MCP_NAMES + Object.values(poolTools).forEach), add per-tool scopes
(read:health for status/sessions/health, write:resilience for reset/warm) to
both the inline tool defs and the canonical MCP_TOOL_SCOPES map. The forEach is
typed structurally (no new no-explicit-any) since the shape is pinned by tests.

Tests: extend mcp-tool-collections-shape with poolTools; add a dedicated guard
pinning the server.ts wiring (import/registration/reserved-name), the scope
contract (inline == MCP_TOOL_SCOPES, all in MCP_SCOPE_LIST), and live handler
behavior against the in-memory PoolRegistry.
…phase slice) (#4392)

Segundo incremento da decomposição do chatCore.ts (5127 LOC). Extrai a primeira
fatia PURA da fase de request-setup do handleChatCore: a metadata de roteamento de
modelo (apiFormat + targetFormat via structural-narrowing sem 'as any', + o
requestedModel com fallback) — para um leaf testável em chatCore/requestSetup.ts.

- resolveChatCoreRequestSetup(modelInfo, body, model): ChatCoreRequestSetup — puro,
  byte-idêntico ao bloco inline. Tipo ChatCoreRequestSetup é a semente do futuro
  ChatCoreContext (as fases maiores — setup/dispatch/streaming — virão depois).
- chatCore.ts: bloco inline (~21 ln) vira 5 linhas.
- 5 testes de caracterização (narrowing + fallback do requestedModel).

185/185 chatcore tests (caracterização do handleChatCore preservada), typecheck 0,
complexity OK 1896 (mudança neutra), file-size encolhe.

Timing: feito como leaf PURO seguro porque a v3.8.30 está finalizando — o
ChatCoreContext mutável (carrega body + 20 inputs no hot-path) cabe melhor em v3.8.31.

Follow-up: atribuir requestSetup.ts a um batch de mutação no nightly (como os demais
chatCore leaves); 5 testes cobrem as branches por ora.
…ale seed (#4395)

api-airforce carries a real live https://api.airforce/v1/models catalog but was
left out of NAMED_OPENAI_STYLE_PROVIDERS, so the import route served its small
hardcoded seed (grok-3, grok-2-1212, claude-3.7-sonnet …) — models that no
longer exist upstream, so chat failed even though the connection test passed.

Add api-airforce to NAMED_OPENAI_STYLE_PROVIDERS (same shape as #4249/#4202/#3976
and the provider-model-sweep rows) so import does a live <baseUrl>/models fetch.
The registry seed stays as the offline fallback, so import never breaks if the
upstream is unreachable.

TDD: extends tests/unit/provider-sweep-live-discovery.test.ts (red proven: the
api-airforce case did not probe upstream and served local_catalog).
…ass missing variants (#4394)

OmniRoute already references gpt-4.1-mini, gpt-4.1-nano, o3-mini, and o4-mini
throughout sibling subsystems (cost estimator, modelCapabilities, taskFitness,
free-model catalog, multiple aggregator registries) but the direct openai
provider entry exposes only o3/gpt-4.1. Symmetrically, the gemini entry has
gemini-2.0-flash and gemini-3.1-flash-lite-preview but not the lite/lite-preview
siblings that pair with them. Inspired by an upstream effort to refresh both
static lists; scope is intentionally minimal:

- ADD openai: gpt-4.1-mini, gpt-4.1-nano, o3-mini, o4-mini
  (o3-mini/o4-mini get REASONING_UNSUPPORTED, matching the o3 entry)
- ADD gemini: gemini-2.0-flash-lite, gemini-3-flash-lite-preview

SKIP (already present elsewhere — kept in their typed registries):
- TTS (tts-1, tts-1-hd, gpt-4o-mini-tts)            -> audioRegistry.ts
- OpenAI embeddings (text-embedding-3-small/-large/ada-002) -> embeddingRegistry.ts
- Gemini embeddings (text-embedding-00*, gemini-embedding-*) -> embeddingRegistry.ts
- gemini-3.1-flash-image-preview                   -> imageRegistry.ts

SKIP (deliberate OmniRoute curation):
- gpt-5/gpt-5-mini/gpt-5-nano/gpt-5.1/gpt-5.2 (superseded by gpt-5.4 series)
- o1, o1-mini (superseded by o3)
- gpt-4-turbo (legacy; gpt-4o is the canonical mid-tier)
- o3-pro (no references anywhere in OmniRoute)

TDD: tests/unit/provider-registry-openai-gemini-expanded-models.test.ts
asserts the new ids and pins the existing ids so a future refactor cannot
silently drop them.


Inspired-by: decolua/9router#398

Co-authored-by: Ibrahim Ryan <ryan@nuevanext.com>
…] is empty (#4393)

A client (e.g. Fabric-AI) calling POST /v1/responses with `input: []`
used to be translated into `messages: []`, which every upstream Chat-
Completions provider rejects with `400: at least one message is required`
(surfaced to the client as a confusing 406).

Mirror the existing empty-string handling: when input[] is empty, inject
a single placeholder user message so the request is always valid. Touches
`openaiResponsesToOpenAIRequest` only (the same translator that already
covers the empty-string case at the top of the function), preserves any
caller-supplied `instructions` as a system message, and leaves every
non-empty input path byte-identical.

Adds a TDD test reproducing the empty input[] case (RED before fix, GREEN
after) and updates one pre-existing test that enshrined the broken
behavior (`assert.deepEqual(call.body.messages, [])`) to assert the new
correct shape.


Inspired-by: decolua/9router#419

Co-authored-by: anuragg-saxenaa <anuragg.saxenaa@gmail.com>
… 1) (#3455)

Integrated into release/v3.8.32 — operational docs (usage/quota, open-sse architecture, database, monitoring).
…4430)

Future-reference 'Quick end-to-end check' in REMOTE-MODE.md: connect -> mint scoped
token -> route a command -> switch back -> tear down, with a placeholder host. Makes
explicit (and verified live) that `contexts remove` is LOCAL-only — it drops the saved
credential but does NOT revoke the server-side token; use `tokens revoke <id>` to kill
access. Also documents that `--yes` is required for non-interactive removal and that
removing the active context falls back to default.
…07 (#4023) (#4023)

Integrated into release/v3.8.32 — web-cookie + no-auth provider validation (AUTH_007 SESSION_EXPIRED detection).
…/off + level (Phase 1) (#4432)

* docs(compression): design spec for the unified compression config panel

Engine-centric central panel (single source for master + per-engine on/off + level),
default pipeline derived from the engines map; Combos page owns named ordered pipelines +
active-profile selection; per-engine pages keep only detailed config. Phased: (1) core
consolidation + Model A migration, (2) named profiles + active selector, (3) per-request
x-omniroute-compression header.

* docs(compression): Phase 1 implementation plan for the unified config panel

12 TDD tasks: engine catalog, engines map + activeComboId, deriveDefaultPlan,
migration 102 + backfill, resolveCompressionPlan, runtime wiring, API, default-combo
shim, the engine-grid panel, consolidation, menu. Ends with a recap + the pending
follow-ups (Phase 2 named profiles/active selector, Phase 3 header, deferred items).

* feat(compression): engine catalog metadata (levels, single-mode, order)

* feat(compression): add engines map + activeComboId to CompressionConfig

* feat(compression): deriveDefaultPlan (engines map -> mode/pipeline)

Pure function that converts a per-engine EngineToggle map + masterEnabled
flag into a { mode, stackedPipeline } plan: off when master is off or no
engines on; single-mode when exactly one single-mode engine is enabled;
stacked (sorted by stackPriority) otherwise.

* feat(compression): persist+backfill engines map and activeComboId (migration 102)

* feat(compression): resolveCompressionPlan precedence resolver (header>override>active>default)

* feat(compression): selectCompressionStrategy uses resolveCompressionPlan

* feat(api): settings/compression carries engines map + activeComboId

Extends compressionSettingsUpdateSchema with engines (Record<string,{enabled:boolean,level?:string}>)
and activeComboId (string|null) so the PUT route accepts and persists these fields.
GET already returns the full getCompressionSettings() object which includes both fields.

TDD: added route round-trip tests (PUT+GET) for engines, activeComboId, null clear,
and schema rejection of invalid engines shape.

* refactor(compression): default-combo route is a read-only shim (default derived from engines)

* feat(dashboard): engine-grid compression panel (single source for on/off + level)

* refactor(dashboard): remove duplicate compression toggles; per-engine pages keep only detailed config

* feat(compression): unified panel menu order + derived-pipeline integration coverage

* fix(compression): import ENGINE_IDS via types re-export so it resolves under vitest

The bare "@omniroute/open-sse/.../engineCatalog.ts" specifier resolves under tsc/tsx
but not under vitest's MCP config: Vite externalizes a brand-new open-sse module to
Node, which can't load the .ts subpath. types.ts is already in Vite's graph, so route
ENGINE_IDS through its re-export. Fixes 3 failing MCP vitest suites (cacheTools,
dbHealthTool, essentialTools).

* fix(compression): engines map drives dispatch only when explicitly panel-saved

Code-review finding: the legacy seeded default combo (present on every install via
migration 042/043) was silently overriding a panel-configured engines map — the
default-combo block in chatCore set compressionComboApplied=true, skipping the
engines-map override, so an operator's panel toggles were ignored.

Gate the engines-map path on a new runtime-only CompressionConfig.enginesExplicit
(true when a stored engines row exists). Panel-saved installs: the engines map is
authoritative (deriveDefaultPlanFromConfig + new enginesMapDerivesStackedPipeline
guard the chatCore default-combo block). Legacy/backfilled installs: the map is
display-only and dispatch stays on the historical defaultMode/default-combo path —
zero behaviour change until the operator opts in via the panel.

Also documents why resolveCacheAwareConfig's getCacheAwareStrategy(config.defaultMode)
arg is safe (skipSystemPrompt is mode-independent).

* docs(compression): add MDX frontmatter + escape inline brace to fix fumadocs build

docs/compression/**/*.md is compiled as MDX by the Next build (fumadocs,
source.config.ts). The two planning docs lacked the required `title`
frontmatter (build error: 'title: expected string, received undefined') and the
design doc had a bare {rtk,caveman} that MDX parsed as a JSX expression. Adds
frontmatter matching the sibling docs and backticks the brace. Fixes the
dast-smoke build step.
@KooshaPari

Copy link
Copy Markdown
Contributor

4x sessions open at home either proactively finding items to PR or making PRS for open issues. This release will be interesting if they finish on time.

)

* fix(cli): non-interactive-safe prompts + singular `context` alias

Two CLI-polish follow-ups (the third — an async-warn refactor for #2807 — is moot:
#4373 replaced the Gemini remote-URL drop+warn with fileData pass-through, so no
async warn remains).

1) createPrompt (io.mjs) — the shared interactive helper used rl.question with no
   EOF guard, so a non-interactive stdin (pipe, CI, `< /dev/null`) left the await
   pending and Node warned about an 'unsettled top-level await' while the command
   hung. ask/askSecret now resolve on the readline `close` event (fired on EOF)
   with the default / empty string. A genuinely piped line still arrives via the
   question callback first, so `echo value | omniroute …` keeps working — only the
   no-input EOF case falls back. This fixes every command that prompts, centrally
   (mirrors the contexts `confirm()` fix from #4397).

2) `contexts` gains a singular `context` alias — the connect output and older docs
   said `omniroute context current`; the alias keeps that muscle-memory working.

Tests: cli-io.test.ts (ask default/empty + askSecret resolve on EOF, no hang) and a
cli-contexts.test.ts case asserting the `context` alias is registered. 11/11.

* test(cli): mock .alias() in the fake program for the contexts subcommand test

The existing 'registers a current subcommand' test uses a minimal fake commander
program; registerContexts now calls .alias("context"), so the fake needs an alias()
stub (returns this) or the chain throws. Add it. (Self-introduced by the alias in
this branch; my 4 new tests already pass in CI.)
…rows (restore-green) (#4447)

Integrated into release/v3.8.32 (restore-green: OpenAI gpt-4.1-mini/nano + o3-mini/o4-mini pricing rows)
…envelope (port from 9router#1944) (#4431)

Integrated into release/v3.8.32
…iew-prs mine r2 (#4461)

Baseline reconcile for /review-prs mine r2 batch
Comment thread src/shared/utils/secureRandom.ts Fixed
diegosouzapw and others added 5 commits June 20, 2026 21:54
…ider validators (#4467)

#4023 added validateWebCookieProvider (generic /models session ping -> AUTH_007
SESSION_EXPIRED on 401/403) and dispatched ALL web-cookie providers to it at the TOP
of validateProviderApiKey, BEFORE the SPECIALTY_VALIDATORS table. That shadowed the
rich per-provider validators (validateGrokWebProvider with #3474 IP-reputation/
Cloudflare guidance, validateChatGptWebProvider cf-mitigated, claude/gemini/copilot/
qwen/t3-web), which became dead code and broke all 41 web-cookie assertions in
provider-validation-specialty.test.ts (latent on release/v3.8.32; surfaced by a
__RUN_ALL__ unit run).

Move the generic dispatch to AFTER SPECIALTY_VALIDATORS so it is a FALLBACK only for
web-cookie providers without a dedicated validator. Rich validators run first
(provider-validation-specialty 112/112 restored); the generic + its AUTH_007
capability are preserved (web-cookie-auth007 stays 5/5, it calls the function directly).

Rebaselines validation.ts file-size 4518->4522 (+4, justified). Note: the sibling
pricing half of this restore-green already landed via #4447; this PR carries only the
stranded web-cookie validator fix.
…nfinite-scroll load-more (#4269) (#4478)

* fix(dashboard): resume request-log auto-refresh by gating the ghost infinite-scroll load-more (#4269)

* chore(quality): rebaseline RequestLoggerV2.tsx file-size 1287->1316 (#4269 ghost-load-more fix)
@sonarqubecloud

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
0.0% Coverage on New Code (required ≥ 80%)
B Maintainability Rating on New Code (required ≥ A)
D Reliability Rating on New Code (required ≥ A)

See analysis details on SonarQube Cloud

Catch issues before they fail your Quality Gate with our IDE extension SonarQube for IDE

diegosouzapw and others added 16 commits June 21, 2026 07:49
Thanks @CleanDev-Fix! Rebased onto release/v3.8.32 and integrated. Kiro overage quotas now stay routable (verified the buildKiroUsageResult → isExhausted → auth filter chain). Ships in the upcoming release.
…age (#4472)

Thanks @adivekar-utexas! Rebased onto release/v3.8.32 (resolved the additive conflict with #3872's queueDepth). Per-combo stickyRoundRobinLimit override now ships.
…4473)

Thanks @adivekar-utexas! Rebuilt on release/v3.8.32 with just the command-code commits (the combo half landed via #4472). reasoning/thinking passthrough + body.model rewrite now ship.
Thanks @KooshaPari! Rebased onto release/v3.8.32. Recharts now lazy-loads (public analytics API preserved, source-guard test included).
…#4433)

Thanks @KooshaPari! Rebased onto release/v3.8.32. Opt-in memory/bifrost compose profiles + BIFROST_ENABLED killswitch land (fixed the doc image to maximhq + frontmatter on review). qdrant-wiring 17/17.
Thanks @janeza2! Rebased onto release/v3.8.32 and completed the end-to-end wiring on review: kimi-coding-apikey was missing from USAGE_SUPPORTED_PROVIDERS (dashboard 400) and USAGE_FETCHER_PROVIDERS (preflight) — added both + header-selection/wiring tests. Now fully routable.
…4479)

Thanks @TF0rd! Rebased onto release/v3.8.32. Added the missing regression test for the model-aware redacted_thinking branch (Rule #18): minimax-m3 (claude format) gets redacted_thinking, glm-5.1 (openai) stays plain. 3/3.
…oop block (#4459)

Thanks @artickc! Rebased onto release/v3.8.32. Added a regression guard for the terminal type-scan shortcut (the one change with correctness risk) — typed/event-line/OpenAI/[DONE] shapes all verified. CRC gate stays opt-in, busy_timeout is tuning, compact artifacts round-trip via JSON.parse. 4/4.
Thanks @Rahulsharma0810! Rebased onto release/v3.8.32. Opt-in MODELS_CATALOG_PREFIX_MODE (dual default = no change) ships; restored 4 incidental comments + reconciled catalog.ts file-size baseline (1478->1486) on review. 48/48.
Thanks @adivekar-utexas! Rebuilt on release/v3.8.32 with just the UI commits (combo+command-code landed via #4472/#4473). Extracted a pure targetFormatBadgeI18nKey helper + added its unit test (Rule #18 gap) and reconciled the file-size baseline. targetFormat selector now ships.
@diegosouzapw
diegosouzapw merged commit bfaf459 into main Jun 21, 2026
7 checks passed
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
Release v3.8.32 — see CHANGELOG.md [3.8.32] for the full list. Merged via --admin over documented non-blocking checks: CodeQL alerts ratchet (diegosouzapw#665 fixed by diegosouzapw#4457/diegosouzapw#4462, auto-closes on main rescan), Integration Tests (env-flaky batch-upstream), SonarCloud/SonarQube (advisory new-code).
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
Release v3.8.32 — see CHANGELOG.md [3.8.32] for the full list. Merged via --admin over documented non-blocking checks: CodeQL alerts ratchet (diegosouzapw#665 fixed by diegosouzapw#4457/diegosouzapw#4462, auto-closes on main rescan), Integration Tests (env-flaky batch-upstream), SonarCloud/SonarQube (advisory new-code).
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
Release v3.8.32 — see CHANGELOG.md [3.8.32] for the full list. Merged via --admin over documented non-blocking checks: CodeQL alerts ratchet (diegosouzapw#665 fixed by diegosouzapw#4457/diegosouzapw#4462, auto-closes on main rescan), Integration Tests (env-flaky batch-upstream), SonarCloud/SonarQube (advisory new-code).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants