Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
ab8d9ed
feat(resilience): operator-configurable global credential health chec…
adivekar-utexas Aug 29, 2026
f89b48b
fix(i18n): add credential health check strings to pt-BR and vi locales
adivekar-utexas Aug 30, 2026
6f953f1
fix(i18n): fill missing locale strings and fix ICU plurals (vi, pt-BR)
adivekar-utexas Aug 30, 2026
4d20d37
fix(sse): keep cache-write tokens in OpenAI-shaped usage (#11814)
TheDemonTuan Aug 30, 2026
5d07bf3
feat(catalog): add feature flag to disable thinking level variants in…
b3nw Aug 30, 2026
fe8ef4f
fix(providers): use v1beta1 Model Garden publisher list for Vertex An…
fabioluissilva Aug 30, 2026
476b20b
fix(providers): cloudflare-ai flattens message content unconditionall…
davidlinfr Aug 30, 2026
51e4930
fix(build): prune non-production trees in NFT trace excludes and tsco…
Chewji9875 Aug 30, 2026
ff4ac6c
fix(provider/nous): inject required user tag into inference requests …
Karan825 Aug 30, 2026
6a41a78
fix(catalog): derive vision/modalities for built-in auto combos from …
Prajeeth-12 Aug 30, 2026
09428da
fix(dashboard): prevent provider icons collapsing to zero size (#12054)
ponkcore Aug 30, 2026
82f09f4
fix(api): align combo body and legacy key access (#12070)
marcelokarval Aug 30, 2026
54a1114
fix(sse): keep unavailable forced connections scoped (#12080)
keeltrace Aug 30, 2026
6096ea5
test(ui): correct inactive auto-fetch expectation (#12098)
RaviTharuma Aug 30, 2026
2da9ade
fix(sse): honor CLIProxyAPI environment API key (#12099)
RaviTharuma Aug 30, 2026
14dc6e8
feat(providers): add Perplexity Agent API provider (#12103)
AIB1TAL0S Aug 30, 2026
039a425
fix(oauth): bind Google refresh to the client that issued the token (…
HouMinXi Aug 30, 2026
55f6b98
fix(executors): DuckDuckGo ERR_BN_LIMIT without blind retry + proxy p…
oyi77 Aug 30, 2026
3d15294
fix(leases): project status lease row to lease columns so joined conn…
geek007git Aug 30, 2026
00bc397
fix(plugins): do not kill the plugin process when a fire-and-forget h…
geek007git Aug 30, 2026
1b2c6f4
fix(guardrails): restore injection-guard logging on middleware-only r…
geek007git Aug 30, 2026
d812585
fix(plugins): refresh stored manifest from disk on activate so new ho…
geek007git Aug 30, 2026
908c1b8
fix(codex): preserve existing provider state when bulk-import upserts…
geek007git Aug 30, 2026
a2c5d8a
feat(quota): use official OpenCode Go usage API (#12124)
ddarkr Aug 30, 2026
838fc00
fix(resilience): decouple rate-limit execution expiration from queue-…
alvinveroy Aug 30, 2026
9b9ea88
fix(migrations): add renamed migration compatibility for 056/073/077/…
oyi77 Aug 30, 2026
d13c6cb
fix(sse): emit native web_search_call for Responses web_search fallba…
watchingdogs Aug 30, 2026
e93c5e7
fix(diagnostics): keep the call-log error when the size limit strips …
ntdat812 Aug 30, 2026
6f914b7
feat(zai): add GLM-5.3-Flash Coding Plan support (#11801)
Neuron-Mr-White Aug 30, 2026
9ad41c9
board #12043
diegosouzapw Aug 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 11 additions & 13 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -673,21 +673,11 @@ NEXT_PUBLIC_CLOUD_URL=
# open-sse/services/usage.ts.
#OMNIROUTE_CROF_USAGE_URL=https://crof.ai/usage_api/
#OMNIROUTE_CODEWHISPERER_BASE_URL=https://codewhisperer.us-east-1.amazonaws.com
#OMNIROUTE_OPENCODE_QUOTA_URL=https://opencode.ai/zen/go/v1/quota
# OpenCode Go has no public quota API — this has no default and stays
# unset unless you explicitly opt in to a self-hosted/mirrored endpoint:
#OMNIROUTE_OPENCODE_GO_QUOTA_URL=
#OMNIROUTE_OPENCODE_GO_DASHBOARD_URL=https://opencode.ai/workspace
# Official OpenCode Go usage endpoint, authenticated with the connection API key.
# Override only for relays or test fixtures.
#OMNIROUTE_OPENCODE_QUOTA_URL=https://opencode.ai/zen/go/v1/usage
#OMNIROUTE_OLLAMA_CLOUD_USAGE_URL=https://ollama.com/settings

# OpenCode Go dashboard quota scraping. Prefer configuring these per connection
# in Dashboard → Providers → OpenCode Go. Env vars are useful for headless
# deployments or shared server defaults. The cookie is sensitive.
#OPENCODE_GO_WORKSPACE_ID=wrk_...
#OMNIROUTE_OPENCODE_GO_WORKSPACE_ID=wrk_...
#OPENCODE_GO_AUTH_COOKIE=auth=...
#OMNIROUTE_OPENCODE_GO_AUTH_COOKIE=auth=...

# OpenCode Go/Zen VPS egress (#5997): on a datacenter VPS, Cloudflare in front of
# opencode.ai/zen/go 403s chat requests that lack OpenCode CLI identity headers.
# When your clients don't already send them, set this to synthesize the CLI headers
Expand Down Expand Up @@ -2042,6 +2032,8 @@ APP_LOG_TO_FILE=true
# CLIPROXYAPI_HOST=127.0.0.1
# CLIPROXYAPI_PORT=5544
# CLIPROXYAPI_CONFIG_DIR=~/.cli-proxy-api
# Data-plane key fallback; the cliproxyapi_api_key setting takes precedence.
# CLIPROXYAPI_API_KEY=
# Management key for an externally managed instance. Embedded instances use
# OmniRoute's encrypted service key.
# CLIPROXYAPI_MANAGEMENT_KEY=
Expand Down Expand Up @@ -2145,6 +2137,12 @@ APP_LOG_TO_FILE=true
# Used by: open-sse/services/rateLimitManager.ts
# RATE_LIMIT_MAX_WAIT_MS=15000

# Limiter-managed execution backstop (Bottleneck `expiration`): bounds a job's
# post-dispatch execution, never queue wait. Must stay ABOVE upstream
# fetch-start timeouts on non-incremental gateways. Default: 600000 (10 min)
# Used by: open-sse/services/rateLimitManager.ts
# RATE_LIMIT_EXECUTION_MAX_WAIT_MS=600000

# Rate limit queue admission cap: reject with 429 queue_full once this many requests
# are already queued (0 = disabled/unbounded, the default). Used by: open-sse/services/rateLimitManager.ts
# RATE_LIMIT_MAX_QUEUE_DEPTH=0
Expand Down
8 changes: 8 additions & 0 deletions .github/workflows/electron-release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,9 @@ jobs:
- uses: actions/checkout@v7
with:
persist-credentials: false
# workflow_dispatch: build the tag being (re)built, not the dispatching branch. On a
# tag push this resolves to the same commit.
ref: ${{ needs.validate.outputs.version }}
- name: Setup Node
uses: actions/setup-node@v7
with:
Expand Down Expand Up @@ -170,6 +173,9 @@ jobs:
- uses: actions/checkout@v7
with:
persist-credentials: false
# workflow_dispatch: build the tag being (re)built, not the dispatching branch. On a
# tag push this resolves to the same commit.
ref: ${{ needs.validate.outputs.version }}
- name: Setup Node
uses: actions/setup-node@v7
with:
Expand Down Expand Up @@ -356,6 +362,8 @@ jobs:
with:
persist-credentials: false
fetch-depth: 0
# Source archives + SBOM come from the tag being released, not the dispatching branch.
ref: ${{ needs.validate.outputs.version }}

# `merge-multiple` is deliberately OFF. It resolves same-name collisions by ARRIVAL
# ORDER, and the two macOS jobs each emit their own `latest-mac.yml` listing only their
Expand Down
15 changes: 11 additions & 4 deletions bin/cli/api-commands/combos.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ export function register_combos(parent) {
});
tag.command("post-api-combos")
.description("Create routing combo")
.option("--body <jsonOrPath>", "JSON body or @path/to/file.json")
.requiredOption("--body <jsonOrPath>", "JSON body or @path/to/file.json")
.action(async (opts, cmd) => {
const gOpts = cmd.optsWithGlobals();
let url = "/api/combos";
Expand Down Expand Up @@ -44,7 +44,7 @@ export function register_combos(parent) {
tag.command("put-api-combos-id-")
.description("Update combo")
.requiredOption("--id <id>", "")
.option("--body <jsonOrPath>", "JSON body or @path/to/file.json")
.requiredOption("--body <jsonOrPath>", "JSON body or @path/to/file.json")
.action(async (opts, cmd) => {
const gOpts = cmd.optsWithGlobals();
let url = "/api/combos/{id}";
Expand All @@ -62,7 +62,7 @@ export function register_combos(parent) {
tag.command("patch-api-combos-id-")
.description("Update combo")
.requiredOption("--id <id>", "")
.option("--body <jsonOrPath>", "JSON body or @path/to/file.json")
.requiredOption("--body <jsonOrPath>", "JSON body or @path/to/file.json")
.action(async (opts, cmd) => {
const gOpts = cmd.optsWithGlobals();
let url = "/api/combos/{id}";
Expand Down Expand Up @@ -99,10 +99,17 @@ export function register_combos(parent) {
});
tag.command("post-api-combos-test")
.description("Test a combo configuration")
.requiredOption("--body <jsonOrPath>", "JSON body or @path/to/file.json")
.action(async (opts, cmd) => {
const gOpts = cmd.optsWithGlobals();
let url = "/api/combos/test";
const res = await apiFetch(url, { method: "POST", baseUrl: gOpts.baseUrl, apiKey: gOpts.apiKey });
let body;
if (opts.body) {
body = opts.body.startsWith("@")
? JSON.parse(readFileSync(opts.body.slice(1), "utf8"))
: JSON.parse(opts.body);
}
const res = await apiFetch(url, { method: "POST", body, baseUrl: gOpts.baseUrl, apiKey: gOpts.apiKey });
const data = res.ok ? await res.json() : await res.text();
emit(data, gOpts);
});
Expand Down
1 change: 1 addition & 0 deletions changelog.d/features/11801-zai-glm-53-flash.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **feat(zai):** add GLM-5.3-Flash Coding Plan support (1M context, 128K output, vision, `low|high|max` reasoning) and route `zai` GLM-5.3-family API-key traffic through the OpenAI-compatible Coding Plan endpoint with native thinking defaults ([#11801](https://github.com/diegosouzapw/OmniRoute/pull/11801)) — thanks @Neuron-Mr-White
1 change: 1 addition & 0 deletions changelog.d/features/disable-thinking-level-variants.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **feat(catalog):** add `OMNIROUTE_DISABLE_THINKING_LEVEL_VARIANTS` feature flag to optionally filter out thinking level variants from model catalog ([#PR_NUMBER](https://github.com/diegosouzapw/OmniRoute/pull/PR_NUMBER))
1 change: 1 addition & 0 deletions changelog.d/features/perplexity-agent-provider.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **feat(providers):** add a Perplexity Agent API provider (`perplexity-agent` / `pplx-agent`) for Perplexity `/v1/responses`, including the documented Anthropic, OpenAI, Google, xAI, DeepSeek, Z.AI, Moonshot/Kimi, NVIDIA, and Perplexity model IDs plus Anthropic-model `max_output_tokens` compatibility.
1 change: 1 addition & 0 deletions changelog.d/fixes/0000-provider-icon-zero-size.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(dashboard):** Keep local and theme-aware provider SVG icons at a definite layout size so Chromium does not collapse them to 0×0 after the v3.8.50 image-rendering change ([#12054](https://github.com/diegosouzapw/OmniRoute/pull/12054)) — thanks @ponkcore
1 change: 1 addition & 0 deletions changelog.d/fixes/11861-nous-tags-user-injection.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(provider/nous):** inject required user= tag into Nous Research inference requests to resolve upstream 400 "missing tags" error ([#11861](https://github.com/diegosouzapw/OmniRoute/issues/11861)) — thanks @Karan825
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(providers):** Vertex AI Anthropic partner-model discovery now calls the Model Garden `v1beta1` publisher list (`/v1beta1/publishers/anthropic/models`, global) and parses its `publisherModels` envelope, so Claude models auto-synced from Vertex populate the active live catalog and route at request time instead of returning `Model '<id>' is not available in the active live catalog` ([#11991](https://github.com/diegosouzapw/OmniRoute/issues/11991)) — thanks @fabioluissilva
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(providers):** `cloudflare-ai` no longer refuses image content parts for every Workers AI model ([#12002](https://github.com/diegosouzapw/OmniRoute/pull/12002)) — the plain-string `content` requirement behind #2539 is carried by the _model_ schema, not by the `/ai/v1/chat/completions` endpoint (measured: an all-text part array returns 200 on `@cf/mistralai/mistral-small-3.1-24b-instruct`, `@cf/meta/llama-4-scout-17b-16e-instruct` and `@cf/meta/llama-3.3-70b-instruct-fp8-fast`, and 400 on the text-only `@cf/qwen/qwen2.5-coder-32b-instruct`). `transformRequest()` flattened every array and threw on the first non-text part (#6390), so image input was refused for vision-capable Cloudflare models that accept it. All-text arrays are still flattened — the one shape every model accepts — while an array carrying a non-text part is passed through untouched, so the attachment is still never silently dropped. Regression guards: `tests/unit/cloudflare-ai-image-parts-6390.test.ts`.
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(resilience):** decouple the limiter-managed execution backstop from the queue-wait budget — new `requestQueue.executionMaxWaitMs` (env `RATE_LIMIT_EXECUTION_MAX_WAIT_MS`, default 600000 = 10 min) now feeds Bottleneck's post-dispatch `expiration`, while `requestQueue.maxWaitMs` keeps its documented queue-wait semantics. Previously the queue-wait budget doubled as the execution expiration, so legitimate long-running calls on non-incremental gateways (whole generation buffered before the first upstream byte, e.g. Console Go / Command Code tiers serving GLM models) were killed mid-flight at the queue budget with a false 504 `RATE_LIMIT_EXECUTION_TIMEOUT` — the local limiter undercut the provider-aware upstream fetch-start timeouts. The surfaced 504 message now names `requestQueue.executionMaxWaitMs`; the error keeps the #4165 guarantees (disclaims an upstream timeout, preserves the Bottleneck error as `cause`, branded code + trusted provenance, classified request-scoped so combo falls back). A real queue-wait bound (the `Promise.race` around `limiter.schedule()` sketched in #9533) remains future work. (#12025)
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **Call logs:** keep the `error` field when an artifact exceeds the storage cap, instead of replacing it with the omission marker. The error is the only field that says *why* a request failed and is typically ~90 bytes next to the multi-hundred-KB bodies that trip the cap, so dropping it left a size-limited row undiagnosable — a provider outage, a local timeout and an upstream 400 all rendered identically. It is now preserved at every fallback stage, truncated to 4KB if it is itself large ([#12026](https://github.com/diegosouzapw/OmniRoute/issues/12026)).
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(diagnostics):** preserve the error field (truncated to 4KB with a `[truncated: …]` suffix) in every call-log artifact size-limit fallback stage. Previously the minimal fallback replaced the error with `[omitted: call log artifact size limit exceeded]`, so an oversized artifact row showed nothing about WHY the request failed — e.g. 91 of 847 opencode-go 504 rows on one production instance were undiagnosable from the dashboard. Oversized request/response bodies are still omitted exactly as before; the error cap is independent of the payload sizes that tripped the fallback. (#12026)
1 change: 1 addition & 0 deletions changelog.d/fixes/12031-web-search-call-emission.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(sse):** OpenAI Responses clients that declare the native `web_search` tool now receive a spec-shaped `web_search_call` output item with `action.sources` alongside the preserved function-call round-trip, so search results executed through OmniRoute's own search backend are consumable by standard Responses clients (Codex, pi-web-access, …).
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(executors):** handle DuckDuckGo ERR_BN_LIMIT (418) without retrying — when the upstream returns `418 ERR_BN_LIMIT` (rate-limit/ban), the executor now returns the error immediately instead of burning another VQD acquisition that would only count against the IP limit. The retry logic for `418 ERR_CHALLENGE` (unsolved challenge) remains unchanged. ([#11598](https://github.com/diegosouzapw/OmniRoute/pull/11598))
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(api):** Generated API CLI commands now enforce required OpenAPI request bodies; Combo test commands forward the required `comboName` body, while API keys created by older writers after migration 149 preserve legacy allow-all Combo access without widening explicit empty allowlists — thanks @marcelokarval
15 changes: 0 additions & 15 deletions config/quality/eslint-suppressions.json
Original file line number Diff line number Diff line change
Expand Up @@ -29,11 +29,6 @@
"count": 1
}
},
"open-sse/config/providers/registry/zai/index.ts": {
"@typescript-eslint/no-unused-vars": {
"count": 1
}
},
"open-sse/config/providers/shared.ts": {
"@typescript-eslint/no-unused-vars": {
"count": 1
Expand Down Expand Up @@ -618,11 +613,6 @@
"count": 1
}
},
"open-sse/services/opencodeQuotaFetcher.ts": {
"no-restricted-syntax": {
"count": 1
}
},
"open-sse/services/providerCostData.ts": {
"@typescript-eslint/no-unused-vars": {
"count": 1
Expand Down Expand Up @@ -4477,11 +4467,6 @@
"count": 6
}
},
"tests/unit/opencode-quota-fetcher.test.ts": {
"@typescript-eslint/no-explicit-any": {
"count": 1
}
},
"tests/unit/openrouter-vision-sync-4264.test.ts": {
"@typescript-eslint/no-explicit-any": {
"count": 4
Expand Down
8 changes: 1 addition & 7 deletions docs/i18n/zh-CN/docs/reference/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -269,13 +269,7 @@ OmniRoute 提供两层防护:请求侧的注入扫描和响应侧的 PII 脱
| `OMNIROUTE_KIE_CALLBACK_URL` | _(未设置)_ | `open-sse/utils/kieTask.ts` | `KIE_CALLBACK_URL` 的替代写法。主变量未设置时的回退。 |
| `OMNIROUTE_PUBLIC_URL` | _(未设置)_ | `open-sse/utils/kieTask.ts` | 用于组合异步回调 URL 的公共源。kie.ai 回调的最低优先级回退;也用作其他中继的通用公共 URL。 |
| `OMNIROUTE_CROF_USAGE_URL` | `https://crof.ai/usage_api/` | `open-sse/services/usage.ts` | Usage 页面使用的 CrofAI 配额查询端点。可覆盖为中继/测试固定件。 |
| `OMNIROUTE_OPENCODE_QUOTA_URL` | `https://opencode.ai/zen/go/v1/quota` | `open-sse/services/opencodeQuotaFetcher.ts` | Usage 页面使用的 OpenCode (zen/go) 配额查询端点。可覆盖为中继/测试固定件。 |
| `OMNIROUTE_OPENCODE_GO_QUOTA_URL` | _(未设置)_ | `open-sse/services/opencodeOllamaUsage.ts` | Usage 页面使用的 OpenCode Go 配额查询端点。OpenCode Go 没有公开的配额 API,因此没有默认值;除非运维人员显式设置该变量选择接入自建/镜像端点,否则不会发起网络请求。 |
| `OMNIROUTE_OPENCODE_GO_DASHBOARD_URL` | `https://opencode.ai/workspace` | `open-sse/services/usage.ts` | 配置了 workspace ID 和 auth Cookie 时用于配额抓取的 OpenCode Go Dashboard 基础 URL。可覆盖为中继/测试固定件。 |
| `OPENCODE_GO_WORKSPACE_ID` | _(未设置)_ | `open-sse/services/usage.ts` | 用于 Dashboard 配额抓取的 OpenCode Go workspace ID。配置多个账户时,推荐使用每个连接的 Dashboard 字段。 |
| `OMNIROUTE_OPENCODE_GO_WORKSPACE_ID` | _(未设置)_ | `open-sse/services/usage.ts` | OpenCode Go workspace ID 环境变量的备选名,在较短的别名之前使用。配置多个账户时,推荐使用每个连接的 Dashboard 字段。 |
| `OPENCODE_GO_AUTH_COOKIE` | _(未设置)_ | `open-sse/services/usage.ts` | 用于 Dashboard 配额抓取的 OpenCode Go `auth` Cookie。敏感信息;配置多个账户时,推荐使用每个连接的 Dashboard 字段。 |
| `OMNIROUTE_OPENCODE_GO_AUTH_COOKIE` | _(未设置)_ | `open-sse/services/usage.ts` | OpenCode Go `auth` Cookie 环境变量的备选名,在较短的别名之前使用。敏感信息;配置多个账户时,推荐使用每个连接的 Dashboard 字段。 |
| `OMNIROUTE_OPENCODE_QUOTA_URL` | `https://opencode.ai/zen/go/v1/usage` | `open-sse/services/opencodeQuotaFetcher.ts` | Usage 页面使用的 OpenCode Go 官方 API key 认证用量端点。可覆盖为中继/测试固定件。 |
| `OMNIROUTE_OLLAMA_CLOUD_USAGE_URL` | `https://ollama.com/settings` | `open-sse/services/usage.ts` | 用于配额抓取的 Ollama Cloud settings URL。可覆盖为中继/测试固定件。 |
| `OLLAMA_USAGE_COOKIE` | _(未设置)_ | `open-sse/services/usage.ts` | 用于设置页面配额抓取的 Ollama Cloud `__Secure-session` Cookie。敏感信息;配置多个账户时,推荐使用每个连接的 Dashboard 字段。 |
| `OLLAMA_CLOUD_USAGE_COOKIE` | _(未设置)_ | `open-sse/services/usage.ts` | Ollama Cloud `__Secure-session` Cookie 环境变量的备选名。敏感信息;配置多个账户时,推荐使用每个连接的 Dashboard 字段。 |
Expand Down
15 changes: 15 additions & 0 deletions docs/openapi.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2350,9 +2350,24 @@ paths:
post:
tags: [Combos]
summary: Test a combo configuration
requestBody:
required: true
content:
application/json:
schema:
type: object
required: [comboName]
properties:
comboName:
type: string
minLength: 1
responses:
"200":
description: Test result
"400":
description: Missing or invalid combo name
"404":
description: Combo not found

/api/settings:
get:
Expand Down
Loading
Loading