Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
38ba856
fix(resilience): route chat by per-connection synced model inventory …
pacocartones Aug 22, 2026
f387575
fix(memory): treat TokenRouter as system-must-be-first (live HTTP 400…
ggdayup Aug 23, 2026
b384455
fix(api): refuse creating a routing combo without any model (#11162)
maxmad64bis Aug 23, 2026
1058426
fix(executors): rotate on upstream 400 empty-body rejections (opencod…
maxmad64bis Aug 23, 2026
60f25ef
fix(sse): merge purify_history compression notice into the leading sy…
ggdayup Aug 23, 2026
80b8d2a
fix(sse): resume stream recovery after a clean stop with reasoning-on…
maxmad64bis Aug 23, 2026
62ab93d
fix(sse): reject a low-overlap stream-recovery continuation instead o…
maxmad64bis Aug 23, 2026
eb9fa33
feat(api): structured ?format=json for the self-service usage endpoin…
diegosouzapw Aug 23, 2026
3ef54fc
feat(api): return every connection's snapshot under providers[] in om…
diegosouzapw Aug 23, 2026
eb57973
feat(dashboard): add CheaperInference sponsor banner and route banner…
diegosouzapw Aug 23, 2026
d888f1a
fix(combo): accept SSE comment lines (e.g. OpenRouter keep-alives) in…
asorourx Aug 23, 2026
e73ab00
fix(codex): make remote compaction V2 complete reliably (#11041)
jackjinke Aug 23, 2026
2300171
fix(resilience): drain heavyweight SSE on SIGTERM (#11020)
RaviTharuma Aug 23, 2026
e52d2db
fix(compression): CCR must not strand prompts for callers without the…
HouMinXi Aug 23, 2026
c018bb4
fix(catalog): declare GLM reasoning effort tiers (#10963)
xz-dev Aug 23, 2026
9689dce
fix(reasoning): preserve mixed plaintext and drop incompatible state …
jackjinke Aug 23, 2026
2dd2033
feat(providers): add Logfare as a free OpenAI-compatible provider (#1…
jonlwheat2-gif Aug 23, 2026
79c5bdf
fix(release): repair v3.8.50 base-red tail after latest root lift (#1…
backryun Aug 23, 2026
f131b64
fix(security): add test coverage for Tier 1 local-only route guard pr…
rqzbeh Aug 23, 2026
b25b2ea
refactor(dashboard): format custom provider quota keys into title-cas…
rqzbeh Aug 23, 2026
47147e0
fix(auth): replace router.push with window.location navigation after …
rqzbeh Aug 23, 2026
8fa3e31
fix(sse): unpin static Antigravity sessionId and add DNS retry classi…
rqzbeh Aug 23, 2026
c92bd40
perf(proxy): implement non-blocking async proxy log batching and perf…
rqzbeh Aug 23, 2026
376b49d
fix(ratelimit): keep operator minTime floor when relaxing on headroom…
pacocartones Aug 23, 2026
b6fc559
fix(providers): g4f.space sub-providers no longer advertise a free ti…
pacocartones Aug 23, 2026
8c73386
fix(routing): forward lkgpEnabled into RoutingContext so the LKGP tog…
pacocartones Aug 23, 2026
a966c75
fix(routing): keep keyless custom-compatible connections in the auto/…
pacocartones Aug 23, 2026
8d6870f
fix(analytics): classify opencode-go as a flat-rate subscription (#11…
pacocartones Aug 23, 2026
100b60a
board #11186
hartmark Aug 23, 2026
1722c83
chore(quality): freeze auth.ts at 3337 with annotation (#11186)
hartmark Aug 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ Repository map and Reference Documentation sections below.

## Project at a Glance

**OmniRoute** — unified AI proxy/router. One endpoint, 348 LLM providers, auto-fallback.
**OmniRoute** — unified AI proxy/router. One endpoint, 349 LLM providers, auto-fallback.

| Layer | Location | Purpose |
| ------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
Expand Down
447 changes: 447 additions & 0 deletions PROVIDER_REFERENCE.md

Large diffs are not rendered by default.

25 changes: 13 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@

# 🚀 OmniRoute — The Free AI Gateway

<img src="./docs/diagrams/readme-hero.svg" width="100%" alt="OmniRoute — Never stop coding. Every AI tool → 348 providers — 90+ free — through one endpoint. Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini with auto-fallback. RTK + Caveman stacked compression saves 15–95% tokens (~89% avg) — never hit limits. 348 AI providers · 90+ free tiers · ~1.51B free tokens/mo · 19 routing strategies · $0 to start."/>
<img src="./docs/diagrams/readme-hero.svg" width="100%" alt="OmniRoute — Never stop coding. Every AI tool → 349 providers — 90+ free — through one endpoint. Claude Code, Codex, Cursor, Cline, Copilot & Antigravity into FREE Claude / GPT / Gemini with auto-fallback. RTK + Caveman stacked compression saves 15–95% tokens (~89% avg) — never hit limits. 349 AI providers · 90+ free tiers · ~1.51B free tokens/mo · 19 routing strategies · $0 to start."/>

</div>

Expand Down Expand Up @@ -101,7 +101,7 @@
<tr>
<td align="right"><b>⚙️ Features</b></td>
<td align="center"><a href="#-combos--the-flagship">🎯 Combos</a></td>
<td align="center"><a href="#-348-ai-providers--90-free">🌐 Providers</a></td>
<td align="center"><a href="#-349-ai-providers--90-free">🌐 Providers</a></td>
<td align="center"><a href="#-full-cli--a2a--mcp">🔌 CLI &amp; MCP</a></td>
</tr>
<tr>
Expand Down Expand Up @@ -210,7 +210,7 @@ curl http://localhost:20128/v1/chat/completions \

</div>

<img src="./docs/diagrams/promise-pillars.svg" width="100%" alt="The Promise — One endpoint. 348 providers. Never stop building — OmniRoute picks the cheapest one that works. Six pillars: Never hit limits (auto-fallback across 348 providers in milliseconds, zero downtime) · Save up to 95% tokens (RTK + Caveman stacked compression cuts 15–95%, ~89% avg on tool-heavy sessions) · $0 to start (90+ free tiers, 56 free forever — no card needed) · Every tool works (33 coding agents through one config) · One endpoint (OpenAI ↔ Claude ↔ Gemini ↔ Responses API at /v1) · Production-grade (circuit breakers, TLS stealth, MCP 110 tools, A2A, memory, guardrails, evals — 25,000+ tests)."/>
<img src="./docs/diagrams/promise-pillars.svg" width="100%" alt="The Promise — One endpoint. 349 providers. Never stop building — OmniRoute picks the cheapest one that works. Six pillars: Never hit limits (auto-fallback across 349 providers in milliseconds, zero downtime) · Save up to 95% tokens (RTK + Caveman stacked compression cuts 15–95%, ~89% avg on tool-heavy sessions) · $0 to start (90+ free tiers, 56 free forever — no card needed) · Every tool works (33 coding agents through one config) · One endpoint (OpenAI ↔ Claude ↔ Gemini ↔ Responses API at /v1) · Production-grade (circuit breakers, TLS stealth, MCP 110 tools, A2A, memory, guardrails, evals — 25,000+ tests)."/>

<br/>
<br/>
Expand Down Expand Up @@ -461,7 +461,7 @@ All **19** strategies — mix & match per combo step:

</div>

<img src="./docs/diagrams/comparison-table.svg" width="100%" alt="What sets OmniRoute apart — comparison table vs 9router, OpenRouter, CLIProxyAPI and LiteLLM across 13 capabilities. OmniRoute: 348 providers, 90+ free providers built-in, 19 routing strategies, 12-engine token compression, built-in MCP server with 110 tools, A2A agent protocol, persistent memory, guardrails, cloud agents, TLS fingerprint stealth, Desktop/Termux/PWA, 43 i18n UI locales, 100% MIT self-hosted. OmniRoute is the only one with the full set; competitors show a mix of checks, partials and crosses. Verified from each project&apos;s docs."/>
<img src="./docs/diagrams/comparison-table.svg" width="100%" alt="What sets OmniRoute apart — comparison table vs 9router, OpenRouter, CLIProxyAPI and LiteLLM across 13 capabilities. OmniRoute: 349 providers, 90+ free providers built-in, 19 routing strategies, 12-engine token compression, built-in MCP server with 110 tools, A2A agent protocol, persistent memory, guardrails, cloud agents, TLS fingerprint stealth, Desktop/Termux/PWA, 43 i18n UI locales, 100% MIT self-hosted. OmniRoute is the only one with the full set; competitors show a mix of checks, partials and crosses. Verified from each project&apos;s docs."/>

<sub>📊 Full methodology &amp; per-feature detail vs 9router, OpenRouter, CLIProxyAPI &amp; LiteLLM → [`docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md`](docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md)</sub>

Expand Down Expand Up @@ -559,7 +559,7 @@ the current catalog at **[radar.omniroute.online/planos](https://radar.omniroute
- **🖼️ New endpoints** — `/v1/ocr` (Mistral OCR) and `/v1/audio/translations` (Whisper-style) round out the media surface. → [API Reference](docs/reference/API_REFERENCE.md)
- **🎨 Image / video / audio generation** — one API for media: xAI Grok Imagine & Novita AI video, ComfyUI, Freepik, Adobe Firefly, Microsoft Designer, Segmind, EdgeTTS. → [API Reference](docs/reference/API_REFERENCE.md)
- **🌍 Deployment & ops** — reverse-proxy `basePath`, browser-language auto-detect, per-key device tracking, root-less MITM trust, zh-TW localization. → [Environment](docs/reference/ENVIRONMENT.md)
- **🤝 More providers & agents** — Cursor Cloud Agent, Grok Build (xAI) with browser + OAuth login, Ollama first-class card, Claude Opus 5 & Sonnet 5, Kimi official partnership (Code/Web/Moonshot), Zed, Requesty, SenseNova, Yuanbao, Agnes AI… and a refreshed **348-provider catalog**. → [Providers](docs/reference/PROVIDER_REFERENCE.md)
- **🤝 More providers & agents** — Cursor Cloud Agent, Grok Build (xAI) with browser + OAuth login, Ollama first-class card, Claude Opus 5 & Sonnet 5, Kimi official partnership (Code/Web/Moonshot), Zed, Requesty, SenseNova, Yuanbao, Agnes AI… and a refreshed **349-provider catalog**. → [Providers](docs/reference/PROVIDER_REFERENCE.md)
- **📡 Routing transparency** — every response carries an `X-OmniRoute-Decision` header naming the strategy/provider/latency that served it, a new `cache-optimized` combo strategy + Auto-Combo `cacheAffinity` factor route repeat requests back to the connection holding the cached prefix, and a read-only `/v1/auto-combo/{channel}/candidates` endpoint exposes an `auto/*` channel's live candidate pool. → [Auto-Combo](docs/routing/AUTO-COMBO.md)
- **⚡ Local performance & infra** — one-click local Redis, Cloudflare Workers / Deno Deploy relay deployers, Bifrost & Mux as supervised embedded services. → [Embedded Services](docs/frameworks/EMBEDDED-SERVICES.md)

Expand Down Expand Up @@ -642,11 +642,11 @@ of your shell history. → [CLI Integrations](docs/guides/CLI-INTEGRATIONS.md)

<div align="center">

## 🌐 348 AI Providers — 90+ Free
## 🌐 349 AI Providers — 90+ Free

</div>

> The most complete catalog of any open-source router: **348 providers**, **90+ with a free tier**, **56 free forever**.
> The most complete catalog of any open-source router: **349 providers**, **90+ with a free tier**, **56 free forever**.

<div align="center">

Expand Down Expand Up @@ -990,11 +990,11 @@ docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \

`:latest` follows the highest **published** stable SemVer. It does not track git `main`. Pin `:X.Y.Z` for GitOps. See [Docker Release Channels](docs/guides/DOCKER_GUIDE.md#release-channels).The image pins **`OMNIROUTE_MEMORY_MB=1024`**. That is enough for the dashboard and a light chat. **Coding agents** (`POST /v1/responses` from Claude Code, Codex, Grok, …) need a much larger V8 heap or the process `FATAL ERROR`s at ~12 GiB under two overlapping long contexts. Size the container above the heap (native buffers sit outside V8):

| Workload | Heap (`-e OMNIROUTE_MEMORY_MB`) | Container (`--memory`) |
| --- | --- | --- |
| Dashboard / light chat | `1024` (image default) | ≥2 g |
| One coding agent | `8192` | ≥10 g |
| Two concurrent long `/v1/responses` | `10240`–`12288` | ≥12–16 g |
| Workload | Heap (`-e OMNIROUTE_MEMORY_MB`) | Container (`--memory`) |
| ----------------------------------- | ------------------------------- | ---------------------- |
| Dashboard / light chat | `1024` (image default) | ≥2 g |
| One coding agent | `8192` | ≥10 g |
| Two concurrent long `/v1/responses` | `10240`–`12288` | ≥12–16 g |

```bash
docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \
Expand All @@ -1003,6 +1003,7 @@ docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \
```

Full table: [Docker Guide — runtime RAM](docs/guides/DOCKER_GUIDE.md#runtime-ram-for-coding-agents).

> **Pre-release Docker channel:** `diegosouzapw/omniroute:next` and
> `diegosouzapw/omniroute:next-web` follow the current default `release/v*`
> branch. These mutable tags are intended only for testing unreleased fixes and
Expand Down
6 changes: 6 additions & 0 deletions bin/cli/commands/combo.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -307,6 +307,12 @@ export async function runComboCreateCommand(name, strategy = "priority", opts =
}

const models = Array.isArray(opts.models) ? opts.models : [];
if (!models.length) {
console.error(
"combo create requires at least one target. Pass --models <provider/model,...> and/or repeat --model <provider/model>."
);
return 1;
}

try {
return await withRuntime(async ({ kind, api, db }) => {
Expand Down
1 change: 1 addition & 0 deletions changelog.d/features/10987-logfare-free-provider.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **feat(providers):** add Logfare as a free OpenAI-compatible provider — dashboard card with a Free badge and request-logging disclosure (every prompt/completion is logged for research; opt out at logfare.ai/consent), live model discovery from `https://logfare.ai/v1/models` (20 models, 11 chat-capable: kimi-k3, deepseek-v4-pro, glm-5.2, gpt-5.6-luna, minimax-m3…), full chat/streaming through the existing OpenAI-compatible path, the real Logfare logo on the card, and a listing in the free-tiers guide. ([#10987](https://github.com/diegosouzapw/OmniRoute/pull/10987))
1 change: 1 addition & 0 deletions changelog.d/features/11190-usage-command-json.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **feat(api):** `/api/usage/om-usage` gains a structured form — `?format=json` returns the key's own usage as `ApiKeyUsageLimitStatus` + `UsageSnapshot` instead of `text/plain`. This is the surface a UI (the OmniCopilot panel) consumes to show a key holder their daily/weekly spend and quota reset. The route is self-service (the caller's own key, gated by `allowUsageCommand`), not the management surface; refusals come back as a discriminated `{ "allowed": false, "error": … }` so a UI can tell "not allowed" apart from "allowed but nothing cached yet". The endpoint was previously undocumented in `API_REFERENCE.md`; it now has a section ([#11190](https://github.com/diegosouzapw/OmniRoute/pull/11190))
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **feat(api):** `/api/usage/om-usage?format=json` now returns `providers[]` — every connection's quota snapshot, not just the single selected one — so a panel can render Codex / Claude / OpenCode side by side. The collector already gathered all of them; the single-pick `provider` field (kept) is a terminal presentation choice. Closes the per-connection gap from OmniCopilot #8 ([#11192](https://github.com/diegosouzapw/OmniRoute/pull/11192))
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(providers):** the five g4f.space sub-providers (Groq, Gemini, Pollinations, Ollama, NVIDIA) no longer advertise a free tier — a keyless `POST /v1/chat/completions` now returns `402 insufficient_credits` behind a proof-of-work "cake" wall (re-verified live 2026-08-22), so `hasFree` is `false` and the notes point at `g4f.dev/members.html`. The gateway still works with a member key, so its registry wiring and `authType: "optional"` are unchanged ([#10071](https://github.com/diegosouzapw/OmniRoute/issues/10071)) — thanks @chirag127
2 changes: 1 addition & 1 deletion changelog.d/fixes/10550-responses-reasoning-transport.md
Original file line number Diff line number Diff line change
@@ -1 +1 @@
- Preserve portable plaintext reasoning by default across streaming and non-streaming Chat Completions and Responses routes while keeping provider-bound opaque state target-compatible. Combos now drop incompatible continuation reasoning by default and can explicitly skip incompatible targets, while known providers no longer show redundant encrypted-reasoning controls. (#10550)
- Preserve portable plaintext reasoning by default across streaming and non-streaming Chat Completions and Responses routes while keeping provider-bound opaque state target-compatible. Direct requests drop incompatible continuation reasoning by default; combos can explicitly skip incompatible targets without mutating the request. Known providers no longer show redundant encrypted-reasoning controls. (#10550, #10959)
1 change: 1 addition & 0 deletions changelog.d/fixes/10949-mixed-reasoning-plaintext.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- Preserve explicit plaintext reasoning when a Responses reasoning item also carries opaque provider state (rare OpenCode Go `deepseek-v4-flash` responses). Mixed plaintext + opaque input is projected onto the target transport: plaintext targets keep portable text, opaque targets keep provider state. Opaque-only reasoning is dropped when the selected target cannot replay it, allowing cross-model conversations to continue. (#10949, #10959)
1 change: 1 addition & 0 deletions changelog.d/fixes/11015-shutdown-track-sse.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(resilience):** count heavyweight `/v1` admission leases in the SIGTERM drain and send `Retry-After` on shutdown 503s so Recreate no longer looks like an empty 502 ([#11015](https://github.com/diegosouzapw/OmniRoute/issues/11015)) — thanks @RaviTharuma
1 change: 1 addition & 0 deletions changelog.d/fixes/11089-chat-routing-synced-inventory.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(resilience):** filter chat connection selection by each connection's *synced* model inventory on multi-host self-hosted providers (`ollama-local`, `lm-studio`, `vllm`, …), so a request for a model only one host advertises is pinned to that host instead of failing over onto a host that never had it ([#11089](https://github.com/diegosouzapw/OmniRoute/issues/11089))
1 change: 1 addition & 0 deletions changelog.d/fixes/11149-opencode-go-flat-rate.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(analytics):** `opencode-go` is now classified as a flat-rate subscription, so cost analytics shows $0 for it instead of billing every call at the underlying model’s metered rate — it resells GLM, Kimi, Grok, DeepSeek, MiniMax, Qwen and GPT-5.x under one flat monthly fee, which made the overstatement large rather than marginal ([#11149](https://github.com/diegosouzapw/OmniRoute/pull/11149)) — thanks @electrumguy
1 change: 1 addition & 0 deletions changelog.d/fixes/11162-combo-create-requires-model.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **Combo create:** creating a routing combo without any model is now refused (`400`) — the CLI requires `--models`/`--model` on `combo create`, matching the dashboard which already rejected empty combos.
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(routing):** a custom `openai-compatible-*` / `anthropic-compatible-*` connection pointing at a keyless self-hosted backend (llama.cpp, Ollama, vLLM started without an API key) now stays in the `auto/*` candidate pool instead of being silently dropped by the credential gate — for those IDs "no credential" is the normal configuration, not an unconfigured connection ([#11180](https://github.com/diegosouzapw/OmniRoute/pull/11180)) — thanks @marcs7
1 change: 1 addition & 0 deletions changelog.d/fixes/11181-lkgp-enabled-context.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(routing):** the Routing tab's "last known good provider" toggle now actually takes effect — `lkgpEnabled` was persisted and the `lkgp` strategy guarded on it, but the setting was never forwarded into the `RoutingContext` built in `resolveAutoStrategyOrder()`, so `context.lkgpEnabled` was always `undefined` and the off-switch was unreachable ([#11181](https://github.com/diegosouzapw/OmniRoute/issues/11181))
1 change: 1 addition & 0 deletions changelog.d/fixes/9763-ratelimit-mintime-floor.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(ratelimit):** respect operator `minTimeBetweenRequestsMs` floor when relaxing the limiter on headroom — the adaptive rate-limit learning no longer silently erases a configured minimum gap between requests when the upstream reports plenty of remaining capacity ([#9763](https://github.com/diegosouzapw/OmniRoute/issues/9763)).
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(executors):** OpencodeExecutor rotates (or retries once on a single-account direct path) on upstream 400 empty-body rejections — malformed completion envelopes with no error field were propagated as success and killed client sessions. Bounded +1 attempt per request; body reads are conditioned on status 400 so successful/streaming responses are never buffered. 400s carrying an error field keep propagating immediately.
12 changes: 1 addition & 11 deletions config/quality/eslint-suppressions.json
Original file line number Diff line number Diff line change
Expand Up @@ -853,11 +853,6 @@
"count": 1
}
},
"src/app/api/usage/call-logs/route.ts": {
"no-restricted-imports": {
"count": 1
}
},
"src/app/api/usage/quota/route.ts": {
"no-restricted-imports": {
"count": 1
Expand Down Expand Up @@ -953,11 +948,6 @@
"count": 1
}
},
"src/app/api/v1/rerank/route.ts": {
"no-restricted-imports": {
"count": 1
}
},
"src/app/api/v1/vscode/[token]/models/route.ts": {
"no-restricted-syntax": {
"count": 1
Expand Down Expand Up @@ -3259,4 +3249,4 @@
"count": 5
}
}
}
}
Loading
Loading