Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
e307f29
feat(reasoning): adaptive effort via X-OmniRoute-Effort header and de…
patrykkopycinski Sep 12, 2026
8867c1c
refactor(effort): extract adaptive-effort wiring out of chatCore.ts
patrykkopycinski Sep 15, 2026
59a7983
Merge remote-tracking branch 'origin/release/v3.8.51' into rework/13448
diegosouzapw Sep 16, 2026
97b26f3
fix(sse): scope adaptive-effort wiring to OpenAI dispatch, widen defa…
diegosouzapw Sep 16, 2026
135ad15
docs(routing): correct adaptive-effort scope claim to OpenAI-dispatch…
diegosouzapw Sep 16, 2026
7cc019a
feat(reasoning): adaptive effort via X-OmniRoute-Effort header and de…
patrykkopycinski Sep 12, 2026
72c5846
refactor(effort): extract adaptive-effort wiring out of chatCore.ts
patrykkopycinski Sep 15, 2026
447b4fc
refactor(effort): read x-omniroute-effort inside the wiring module
patrykkopycinski Sep 16, 2026
44fdd36
feat(reasoning): adaptive effort via X-OmniRoute-Effort header and de…
patrykkopycinski Sep 12, 2026
cea96b1
refactor(effort): extract adaptive-effort wiring out of chatCore.ts
patrykkopycinski Sep 15, 2026
2152564
refactor(effort): read x-omniroute-effort inside the wiring module
patrykkopycinski Sep 16, 2026
5c914ec
Merge the contributor's rebased head of #13448 into the maintainer re…
diegosouzapw Sep 16, 2026
6025c5b
Merge the maintainer rework of #13448 (OpenAI-dispatch scoping, auto …
diegosouzapw Sep 16, 2026
88d5e0c
fix(db): gate the auto-cleanup VACUUM on reclaimable space, not row c…
hartmark Sep 16, 2026
421d1ff
fix(devin): fall back to CLI probe when the HTTP API rejects a CLI-fo…
patrykkopycinski Sep 16, 2026
a2d945f
Merge release/v3.8.51 into #13448
diegosouzapw Sep 16, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 10 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -1161,9 +1161,18 @@ PROVIDER_LIMITS_SYNC_SPACING_MS=1500
#OMNIROUTE_WAL_GUARD_MAX_MB=256

# Minimum rows a cleanup must delete before the post-cleanup VACUUM runs. Default: 1000.
# 0 always vacuums when rows were freed. Used by: src/lib/db/cleanup.ts.
# 0 always vacuums when rows were freed. The post-cleanup VACUUM also runs when the
# reclaimable-space threshold below is met, whichever comes first (either signal fires it).
# Used by: src/lib/db/cleanup.ts::shouldVacuumAfterCleanup().
#OMNIROUTE_VACUUM_MIN_DELETED_ROWS=1000

# Minimum reclaimable space (MB) that alone justifies a full-database VACUUM after a
# cleanup, even when the row-count threshold above was not met (a handful of oversized
# blob rows can free far more space than thousands of tiny rows). VACUUM is synchronous
# and blocks the entire process. Default: 100. 0 always vacuums after any deletion.
# Used by: src/lib/db/cleanup.ts::getVacuumMinReclaimableBytes().
#OMNIROUTE_VACUUM_MIN_RECLAIMABLE_MB=100

# Explicit path to sql-wasm.wasm for the sql.js fallback adapter. Default: auto-detect.
# Used by: src/lib/db/adapters/sqljsAdapter.ts.
#OMNIROUTE_SQLJS_WASM_PATH=
Expand Down
1 change: 1 addition & 0 deletions changelog.d/features/13448-adaptive-reasoning-effort.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **feat(reasoning):** adaptive reasoning effort (`auto`) — the gateway resolves the thinking budget per user turn from deterministic request-shape signals (stateless per-turn pin) instead of forwarding a literal `auto`, applied at the gateway pre-translation for any harness (Claude Code, Cursor, Codex, opencode, Hermes) whose request dispatches to an OpenAI Chat-Completions-shaped upstream (`targetFormat === FORMATS.OPENAI` — `reasoning_effort` is an OpenAI-shaped field, so a Claude- or Gemini-targeted request is unaffected). Opt in via `X-OmniRoute-Effort: auto` or a model's `defaultReasoningEffort: "auto"` (now a valid `ModelSpec` value); any explicit client reasoning field always wins ([#13448](https://github.com/diegosouzapw/OmniRoute/pull/13448))
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(devin):** Fall back to the CLI probe (`devin acp --agent-type summarizer`) when the connection-test HTTP API rejects a CLI-format key, since routing authenticates against the local Devin CLI, not `api.devin.ai` ([#13617](https://github.com/diegosouzapw/OmniRoute/pull/13617))
5 changes: 0 additions & 5 deletions config/quality/eslint-suppressions.json
Original file line number Diff line number Diff line change
Expand Up @@ -1702,11 +1702,6 @@
"count": 1
}
},
"src/lib/providers/validation/webProvidersB.ts": {
"@typescript-eslint/no-unused-vars": {
"count": 1
}
},
"src/lib/quota/quotaAdapters.ts": {
"@typescript-eslint/no-unused-vars": {
"count": 1
Expand Down
3 changes: 2 additions & 1 deletion config/quality/file-size-baseline.json
Original file line number Diff line number Diff line change
Expand Up @@ -451,6 +451,7 @@
"_rebaseline_2026_09_15_13440_daily_reset_tz": "#13440 rework: open-sse/services/accountFallback.ts 2469->2493 (+24): +6 for the operator-clock-first branch in checkFallbackError non-TPD daily quota (nextConfiguredResetMs leaf lives in dailyQuotaReset.ts, under cap) and +18 from the mandatory lint-staged Prettier pass over pre-existing unformatted lines of the touched file (no logic). executeTargetAttempt.ts 1212->1215 and roundRobinCombo.ts 1205->1208 (+3 each): one import plus the rotation/dailyReset arguments at the existing checkFallbackError call site; the lookup itself is the new comboDailyResetClock.ts leaf (under cap). Covered by tests/unit/daily-reset-tz-threading.test.ts.",
"_rebaseline_2026_09_15_13672_retry_after_provenance": "#13672 rework (opt-in RETRY_AFTER_PROVENANCE_ENABLED): open-sse/services/combo/executeTargetAttempt.ts 1212->1220 (+8) and roundRobinCombo.ts 1205->1210 (+5) at the existing drain-path clone/parse block: capture the already-read body text, log an unreadable hint (debug for a non-JSON page, warn for a failed clone) instead of an empty catch, and one flag-gated prose fallback line; the import grows by the two helpers. Parsing, flag read and the Retry-After/provenance logic live in open-sse/utils/error.ts (under cap). Covered by tests/unit/retry-after-provenance.test.ts (flag off and on).",
"_rebaseline_2026_09_15_13439_protected_priority_stop_status": "#13439 rework (opt-in PROTECTED_PRIORITY_INFRA_502_ENABLED): open-sse/services/combo/executeTargetAttempt.ts 1212->1217 (+5): two import lines for the new protectedPriorityStopStatus.ts leaf (where the provably-non-quota cause list and the flag read live) and the predictive_ttft cause argument at the existing stopProtectedPriorityTarget call, which Prettier splits over three lines. Covered by tests/unit/combo/protected-priority-stop-status-13439.test.ts (every stop cause, flag off and on).",
"_rebaseline_2026_09_16_13448_adaptive_effort_targetformat_gate": "#13448 rework: open-sse/handlers/chatCore.ts 6142->6156 (+14, PR's own growth: the X-OmniRoute-Effort header capture near THINKING_MARKER_HEADER and the wireAdaptiveEffort(translatedBody, {...}) call site right after applyDefaultReasoningEffort, plus this rework's +1 targetFormat argument at that same call). The targetFormat gate itself (ctx.targetFormat !== FORMATS.OPENAI short-circuit) lives in the non-frozen open-sse/handlers/chatCore/adaptiveEffortWiring.ts leaf, not here -- irreducible call-site wiring at the existing post-translation reasoning-normalization chokepoint. Covered by tests/unit/adaptive-effort-wiring.test.ts (targetFormat gate, red-on-tip) and tests/unit/adaptive-effort-model-default-13448.test.ts.",
"_rebaseline_pr1043_minimax_tts": "Upstream port decolua/9router#1043 (toanalien) own growth: audioSpeech.ts 965->1061 (+96). Adds MiniMax T2A v2 TTS dispatch (handleMinimaxSpeech + hexToBytes helper) — provider entry was already in audioRegistry (format: minimax-tts) but no handler existed, falling through to the OpenAI-compatible default that fails (T2A has custom shape + hex-encoded audio + base_resp envelope). New branch sits next to the other inline provider branches (xiaomi-mimo, coqui, tortoise, aws-polly) — extracting would just create indirection. Covered by tests/unit/minimax-tts-1043.test.ts (3 tests, GREEN: success, base_resp error, invalid-hex).",
"_rebaseline_pr4592_exclude_exhausted_auto": "Reconcile #4592 already-merged growth: combo.ts 2991->3036 (+45, terminal-status quota-cutoff exclusion in buildAutoCandidates + opt-in gate). Fast-gate PR->release does not run check:file-size.",
"open-sse/executors/antigravity.ts": 1665,
Expand All @@ -459,7 +460,7 @@
"open-sse/executors/codex.ts": 1505,
"open-sse/executors/cursor.ts": 1808,
"open-sse/executors/muse-spark-web.ts": 1405,
"open-sse/handlers/chatCore.ts": 6146,
"open-sse/handlers/chatCore.ts": 6156,
"open-sse/handlers/imageGeneration.ts": 3293,
"open-sse/handlers/search.ts": 1789,
"open-sse/mcp-server/schemas/tools.ts": 1621,
Expand Down
3 changes: 2 additions & 1 deletion docs/reference/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,7 +104,8 @@ OmniRoute uses **SQLite** (via `better-sqlite3`) for all persistence. These vari
| `OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS` | `21600000` (6h) | `src/lib/db/walMaintenance.ts` | Override the periodic `wal_checkpoint(TRUNCATE)` interval (ms). Auto-checkpoint never shrinks the WAL file itself, and a long-running server never closes its DB. `0` disables. |
| `OMNIROUTE_WAL_PASSIVE_INTERVAL_MS` | `300000` (5m) | `src/lib/db/walMaintenance.ts` | Override the frequent `wal_checkpoint(PASSIVE)` interval (ms). Keeps pending WAL frames small so the periodic TRUNCATE never copies a multi-GB backlog on the main thread. `0` disables. |
| `OMNIROUTE_WAL_GUARD_MAX_MB` | `256` | `src/lib/db/walMaintenance.ts` | When a PASSIVE tick finds the WAL file above this size, escalate to `wal_checkpoint(TRUNCATE)` immediately instead of waiting for the slow tick. |
| `OMNIROUTE_VACUUM_MIN_DELETED_ROWS` | `1000` | `src/lib/db/cleanup.ts` | Post-cleanup VACUUM only runs when the cleanup deleted at least this many rows. `0` means always VACUUM when a cleanup freed any rows; `1` effectively disables the gate. VACUUM rewrites the entire database (multi-GB WAL + I/O burst on large DBs), so tiny cleanups skip it; the Storage page's scheduled VACUUM (default weekly, #4437) and manual VACUUM still reclaim space. |
| `OMNIROUTE_VACUUM_MIN_DELETED_ROWS` | `1000` | `src/lib/db/cleanup.ts` | Post-cleanup VACUUM runs when the cleanup deleted at least this many rows, OR when `OMNIROUTE_VACUUM_MIN_RECLAIMABLE_MB` below is met (whichever fires first). `0` means always VACUUM when a cleanup freed any rows; `1` effectively disables the row-count gate. VACUUM rewrites the entire database (multi-GB WAL + I/O burst on large DBs), so tiny cleanups skip it; the Storage page's scheduled VACUUM (default weekly, #4437) and manual VACUUM still reclaim space. |
| `OMNIROUTE_VACUUM_MIN_RECLAIMABLE_MB` | `100` | `src/lib/db/cleanup.ts` | Minimum reclaimable space (SQLite's own free-page count, in MB) that alone triggers the post-cleanup `VACUUM`, even when `OMNIROUTE_VACUUM_MIN_DELETED_ROWS` was not met -- a handful of oversized blob rows can free far more space than thousands of tiny rows. `VACUUM` is synchronous and blocks the entire process (15-20 min on a multi-GB database). `0` always vacuums after any deletion. |
| `OMNIROUTE_PRESSURE_SELF_RESTART` | `false` | `open-sse/utils/resourcePressure.ts` | Set to `1`/`true`/`yes`/`on` to exit the process after critical resource pressure is sustained for `OMNIROUTE_PRESSURE_SELF_RESTART_AFTER_MS`, letting a supervisor (systemd `Restart=always`, Docker restart policy) bring back a clean process instead of serving 503s indefinitely. |
| `OMNIROUTE_PRESSURE_SELF_RESTART_AFTER_MS` | `120000` (2m) | `open-sse/utils/resourcePressure.ts` | How long critical pressure must persist before the self-restart exit fires. |
| `OMNIROUTE_SQLJS_WASM_PATH` | _(auto-detect)_ | `src/lib/db/adapters/sqljsAdapter.ts` | Explicit path (absolute or relative to cwd) to `sql-wasm.wasm` when using the `sql.js` WASM fallback adapter. Auto-detected via package dependencies and candidate layouts when unset. |
Expand Down
11 changes: 6 additions & 5 deletions docs/routing/AUTO-COMBO.md
Original file line number Diff line number Diff line change
Expand Up @@ -252,11 +252,12 @@ combo's stored config. These apply only to the `auto` strategy and only for the
that carries them; the combo's saved `modePack`/`budgetCap`/`budgetFallback` are used
when the header is absent.

| Header | Accepts | Effect |
| :---------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `X-OmniRoute-Mode` | a preset alias (`fast`, `balanced`, `quality`, `cheap`, `reliable`, `offline`) or a raw pack name (`ship-fast`, `cost-saver`, `quality-first`, `offline-friendly`, `reliability-first`) | Overrides the scoring weights for this request. `balanced`/`default` force the default weights (no pack). Unknown values are ignored (config preserved). |
| `X-OmniRoute-Budget` | a positive number (max USD per request) | Hard cost ceiling: candidates whose estimated cost exceeds it are filtered before selection. What happens when **every** candidate exceeds it is controlled by `X-OmniRoute-Budget-Fallback` below. |
| `X-OmniRoute-Budget-Fallback` | `cheapest` (default, aliases: `cheapest-viable`, `soft`) or `strict` (aliases: `block`, `hard`) | `cheapest`: falls back to the globally cheapest candidate even though it still exceeds the cap (legacy behavior). `strict`: refuses to select — the request fails fast with `HTTP 402` instead of silently overspending. Unknown values are ignored. |
| Header | Accepts | Effect |
| :---------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `X-OmniRoute-Mode` | a preset alias (`fast`, `balanced`, `quality`, `cheap`, `reliable`, `offline`) or a raw pack name (`ship-fast`, `cost-saver`, `quality-first`, `offline-friendly`, `reliability-first`) | Overrides the scoring weights for this request. `balanced`/`default` force the default weights (no pack). Unknown values are ignored (config preserved). |
| `X-OmniRoute-Budget` | a positive number (max USD per request) | Hard cost ceiling: candidates whose estimated cost exceeds it are filtered before selection. What happens when **every** candidate exceeds it is controlled by `X-OmniRoute-Budget-Fallback` below. |
| `X-OmniRoute-Budget-Fallback` | `cheapest` (default, aliases: `cheapest-viable`, `soft`) or `strict` (aliases: `block`, `hard`) | `cheapest`: falls back to the globally cheapest candidate even though it still exceeds the cap (legacy behavior). `strict`: refuses to select — the request fails fast with `HTTP 402` instead of silently overspending. Unknown values are ignored. |
| `X-OmniRoute-Effort` | `auto` (other values reserved) | Adaptive thinking budget: when the request carries **no** reasoning field of any shape (`reasoning_effort`, `reasoning`, `thinking`), the gateway resolves `auto` to `low`/`medium`/`high` from deterministic request-shape signals (last-user-message length, context size up to the last user message, prior tool results, tool-loop depth). Signals are scoped to the current turn — everything after the last user message is ignored — so every request in a tool loop resolves to the same level (stateless per-turn pin, no session state, no mid-loop escalation that would break upstream prompt-cache prefixes). An explicit client reasoning field always wins. Scoped to requests whose upstream dispatch resolves to the OpenAI Chat Completions shape (`targetFormat === FORMATS.OPENAI`) — `reasoning_effort` is an OpenAI-shaped field, so the header is a no-op on a Claude- or Gemini-targeted request (see `open-sse/handlers/chatCore/adaptiveEffortWiring.ts`). |

```bash
# Force the fastest profile, cap this request at $0.05, and hard-block instead of overspending
Expand Down
Loading
Loading