Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,25 @@

---

## [3.7.1] — 2026-04-27

### 🐛 Bug Fixes

- **Earliest-Reset-First Routing — Production Hotfix Bundle (4 fixes):** v3.7.0 shipped the `earliest-reset-first` strategy with an additive `S = 0.85·T + 0.15·min(Q, 30)` score. Production observation revealed four regressions, fixed in this patch:
- **F1 (self-hosted fallback):** Accounts with no quota cache (e.g. user-added `openai-compatible` providers like local llama.cpp) returned `excluded: "no_quota_data"` and produced an `all candidates excluded` 503. They now fall back to `score = 0` and remain eligible — paid candidates with positive scores still beat them, but they are usable when no paid candidate exists. `markAccountExhaustedFrom429()` empty caches are still excluded.
- **F2 (per-model windows ignored):** Removed `modelWindowMapping.ts` and the model-specific weekly-window lookup. Routing now reads only the overall `weekly` window — per-model `weekly Sonnet` / `weekly Omelette` buckets no longer split routing across requests for the same conversation. If Anthropic 429s an opus request because per-model quota is exhausted, `accountFallback` retries on another account.
- **F3 (fresh quota max-urgency):** Anthropic omits `resetAt` for accounts where the timer hasn't started yet (always at 100% remaining). v3.7.0 treated that as `missing` and excluded fresh accounts. v3.7.1 detects the strict pair `Q === 100 AND resetAt === null` and scores the track at `100 × 100 = 10000` (max urgency), so the very first request wakes the timer instead of further burning an already-active account.
- **F4 (multiplicative burn-down):** Replaced the additive score with `S = T_pts × Q_remain` (`score ∈ [0, 10000]`). The new formula makes "quota likely lost if not used" the actual scoring objective: `APEX(17%/2h) avg=690` now correctly loses to `GNUMAX(87%/4h) avg=1460` even though APEX's session is closer to reset. Penalties rescaled 100× (`PENALTY_DEGRADED 25→2500`, `PENALTY_BACKOFF_WEIGHT 1→100`, `PENALTY_ERROR_WEIGHT 0.4→40`) to preserve their relative impact across the new score range. `Q_SATURATION_CAP` removed (multiplicative scoring naturally weights remaining quota).
- **Plan & Review:** `.sisyphus/plans/routing-strategy-v6.md` documents derivation, hand-math, and Oracle/Momus review history (4 logic blockers + 9 process items resolved). `routing-strategy-v5.md` is marked superseded.

### 🔧 Internal

- `selectByEarliestResetFirst(candidates, sessionId)` — `modelHint` parameter dropped (also from `scoreAccount`, `scoreWeeklyTrack`, `isAffinityValid`).
- `scoreWeeklyTrack(connId)` simplified — single `getQuotaWindowStatus(connId, "weekly")` call, no bottleneck logic.
- New unit regression coverage: 18 new tests (F1×4, F2×2, F3×6, F4×6) plus updated existing assertions.

---

## [3.7.0] — 2026-04-26

### ✨ New Features
Expand Down
2 changes: 1 addition & 1 deletion docs/openapi.yaml
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
openapi: 3.1.0
info:
title: OmniRoute API
version: 3.7.0
version: 3.7.1
description: |
OmniRoute is a local-first AI API proxy router. It provides an OpenAI-compatible
endpoint that routes requests to multiple AI providers with load balancing,
Expand Down
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "omniroute",
"version": "3.7.0",
"version": "3.7.1",
"description": "Smart AI Router with auto fallback — route to FREE & cheap models, zero downtime. Works with Cursor, Cline, Claude Desktop, Codex, and any OpenAI-compatible tool.",
"type": "module",
"bin": {
Expand Down
15 changes: 7 additions & 8 deletions src/sse/services/auth.ts
Original file line number Diff line number Diff line change
Expand Up @@ -700,14 +700,13 @@ export async function getProviderCredentials(
const selectedId = getNextFromDeckSync(`conn:${provider}`, ids);
connection = orderedConnections.find((c) => c.id === selectedId) || orderedConnections[0];
} else if (strategy === "earliest-reset-first") {
// Cache-aware burn-down: route to the account whose quota resets soonest,
// with affinity to the same account within a 5-min sliding window. See
// .sisyphus/plans/routing-strategy-v4.md and earliestResetFirst.ts.
const result = selectByEarliestResetFirst(
orderedConnections,
requestedModel,
options.sessionId || null
);
// Cache-aware burn-down: route to the account whose quota is most
// urgent to drain before reset (largest T_pts × Q_remain product), with
// affinity to the same account within a 5-min sliding window. v6
// dropped the modelHint parameter — model-specific weekly windows are
// no longer consulted for ranking or hard exclusion. See
// .sisyphus/plans/routing-strategy-v6.md and earliestResetFirst.ts.
const result = selectByEarliestResetFirst(orderedConnections, options.sessionId || null);
if ("allExcluded" in result) {
log.warn(
"AUTH",
Expand Down
181 changes: 115 additions & 66 deletions src/sse/services/strategies/earliestResetFirst.ts
Original file line number Diff line number Diff line change
@@ -1,18 +1,24 @@
// Earliest-reset-first burn-down routing strategy.
// Earliest-reset-first burn-down routing strategy (v6).
//
// Selects the provider account whose quota will reset SOONEST, with a
// saturating quota-room cap so accounts about to reset (and about to "lose"
// any unused quota anyway) are preferred. See .sisyphus/plans/routing-strategy-v4.md
// for derivation, simulations, and Oracle review history.
// Selects the provider account whose quota is most urgent to drain before
// reset — i.e. has the largest "T_pts × Q_remain" product. v6 supersedes v4
// after production observation revealed four regressions:
// F1: self-hosted/openai-compatible (no quota cache) returned excluded
// F2: per-model weekly windows (Omelette/Sonnet) split routing per request
// F3: 100% fresh accounts (resetAt=null until first request) were excluded
// F4: additive S = 0.85·T + 0.15·min(Q,30) kept exhausted-but-urgent
// accounts winning over abundant-but-cool accounts
//
// See `.sisyphus/plans/routing-strategy-v6.md` for derivation, hand-math, and
// Oracle / Momus review history.

import { getQuotaWindowStatus } from "@/domain/quotaCache";
import { getQuotaWindowStatus, isAccountQuotaExhausted } from "@/domain/quotaCache";
import {
getSessionConnection,
getSessionInfo,
touchSession,
} from "@omniroute/open-sse/services/sessionManager.ts";
import { isTerminalConnectionStatus } from "@/sse/services/accountTerminalStatus";
import { mapModelToRequiredWeekly } from "@/sse/services/strategies/modelWindowMapping";

interface ConnectionLike {
id: string;
Expand Down Expand Up @@ -93,21 +99,22 @@ export interface SelectionTrace {

const SESSION_AFFINITY_WINDOW_MS = 5 * 60 * 1000;
const SCORE_TIE_EPSILON = 1e-9;
const TRACK_WEIGHT_T = 0.85;
const TRACK_WEIGHT_Q = 0.15;
const Q_SATURATION_CAP = 30;
const MIN_USABLE_REMAINING_PCT = 5;
const PENALTY_ERROR_WEIGHT = 0.4;
const PENALTY_BACKOFF_WEIGHT = 1.0;

// v6 multiplicative score range is [0, 10000] (T_pts ∈ [10,100] × Q ∈ [5,100]).
// Penalties scale 100× from v4 to keep their relative impact: the strongest
// backoff (level 4) still subtracts ~baseScore from a typical mid-range
// account, and a single degraded track still costs ~25% of a typical score.
const PENALTY_ERROR_WEIGHT = 40;
const PENALTY_BACKOFF_WEIGHT = 100;
const PENALTY_BACKOFF_CAP = 100;
const PENALTY_BACKOFF_PER_LEVEL = 25;
const PENALTY_DEGRADED = 25;
const PENALTY_DEGRADED = 2500;
const ERROR_RATE_WINDOW_MS = 15 * 60 * 1000;

// Stepwise piecewise functions. Boundaries are inclusive of the upper bound.
// `null` input (missing resetAt) returns null so callers can branch.
// Past resets (s <= 0) report maximal urgency so stale entries route quickly
// to allow refresh on the next cache tick.
// Stepwise time-urgency tables. `null` input (missing resetAt) is handled by
// callers via the F3 fresh-quota branch. Past resets (s <= 0) report maximal
// urgency so stale entries route quickly to refresh on the next cache tick.
export function sessionTimePoints(secondsRemaining: number | null): number | null {
if (secondsRemaining === null) return null;
if (secondsRemaining <= 0) return 100;
Expand Down Expand Up @@ -153,72 +160,75 @@ export function scoreSessionTrack(connId: string): TrackResult {
}

const sec = deltaSec(status.resetAt);

// F3 fresh quota: Anthropic omits resetAt only when the session is at 100%
// and the timer hasn't started. Treat as max urgency so the very first
// request to that account starts the 5h timer rather than letting an
// already-burning account exhaust further. Q≠100 with null resetAt is
// ambiguous (API hiccup, stale cache, unsupported window) — keep missing.
if (sec === null && Q === 100) {
return {
kind: "known",
score: 100 * 100,
remainingPct: Q,
resetAt: null,
secondsToReset: Number.POSITIVE_INFINITY,
windowName: "session",
};
}
if (sec === null) return { kind: "missing" };

const T = sessionTimePoints(sec);
if (T === null) return { kind: "missing" };

const Qcap = Math.min(Q, Q_SATURATION_CAP);
return {
kind: "known",
score: TRACK_WEIGHT_T * T + TRACK_WEIGHT_Q * Qcap,
score: T * Q,
remainingPct: Q,
resetAt: status.resetAt,
secondsToReset: sec,
windowName: "session",
};
}

export function scoreWeeklyTrack(connId: string, modelHint: string | null): TrackResult {
// F2: model-specific weekly windows (Omelette/Sonnet) are deliberately ignored
// in v6 per user instruction. Routing reads only the overall `weekly` window —
// model-aware quota separation is not modelled. If Anthropic 429s an opus
// request because per-model quota is exhausted, accountFallback handles retry.
export function scoreWeeklyTrack(connId: string): TrackResult {
const overall = getQuotaWindowStatus(connId, "weekly", 90);
const requiredWindow = mapModelToRequiredWeekly(modelHint);
const modelSpecific = requiredWindow ? getQuotaWindowStatus(connId, requiredWindow, 90) : null;

// Hard exclusions MUST take precedence over degraded fallback (Copilot review C1):
// an account with overall weekly < 5% is genuinely out of quota and must be
// excluded even if its model-specific window is missing.
if (overall && overall.remainingPercentage < MIN_USABLE_REMAINING_PCT) {
return { kind: "excluded", reason: "weekly_overall<5%", resetAt: overall.resetAt };
}

if (modelSpecific && modelSpecific.remainingPercentage < MIN_USABLE_REMAINING_PCT) {
return {
kind: "excluded",
reason: `${requiredWindow || "weekly_model"}<5%`,
resetAt: modelSpecific.resetAt,
};
}
if (!overall) return { kind: "missing" };

if (requiredWindow && !modelSpecific) {
return { kind: "degraded", reason: `${requiredWindow}_missing` };
if (overall.remainingPercentage < MIN_USABLE_REMAINING_PCT) {
return { kind: "excluded", reason: "weekly<5%", resetAt: overall.resetAt };
}

if (!overall && !modelSpecific) return { kind: "missing" };

const candidates = [overall, modelSpecific].filter(
(c): c is NonNullable<typeof c> => c !== null && c !== undefined
);
const sec = deltaSec(overall.resetAt);
const Q = overall.remainingPercentage;

let bottleneck = candidates[0];
for (const c of candidates) {
if (c.remainingPercentage < bottleneck.remainingPercentage) bottleneck = c;
// F3 fresh quota for weekly track (same semantics as session).
if (sec === null && Q === 100) {
return {
kind: "known",
score: 100 * 100,
remainingPct: Q,
resetAt: null,
secondsToReset: Number.POSITIVE_INFINITY,
windowName: "weekly",
};
}

const sec = deltaSec(bottleneck.resetAt);
if (sec === null) return { kind: "missing" };

const T = weeklyTimePoints(sec);
if (T === null) return { kind: "missing" };

const Q = bottleneck.remainingPercentage;
const Qcap = Math.min(Q, Q_SATURATION_CAP);
return {
kind: "known",
score: TRACK_WEIGHT_T * T + TRACK_WEIGHT_Q * Qcap,
score: T * Q,
remainingPct: Q,
resetAt: bottleneck.resetAt,
resetAt: overall.resetAt,
secondsToReset: sec,
windowName: bottleneck === overall ? "weekly" : requiredWindow || "weekly_model",
windowName: "weekly",
};
}

Expand Down Expand Up @@ -252,9 +262,9 @@ function earliestKnownReset(tracks: TrackResult[]): string | null {
return new Date(Math.min(...dates)).toISOString();
}

export function scoreAccount(conn: ConnectionLike, modelHint: string | null): ScoredCandidate {
export function scoreAccount(conn: ConnectionLike): ScoredCandidate {
const s = scoreSessionTrack(conn.id);
const w = scoreWeeklyTrack(conn.id, modelHint);
const w = scoreWeeklyTrack(conn.id);

if (s.kind === "excluded" || w.kind === "excluded") {
const ex = s.kind === "excluded" ? s : (w as TrackExcluded);
Expand All @@ -281,7 +291,41 @@ export function scoreAccount(conn: ConnectionLike, modelHint: string | null): Sc
}

if (trackScores.length === 0) {
return { conn, excluded: true, reason: "no_quota_data" };
// F1: distinguish "no cache entry exists" (self-hosted / openai-compatible
// never reports usage → score=0 fallback so paid candidates still beat
// them but they remain eligible) from "cache says 429-exhausted with no
// resetAt" (markAccountExhaustedFrom429 → must stay excluded until
// refresh / TTL).
if (isAccountQuotaExhausted(conn.id)) {
return {
conn,
excluded: true,
reason: "quota_exhausted_unknown_reset",
resetAt: null,
breakdown: {
s,
w,
P_error: 0,
P_backoff: 0,
degraded_pen: 0,
baseScore: 0,
},
};
}
return {
conn,
excluded: false,
score: 0,
earliestReset: null,
breakdown: {
s,
w,
P_error: 0,
P_backoff: 0,
degraded_pen: 0,
baseScore: 0,
},
};
}

const baseScore = trackScores.reduce((a, b) => a + b, 0) / trackScores.length;
Expand Down Expand Up @@ -329,11 +373,7 @@ export function candidateComparator(a: ScoredCandidate, b: ScoredCandidate): num
return a.conn.id.localeCompare(b.conn.id);
}

export function isAffinityValid(
conn: ConnectionLike,
modelHint: string | null,
sessionId: string | null
): AffinityValidity {
export function isAffinityValid(conn: ConnectionLike, sessionId: string | null): AffinityValidity {
if (conn.isActive === false) return { valid: false, reason: "inactive" };

if (conn.rateLimitedUntil) {
Expand All @@ -345,6 +385,16 @@ export function isAffinityValid(

if (isTerminalConnectionStatus(conn)) return { valid: false, reason: "terminal" };

// Mirror the F1 guard from scoreAccount: a connection marked exhausted via
// markAccountExhaustedFrom429() has empty `quotas:{}` and `exhausted:true`,
// making both tracks return `missing` (not `excluded`). Without this check,
// isAffinityValid would keep affinity pinned and the next request to this
// account would be sent right back to a 429-exhausted endpoint. Copilot
// review on PR #23.
if (isAccountQuotaExhausted(conn.id)) {
return { valid: false, reason: "quota_exhausted_unknown_reset" };
}

const session = getSessionInfo(sessionId);
if (!session) return { valid: false, reason: "session_expired" };
if (Date.now() - session.lastActive > SESSION_AFFINITY_WINDOW_MS) {
Expand All @@ -353,7 +403,7 @@ export function isAffinityValid(

const s = scoreSessionTrack(conn.id);
if (s.kind === "excluded") return { valid: false, reason: s.reason };
const w = scoreWeeklyTrack(conn.id, modelHint);
const w = scoreWeeklyTrack(conn.id);
if (w.kind === "excluded") return { valid: false, reason: w.reason };
Comment on lines 404 to 407

Copilot AI Apr 26, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

isAffinityValid can incorrectly keep affinity pinned to a connection that was marked quota-exhausted via markAccountExhaustedFrom429(). In that case both tracks are missing (empty quotas) so this function returns valid: true, while scoreAccount() would exclude it via isAccountQuotaExhausted(). Consider checking isAccountQuotaExhausted(conn.id) here (and returning invalid) so affinity breaks and selection can fall back to another account; add a regression test for this scenario.

Copilot uses AI. Check for mistakes.

return { valid: true };
Expand Down Expand Up @@ -385,21 +435,20 @@ function buildAllExcluded(scored: ScoredCandidate[]): AllExcludedResult {

export function selectByEarliestResetFirst(
candidates: ConnectionLike[],
modelHint: string | null,
sessionId: string | null
): { selected: ConnectionLike } | AllExcludedResult {
if (sessionId) {
const boundId = getSessionConnection(sessionId);
if (boundId) {
const bound = candidates.find((c) => c.id === boundId);
const affinity = bound ? isAffinityValid(bound, modelHint, sessionId) : null;
const affinity = bound ? isAffinityValid(bound, sessionId) : null;
if (bound && affinity?.valid === true) {
return { selected: bound };
}
}
}

const scored: ScoredCandidate[] = candidates.map((c) => scoreAccount(c, modelHint));
const scored: ScoredCandidate[] = candidates.map((c) => scoreAccount(c));
const usable = scored.filter((s) => !s.excluded);

if (usable.length === 0) return buildAllExcluded(scored);
Expand Down
26 changes: 0 additions & 26 deletions src/sse/services/strategies/modelWindowMapping.ts

This file was deleted.

Loading
Loading