diff --git a/docs/ops/MONITORING_GUIDE.md b/docs/ops/MONITORING_GUIDE.md index f69848c4ed4..24172a7a259 100644 --- a/docs/ops/MONITORING_GUIDE.md +++ b/docs/ops/MONITORING_GUIDE.md @@ -1,7 +1,7 @@ --- title: "Monitoring & Observability Guide" version: 3.8.50 -lastUpdated: 2026-08-13 +lastUpdated: 2026-09-18 --- # Monitoring & Observability Guide @@ -421,6 +421,13 @@ previous alert state (no flapping on low traffic). Payloads never contain prompt responses, API keys, connection ids or account ids. Subscribe a webhook to these events (or `*`) in the dashboard. Alerts are off by default: set `slo.alertsEnabled = true` to start evaluation, so existing `*` subscribers do not receive new events after an upgrade. +Turning alerts off forgets the remembered alert state, so turning them back on never replays a +transition that happened while they were off. A tick is skipped while the previous one is still +dispatching webhooks, so slow endpoints cannot cause duplicate or interleaved alerts. + +The alert loop and `GET /api/metrics` read circuit breakers without changing them: an `OPEN` +breaker whose cooldown has elapsed is reported as `HALF_OPEN`, but a scrape never transitions +it, persists it or grants its half-open probe; only live traffic does. --- @@ -499,6 +506,9 @@ Enforced by `open-sse/services/routing/metricLabels.ts`: - Bounded values: first 64 providers, 100 models, 16 strategies, 16 engines are kept; later distinct values collapse into `other`. Each family is additionally capped at 2000 series. Combo names, connection ids, request ids and finish reasons are never labels. + The fixed values `redacted`, `unknown` and `other` never take a slot, and only a model + that served a successful response takes a model slot: failed requests for model ids a + client made up are counted under `other` unless the model is already tracked. - Values are restricted to `[A-Za-z0-9._:/-]`, truncated to 80 chars, and replaced by `redacted` when they look like a credential (API-key prefixes such as `sk-`, `ghp_`), an opaque token (32+ alphanumerics), an e-mail address or a UUID. @@ -528,6 +538,12 @@ defaults by `src/lib/monitoring/sloSettings.ts`. Evaluated by `failed` excludes `cancelled` and `guardrail_blocked` (client or policy decisions). Rate limits and timeouts count against availability but not against `error_rate`. +For `provider_recovery`, an ongoing open episode (breaker still `OPEN` or `HALF_OPEN`) counts +only while that breaker fails, or opens, inside the window. A breaker only leaves `HALF_OPEN` +when traffic probes it, so a provider that stopped receiving traffic would otherwise report a +breach forever. Such an idle open breaker adds no sample; it still shows in +`omniroute_circuit_breakers` and `omniroute_circuit_breaker_open`. + `GET /api/telemetry/summary` `errorRate` (percent) uses the same definition — failed / (success + failed) routed requests in the requested window (clamped to 1–60 min). diff --git a/docs/routing/ROUTING_CONTRACT.md b/docs/routing/ROUTING_CONTRACT.md index aebaaf2930c..bcbf21f08e5 100644 --- a/docs/routing/ROUTING_CONTRACT.md +++ b/docs/routing/ROUTING_CONTRACT.md @@ -1,7 +1,7 @@ --- title: "Routing Contract and Route Explanation" version: 3.8.54 -lastUpdated: 2026-09-14 +lastUpdated: 2026-09-18 --- # Routing Contract and Route Explanation @@ -17,9 +17,9 @@ The types live in `src/shared/contracts/routing.ts` and are shared by `src/` and | Type | Purpose | | ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `RoutingRequest` | What is routed: `requestId`, `model`, `protocol`, optional `capabilities`, `workspaceId`, `policyId`, `stream`, `budget` | -| `RoutingBudget` | Optional `maxCost` (USD) and `maxLatencyMs`, applied to the decision and to every failover attempt | +| `RoutingBudget` | Optional `maxCost` (USD) and `maxLatencyMs`; read by previews and by the `attemptPolicy.ts` library, not by live traffic (see Guarantees) | | `RoutingCandidate` | One provider/model: `score`, weighted `factors`, `eligible`, `exclusionReasons`, `quota`, `circuit`, estimated cost and latency | -| `RoutingDecision` | `decisionId`, `requestId`, `selected`, all `candidates`, `policyVersion`, `generatedAt`, `liveRequestExecuted`, `selectionMode` | +| `RoutingDecision` | `decisionId`, `requestId`, `selected`, `candidates`, `policyVersion`, `generatedAt`, `liveRequestExecuted`, `selectionMode`, `strategy`, optional `omittedCandidates` | | `ProviderAttempt` | One upstream call: provider, model, attempt number, start time, duration, status, `outcome` | | `RoutingExclusionReason` | `not_in_candidate_pool`, `model_not_found`, `capability_missing`, `quota_exhausted`, `circuit_open`, `self_healing_excluded`, `cost_over_budget`, `latency_over_budget` | @@ -27,22 +27,24 @@ Candidates never carry connection or account identifiers, prompts or credentials ## Guarantees -| Guarantee | How it holds | -| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| Preview uses the live selection algorithm | Live selection and preview both run `selectProviderWithTrace()` in `open-sse/services/autoCombo/engine.ts` | -| Preview calls no provider, changes no state | The preview passes `previewSelectionDeps()`: clones of the self-healing and rotation state, exploration off | -| Policy version on every decision | `computeRoutingPolicyVersion()` hashes combo name, candidate pool, weights, mode pack, budget, exploration rate and router strategy (`rp_` + 16 hex characters) | -| Every exclusion has a reason | `hardExclusionReasons()` and the engine trace in `open-sse/services/autoCombo/routingDecision.ts` | -| Unknown quota is not exhausted | `quota: "unknown"` stays eligible and is scored neutral; only a quota cutoff produces `quota_exhausted` | -| No blind retry of permanent errors | `open-sse/services/routing/attemptPolicy.ts`: authentication, model not found, invalid request and exhausted quota are permanent and never retried on the same candidate | -| Failover respects the budget | `checkFailoverBudget()` and `planNextAttempt()`; the live auto strategy keeps its failover chain inside the cost cap (`orderTargetsByCostBudget()`) | -| Deterministic circuit breaker | `src/shared/utils/circuitBreaker.ts` accepts an injected clock; `peekState()` reads the effective state without changing it | -| End-to-end correlation | Live decisions are recorded under the request id the client receives (`x-request-id`) and under their decision id | +| Guarantee | How it holds | +| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Preview uses the live selection algorithm | Live selection and preview both run `selectProviderWithTrace()` in `open-sse/services/autoCombo/engine.ts` | +| Preview calls no provider, changes no state | The preview passes `previewSelectionDeps()`: clones of the self-healing and rotation state, exploration off | +| Policy version on every decision | `computeRoutingPolicyVersion()` hashes combo name, candidate pool, weights, mode pack, budget, exploration rate and router strategy (`rp_` + 16 hex characters) | +| Every exclusion has a reason | `hardExclusionReasons()` and the engine trace in `open-sse/services/autoCombo/routingDecision.ts` | +| Unknown quota is not exhausted | `quota: "unknown"` stays eligible and is scored neutral; only a quota cutoff produces `quota_exhausted` | +| No blind retry of permanent errors | Live: the combo loops retry the same target only on 408, 429, 500, 502, 503 and 504 (`isRetryableAttemptStatus()` in `open-sse/services/routing/attemptPolicy.ts`) | +| Auto combo failover stays in the cost cap | Live, `rules` router strategy only: `orderTargetsByCostBudget()` drops (`strict`) or moves last (`cheapest`) targets whose estimated 1K-token request cost exceeds `budgetCap`. Explicit router strategies (`cost`, `latency`, `lkgp`, ...) ignore `budgetCap` | +| Deterministic circuit breaker | `src/shared/utils/circuitBreaker.ts` accepts an injected clock; `peekState()` reads the effective state without changing it | +| End-to-end correlation | Live decisions are recorded under the request id the client receives (`x-request-id`) and under their decision id | ## Previewing a decision `POST /api/omniroute/route/preview` requires management authentication and never calls a -provider. +provider. The examples on this page read `OMNIROUTE_MANAGE_KEY`, an OmniRoute API key with the +`manage` scope (a dashboard session works too), as in the +[API use cases](../reference/API_USE_CASES.md). - `{ "candidates": [...] }` keeps the original adaptive what-if ranking and its response shape. - `{ "engine": "auto", "request"?, "policy"?, "candidates": [...] }` runs the live auto-combo engine @@ -50,7 +52,7 @@ provider. ```bash curl -s -X POST http://localhost:20128/api/omniroute/route/preview \ - -H "Authorization: Bearer $OMNIROUTE_TOKEN" \ + -H "Authorization: Bearer $OMNIROUTE_MANAGE_KEY" \ -H "Content-Type: application/json" \ -d '{ "engine": "auto", @@ -65,14 +67,18 @@ curl -s -X POST http://localhost:20128/api/omniroute/route/preview \ The answer contains `selected`, `candidates` and `decision` (with `policyVersion` and `liveRequestExecuted: false`), plus the `x-request-id` and `x-omniroute-decision-id` headers. A -candidate without `quotaRemaining` is reported with `quota: "unknown"`. +candidate without `quotaRemaining` is reported with `quota: "unknown"`. The top-level `selected` +is only the provider id, kept for compatibility; `decision.selected` has the provider and the +model. An invalid body gets a 400 whose `error` string lists the invalid fields. ## Explaining a live request 1. The v1 chat completions, responses and messages routes run inside the request id the authorization pipeline stamps, so the router records its decision under that id. -2. Non-streaming responses carry `X-OmniRoute-Decision-Id` and `X-OmniRoute-Policy-Version`. - Streaming responses are sent before routing finishes; use the request id instead. +2. Responses carry `X-OmniRoute-Decision-Id` and `X-OmniRoute-Policy-Version` when a decision + was recorded, streaming responses included (the target is chosen before the stream starts). + The headers are absent when the request was not routed by an `auto` combo or when the response + headers cannot be changed; the request id lookup below still works. 3. `GET /api/omniroute/route/decisions/{id}` returns a decision by decision id or request id. Decisions are kept in memory for 30 minutes; unknown and malformed ids both return 404. 4. `GET /api/usage/route-explain/{id}` adds `routingDecision` to the call-log explanation when a @@ -81,7 +87,7 @@ candidate without `quotaRemaining` is reported with `quota: "unknown"`. ```bash curl -s http://localhost:20128/api/omniroute/route/decisions/ \ - -H "Authorization: Bearer $OMNIROUTE_TOKEN" + -H "Authorization: Bearer $OMNIROUTE_MANAGE_KEY" ``` ## Diagnostic mode @@ -97,6 +103,13 @@ contains prompts, credentials or connection identifiers. decision trace (`/api/usage/combo-trace/{id}`). - Live candidates do not yet distinguish unknown quota from a known value, because the candidate builder lives in a file at its size cap; preview candidates do. -- There is no live latency budget input; latency budgets apply to previews and to - `planNextAttempt()`. -- Decisions are held in memory and are lost on restart. +- Live traffic has no per-request `RoutingBudget` input. The only live cost limit is the auto + combo's `budgetCap` on the `rules` path, checked per attempt against a 1K-token estimate; there + is no cumulative spend check and no live latency budget. +- `classifyAttemptOutcome()`, `isPermanentAttemptOutcome()`, `canRetrySameCandidate()`, + `checkFailoverBudget()` and `planNextAttempt()` are a tested library with no live caller yet. + Previews apply `maxCost` and `maxLatencyMs` as exclusions. +- Decisions are held in memory and are lost on restart. A stored live decision keeps at most 40 + candidates (the selected one always), full factors only for the selected candidate and the 10 + best, and reports the rest as `omittedCandidates`. The store also has a 32 MB size budget, so + under heavy traffic a decision can be evicted before its 30 minutes are up. diff --git a/open-sse/services/autoCombo/routingDecision.ts b/open-sse/services/autoCombo/routingDecision.ts index 663b1a976a1..187334cdc89 100644 --- a/open-sse/services/autoCombo/routingDecision.ts +++ b/open-sse/services/autoCombo/routingDecision.ts @@ -203,6 +203,50 @@ function selectionModeOf(input: BuildRoutingDecisionInput): RoutingDecision["sel return describeRotation(input.outcome.trace.eligible); } +interface DecisionEntry { + key: string; + candidate: DecisionCandidateInput; + routing: RoutingCandidate; +} + +function isStrategyPick(candidate: DecisionCandidateInput, pick: StrategySelection): boolean { + if (candidate.provider !== pick.provider || candidate.model !== pick.model) return false; + return !pick.connectionId || (candidate.connectionId ?? "") === pick.connectionId; +} + +/** + * The candidate an explicit router strategy picked. The strategy chooses among the routable + * candidates on its own rules, so the scoring pass run only to explain the decision does not get + * to exclude its pick: a pick without a hard exclusion is reported eligible. A pick without a + * connection id matches the best-ranked candidate with its provider and model. + */ +function strategyPickEntry( + entries: DecisionEntry[], + input: BuildRoutingDecisionInput, + pick: StrategySelection +): DecisionEntry | undefined { + const entry = entries.find( + (candidateEntry) => + isStrategyPick(candidateEntry.candidate, pick) && + hardExclusionReasons(candidateEntry.candidate, input.request).length === 0 + ); + if (entry && !entry.routing.eligible) { + entry.routing = { ...entry.routing, eligible: true, exclusionReasons: [] }; + } + return entry; +} + +function selectedEntry( + entries: DecisionEntry[], + input: BuildRoutingDecisionInput +): DecisionEntry | undefined { + if (input.budgetExceeded) return undefined; + if (input.strategySelection) return strategyPickEntry(entries, input, input.strategySelection); + const chosen = input.outcome?.selection; + if (!chosen) return undefined; + return entries.find((entry) => entry.key === candidateKey(chosen) && entry.routing.eligible); +} + /** Turn a selection into the shared decision contract. Pure apart from the clock. */ export function buildRoutingDecision( input: BuildRoutingDecisionInput, @@ -213,20 +257,16 @@ export function buildRoutingDecision( if (!scoredByKey.has(candidateKey(scored))) scoredByKey.set(candidateKey(scored), scored); } const eligibleKeys = new Set((input.outcome?.trace.eligible ?? []).map(candidateKey)); - const entries = input.candidates.map((candidate) => ({ + const entries: DecisionEntry[] = input.candidates.map((candidate) => ({ key: candidateKey(candidate), + candidate, routing: toRoutingCandidate(candidate, input, scoredByKey, eligibleKeys), })); entries.sort( (a, b) => Number(b.routing.eligible) - Number(a.routing.eligible) || b.routing.score - a.routing.score ); - const chosen = input.budgetExceeded - ? undefined - : (input.strategySelection ?? input.outcome?.selection); - const selected = chosen - ? entries.find((entry) => entry.key === candidateKey(chosen) && entry.routing.eligible) - : undefined; + const selected = selectedEntry(entries, input); return { decisionId: clock.newDecisionId(), requestId: input.request.requestId, diff --git a/open-sse/services/combo/autoRoutingDecision.ts b/open-sse/services/combo/autoRoutingDecision.ts index 78057cc5d81..a5db71cb396 100644 --- a/open-sse/services/combo/autoRoutingDecision.ts +++ b/open-sse/services/combo/autoRoutingDecision.ts @@ -3,8 +3,9 @@ * * The scoring engine's pick is recorded as a `RoutingDecision` (candidates, scores, exclusion * reasons, policy version) under the request id, so a live request can be explained after it ran. - * The failover chain is kept inside the request cost budget: the budget cap applied to the first - * pick also applies to every later attempt. + * On the scoring-engine ("rules") path the failover chain is kept inside the request cost budget: + * the budget cap applied to the first pick also applies to every later attempt. Explicit router + * strategies ignore the budget cap, as they did before decisions were recorded. */ import type { RoutingDecision } from "@/shared/contracts/routing"; import { getRequestId } from "@/shared/utils/requestId"; @@ -128,7 +129,9 @@ export function recordExplicitStrategyDecision( } /** - * Keep the failover chain inside the request cost budget. With `budgetFallback: "strict"`, + * Keep the failover chain of a scoring-engine ("rules") selection inside the request cost budget. + * The estimate is per attempt (1K tokens at the candidate's price), not cumulative spend. With + * `budgetFallback: "strict"`, * targets whose estimated request cost exceeds `budgetCap` are dropped (unless that would leave * nothing, in which case the engine has already refused the request); otherwise they move behind * the in-budget targets. Targets without a known price are treated as in budget. diff --git a/open-sse/services/combo/resolveAutoStrategy.ts b/open-sse/services/combo/resolveAutoStrategy.ts index 7e1f2c6e10a..25583c76be6 100644 --- a/open-sse/services/combo/resolveAutoStrategy.ts +++ b/open-sse/services/combo/resolveAutoStrategy.ts @@ -366,6 +366,8 @@ export async function resolveAutoStrategyOrder( let selectedModel: string | null = null; let selectedConnectionId: string | null = null; let selectionReason = ""; + // The budget cap applies to the failover chain only when the scoring engine selected. + let failoverBudgetCap: number | null | undefined = null; const autoConfig: AutoComboConfig = { id: combo.id || combo.name, @@ -432,6 +434,7 @@ export async function resolveAutoStrategyOrder( selectedProvider = selection.provider; selectedModel = selection.model; selectedConnectionId = selection.connectionId ?? null; + failoverBudgetCap = budgetCap; selectionReason = `score=${selection.score.toFixed(3)}${selection.isExploration ? " (exploration)" : ""}`; } @@ -461,15 +464,18 @@ export async function resolveAutoStrategyOrder( // routable ranked ones (and, when the cutoff is OFF, makes this identical to // the pre-cutoff behavior), but a quota-blocked target still survives as a // final fallback instead of vanishing — the hard cutoff only de-prioritizes. - // The budget cap that bounded the first pick also bounds every failover attempt. + const failoverTargets = dedupeTargetsByExecutionKey( + [selectedTarget, ...rankedTargets, ...eligibleTargets].filter( + (entry): entry is ResolvedComboTarget => entry !== undefined && entry !== null + ) + ); + // On the scoring-engine path the budget cap that bounded the first pick also bounds every + // failover attempt. An explicit router strategy ignores budgetCap (as in v3.8.53): its pick + // is attempted first, so the recorded decision and the log name the first target tried. orderedTargets = orderTargetsByCostBudget( - dedupeTargetsByExecutionKey( - [selectedTarget, ...rankedTargets, ...eligibleTargets].filter( - (entry): entry is ResolvedComboTarget => entry !== undefined && entry !== null - ) - ), + failoverTargets, candidates, - budgetCap, + failoverBudgetCap, budgetFallback ); diff --git a/open-sse/services/routing/attemptPolicy.ts b/open-sse/services/routing/attemptPolicy.ts index b5c0419807c..6e62f147638 100644 --- a/open-sse/services/routing/attemptPolicy.ts +++ b/open-sse/services/routing/attemptPolicy.ts @@ -4,6 +4,13 @@ * * Permanent failures (bad credentials, unknown model, invalid request, exhausted quota) are never * retried on the same candidate: repeating them cannot succeed and only burns quota and time. + * + * What live traffic uses: only `isRetryableAttemptStatus()`. The combo loops (open-sse/services/ + * combo.ts) retry the same target only on 408, 429, 500, 502, 503 and 504, so 400/401/403/404 + * responses are never retried on the same target. The rest of this module + * (`classifyAttemptOutcome`, `isPermanentAttemptOutcome`, `canRetrySameCandidate`, + * `checkFailoverBudget`, `planNextAttempt`) is a tested library that no live request path calls + * yet: there is no live per-request cost or latency budget, and no cumulative spend check. */ import type { ProviderAttempt, @@ -100,6 +107,7 @@ export interface FailoverBudgetVerdict { /** * Whether one more attempt on `next` still fits the request budget, given what the earlier * attempts cost and how long the request has been running. Unknown estimates do not block. + * Library only: no live request path calls it (see the module comment). */ export function checkFailoverBudget(input: { attempts: readonly ProviderAttempt[]; @@ -136,6 +144,7 @@ export interface FailoverPlan { /** * Next candidate after a failed attempt: the eligible candidates in decision order (selected * first), skipping any already attempted and any whose attempt would exceed the budget. + * Library only: the live combo loops keep their own failover order (see the module comment). */ export function planNextAttempt(input: { decision: RoutingDecision; diff --git a/open-sse/services/routing/decisionStore.ts b/open-sse/services/routing/decisionStore.ts index d0f1268df66..8b35525126e 100644 --- a/open-sse/services/routing/decisionStore.ts +++ b/open-sse/services/routing/decisionStore.ts @@ -1,27 +1,95 @@ /** * Recent routing decisions, looked up by decision id or request id, so a live request can be - * explained after it ran. Bounded in memory (TTL + size cap) like the combo decision trace. + * explained after it ran. Bounded in memory by TTL, entry count and an estimated byte budget. + * + * A decision is stored in a compact form: the selected candidate and the best + * `MAX_CANDIDATES_WITH_FACTORS` keep their full factor breakdown, at most + * `MAX_STORED_CANDIDATES` candidates are kept at all, and `omittedCandidates` says how many were + * dropped. Auto combos over the whole catalog consider hundreds of candidates per request, and + * keeping every factor of every candidate for 2000 requests costs over a gigabyte of heap. * * SAFETY CONTRACT: a stored decision is the `RoutingDecision` contract only — provider/model ids, * scores, exclusion reasons, policy version. Never prompts, bodies, headers or credentials. */ -import type { RoutingDecision } from "@/shared/contracts/routing"; +import type { RoutingCandidate, RoutingDecision } from "@/shared/contracts/routing"; const DECISION_TTL_MS = 30 * 60 * 1000; const MAX_DECISIONS = 2000; +/** Candidates kept per stored decision (the selected one is always kept). */ +export const MAX_STORED_CANDIDATES = 40; +/** Candidates, in decision order, that keep their factor breakdown besides the selected one. */ +export const MAX_CANDIDATES_WITH_FACTORS = 10; +/** Estimated size budget of the whole store. */ +export const MAX_STORE_BYTES = 32 * 1024 * 1024; + +/** Rough per-object overheads used by the size estimate (V8 objects, not JSON). */ +const DECISION_BASE_BYTES = 512; +const CANDIDATE_BASE_BYTES = 320; +const FACTOR_BYTES = 120; interface StoredDecision { decision: RoutingDecision; storedAt: number; + bytes: number; } const decisions = new Map(); const decisionIdByRequestId = new Map(); +let storedBytes = 0; + +function isSameCandidate(a: RoutingCandidate, b: RoutingCandidate | undefined): boolean { + return b !== undefined && a.providerId === b.providerId && a.modelId === b.modelId; +} + +function compactCandidates(decision: RoutingDecision): RoutingCandidate[] { + const { selected, candidates } = decision; + const kept = candidates.slice(0, MAX_STORED_CANDIDATES); + if (selected && !kept.some((candidate) => isSameCandidate(candidate, selected))) { + kept[kept.length - 1] = selected; + } + return kept.map((candidate, index) => + index < MAX_CANDIDATES_WITH_FACTORS || isSameCandidate(candidate, selected) + ? candidate + : { ...candidate, factors: [] } + ); +} + +/** + * The form a decision is retained in: bounded candidates, factors only where they explain the + * choice, and the number of candidates left out. Decisions already within bounds keep their shape. + */ +function compactRoutingDecision(decision: RoutingDecision): RoutingDecision { + const candidates = compactCandidates(decision); + const omitted = + (decision.omittedCandidates ?? 0) + decision.candidates.length - candidates.length; + const unchanged = + omitted === (decision.omittedCandidates ?? 0) && + candidates.every((candidate, index) => candidate === decision.candidates[index]); + if (unchanged) return decision; + return { ...decision, candidates, ...(omitted > 0 ? { omittedCandidates: omitted } : {}) }; +} + +function candidateBytes(candidate: RoutingCandidate): number { + return ( + CANDIDATE_BASE_BYTES + + 2 * (candidate.providerId.length + candidate.modelId.length) + + candidate.factors.length * FACTOR_BYTES + + candidate.exclusionReasons.length * 32 + ); +} + +/** Estimated retained size of a stored decision, in bytes. */ +function estimateDecisionBytes(decision: RoutingDecision): number { + let bytes = DECISION_BASE_BYTES + 2 * (decision.decisionId.length + decision.requestId.length); + for (const candidate of decision.candidates) bytes += candidateBytes(candidate); + return bytes; +} function forget(decisionId: string): void { const stored = decisions.get(decisionId); if (!stored) return; decisions.delete(decisionId); + storedBytes -= stored.bytes; if (decisionIdByRequestId.get(stored.decision.requestId) === decisionId) { decisionIdByRequestId.delete(stored.decision.requestId); } @@ -34,15 +102,23 @@ function pruneExpired(now: number): void { } } -export function recordRoutingDecision(decision: RoutingDecision, now: number = Date.now()): void { - pruneExpired(now); - while (decisions.size >= MAX_DECISIONS) { +function evictOldestUntil(fits: () => boolean): void { + while (decisions.size > 0 && !fits()) { const oldest = decisions.keys().next().value; if (oldest === undefined) break; forget(oldest); } - decisions.set(decision.decisionId, { decision, storedAt: now }); - if (decision.requestId) decisionIdByRequestId.set(decision.requestId, decision.decisionId); +} + +export function recordRoutingDecision(decision: RoutingDecision, now: number = Date.now()): void { + pruneExpired(now); + const compact = compactRoutingDecision(decision); + const bytes = estimateDecisionBytes(compact); + forget(compact.decisionId); + evictOldestUntil(() => decisions.size < MAX_DECISIONS && storedBytes + bytes <= MAX_STORE_BYTES); + decisions.set(compact.decisionId, { decision: compact, storedAt: now, bytes }); + storedBytes += bytes; + if (compact.requestId) decisionIdByRequestId.set(compact.requestId, compact.decisionId); } /** The latest decision with this decision id, or for this request id. */ @@ -57,8 +133,14 @@ export function getRoutingDecision(id: string, now: number = Date.now()): Routin return stored.decision; } +/** Entry count and estimated retained bytes of the store. */ +export function getRoutingDecisionStoreStats(): { entries: number; bytes: number } { + return { entries: decisions.size, bytes: storedBytes }; +} + /** Test hook: clear the store. */ export function resetRoutingDecisionStore(): void { decisions.clear(); decisionIdByRequestId.clear(); + storedBytes = 0; } diff --git a/open-sse/services/routing/metricLabels.ts b/open-sse/services/routing/metricLabels.ts index e20ef9560ab..4624a5575d9 100644 --- a/open-sse/services/routing/metricLabels.ts +++ b/open-sse/services/routing/metricLabels.ts @@ -11,7 +11,10 @@ * an API key, connection id, account id, prompt or response. * - Every dynamic label dimension goes through a `BoundedLabelSet`: the first N * distinct values are kept, everything after that collapses into "other", so - * 10k distinct model ids can never create 10k series. + * 10k distinct model ids can never create 10k series. The fixed values + * ("redacted", "unknown", "other") never take a slot, and a caller can refuse + * new members (e.g. model ids of failed requests, which a client can invent) + * so junk values cannot crowd real ones out of the cap. */ /** Maximum length of a single label value after normalization. */ @@ -59,20 +62,28 @@ export function statusClassOf(status: number | null | undefined): string { return cls >= 1 && cls <= 5 ? `${cls}xx` : "none"; } +const FIXED_LABEL_VALUES: ReadonlySet = new Set([ + OTHER_LABEL_VALUE, + REDACTED_LABEL_VALUE, + UNKNOWN_LABEL_VALUE, +]); + /** * A label dimension with a hard cardinality cap. `resolve()` returns the * sanitized value when it is already tracked or there is room, otherwise * "other". Membership is first-come; the set never grows past `capacity`. + * Fixed values (redacted/unknown/other) pass through without taking a slot. + * With `admit = false` an untracked value is not added and reads as "other". */ export class BoundedLabelSet { private readonly values = new Set(); constructor(readonly capacity: number) {} - resolve(raw: unknown): string { + resolve(raw: unknown, admit = true): string { const value = sanitizeLabelValue(raw); - if (this.values.has(value)) return value; - if (this.values.size < this.capacity) { + if (FIXED_LABEL_VALUES.has(value) || this.values.has(value)) return value; + if (admit && this.values.size < this.capacity) { this.values.add(value); return value; } diff --git a/open-sse/services/routing/metricsSink.ts b/open-sse/services/routing/metricsSink.ts index 366f89538f3..d18394512a0 100644 --- a/open-sse/services/routing/metricsSink.ts +++ b/open-sse/services/routing/metricsSink.ts @@ -306,7 +306,9 @@ export class RoutingMetricsRegistry { const provider = this.providers.resolve(event.provider); const outcome = event.outcome; this.requests.inc([provider, outcome]); - this.modelRequests.inc([this.models.resolve(event.model), outcome]); + // Only a model that served a response takes a model label slot: a client can send any model + // id, and failed requests for invented ids must not crowd real models into "other". + this.modelRequests.inc([this.models.resolve(event.model, outcome === "success"), outcome]); this.statusClasses.inc([statusClassOf(event.status)]); this.strategyRequests.inc([this.strategies.resolve(event.strategy), outcome]); this.attempts.inc([provider]); diff --git a/src/app/(dashboard)/dashboard/analytics/RoutingDecisionLookup.tsx b/src/app/(dashboard)/dashboard/analytics/RoutingDecisionLookup.tsx index 10b033e52dd..fa8f84276e1 100644 --- a/src/app/(dashboard)/dashboard/analytics/RoutingDecisionLookup.tsx +++ b/src/app/(dashboard)/dashboard/analytics/RoutingDecisionLookup.tsx @@ -1,49 +1,116 @@ "use client"; -import { useState, type FormEvent } from "react"; -import { useTranslations } from "next-intl"; +import { useEffect, useRef, useState, type FormEvent } from "react"; +import { useLocale, useTranslations } from "next-intl"; import Badge from "@/shared/components/Badge"; import Card from "@/shared/components/Card"; -import type { RoutingCandidate, RoutingDecision } from "@/shared/contracts/routing"; +import type { + RoutingCandidate, + RoutingCircuitState, + RoutingDecision, + RoutingExclusionReason, + RoutingQuotaState, + RoutingSelectionMode, +} from "@/shared/contracts/routing"; type Translator = ((key: string, values?: Record) => string) & { has?: (key: string) => boolean; }; -type LookupStatus = "idle" | "loading" | "notFound" | "error"; +type LookupStatus = "idle" | "loading" | "found" | "notFound" | "unauthorized" | "error"; function label(t: Translator, key: string, fallback: string): string { return typeof t.has === "function" && t.has(key) ? t(key) : fallback; } +/** Translation keys for the routing contract enums; a missing translation shows the raw value. */ +const QUOTA_KEYS: Record = { + available: "routeDecisionQuotaAvailable", + low: "routeDecisionQuotaLow", + exhausted: "routeDecisionQuotaExhausted", + unknown: "routeDecisionQuotaUnknown", +}; + +const CIRCUIT_KEYS: Record = { + closed: "routeDecisionCircuitClosed", + open: "routeDecisionCircuitOpen", + half_open: "routeDecisionCircuitHalfOpen", +}; + +const SELECTION_MODE_KEYS: Record = { + deterministic: "routeDecisionModeDeterministic", + rotation: "routeDecisionModeRotation", + exploration: "routeDecisionModeExploration", +}; + +const REASON_KEYS: Record = { + not_in_candidate_pool: "routeDecisionReasonNotInCandidatePool", + model_not_found: "routeDecisionReasonModelNotFound", + capability_missing: "routeDecisionReasonCapabilityMissing", + quota_exhausted: "routeDecisionReasonQuotaExhausted", + circuit_open: "routeDecisionReasonCircuitOpen", + self_healing_excluded: "routeDecisionReasonSelfHealingExcluded", + cost_over_budget: "routeDecisionReasonCostOverBudget", + latency_over_budget: "routeDecisionReasonLatencyOverBudget", +}; + +function enumLabel(t: Translator, keys: Record, value: T): string { + const key = keys[value]; + return key ? label(t, key, value) : value; +} + function formatCost(value: number | null): string { return value === null ? "—" : `$${value.toFixed(6)}`; } -function CandidateRow({ candidate, selected }: { candidate: RoutingCandidate; selected: boolean }) { +function formatGeneratedAt(iso: string, locale: string): string { + const date = new Date(iso); + if (Number.isNaN(date.getTime())) return iso; + try { + return new Intl.DateTimeFormat(locale, { dateStyle: "medium", timeStyle: "medium" }).format( + date + ); + } catch { + return iso; + } +} + +function CandidateRow({ + candidate, + selected, + t, +}: { + candidate: RoutingCandidate; + selected: boolean; + t: Translator; +}) { return ( {candidate.providerId}/{candidate.modelId} {selected ? ( - selected + {label(t, "routeDecisionBadgeSelected", "selected")} ) : null} {candidate.score.toFixed(3)} - {candidate.eligible ? "eligible" : "excluded"} + {candidate.eligible + ? label(t, "routeDecisionEligible", "eligible") + : label(t, "routeDecisionExcluded", "excluded")} {candidate.exclusionReasons.length > 0 ? (
- {candidate.exclusionReasons.join(", ")} + {candidate.exclusionReasons + .map((reason) => enumLabel(t, REASON_KEYS, reason)) + .join(", ")}
) : null} - {candidate.quota} - {candidate.circuit} + {enumLabel(t, QUOTA_KEYS, candidate.quota)} + {enumLabel(t, CIRCUIT_KEYS, candidate.circuit)} {formatCost(candidate.estimatedCostUsd)} {candidate.estimatedLatencyMs === null ? "—" : `${candidate.estimatedLatencyMs} ms`} @@ -52,7 +119,15 @@ function CandidateRow({ candidate, selected }: { candidate: RoutingCandidate; se ); } -function DecisionDetails({ decision, t }: { decision: RoutingDecision; t: Translator }) { +function DecisionDetails({ + decision, + t, + locale, +}: { + decision: RoutingDecision; + t: Translator; + locale: string; +}) { const selectedKey = decision.selected ? `${decision.selected.providerId}/${decision.selected.modelId}` : null; @@ -64,9 +139,15 @@ function DecisionDetails({ decision, t }: { decision: RoutingDecision; t: Transl {label(t, "routeDecisionPolicyVersion", "Policy")} {decision.policyVersion} {decision.strategy ? {decision.strategy} : null} - {decision.selectionMode ? {decision.selectionMode} : null} + {decision.selectionMode ? ( + + {enumLabel(t, SELECTION_MODE_KEYS, decision.selectionMode)} + + ) : null} - {decision.liveRequestExecuted ? "live" : "preview"} + {decision.liveRequestExecuted + ? label(t, "routeDecisionLive", "live") + : label(t, "routeDecisionPreview", "preview")}

@@ -74,7 +155,9 @@ function DecisionDetails({ decision, t }: { decision: RoutingDecision; t: Transl ? `${label(t, "routeDecisionSelected", "Selected")}: ${selectedKey}` : label(t, "routeDecisionNoneSelected", "No candidate was selected.")} {" · "} - {decision.generatedAt} +

@@ -95,46 +178,124 @@ function DecisionDetails({ decision, t }: { decision: RoutingDecision; t: Transl key={`${candidate.providerId}/${candidate.modelId}`} candidate={candidate} selected={`${candidate.providerId}/${candidate.modelId}` === selectedKey} + t={t} /> ))}
+ {decision.omittedCandidates ? ( +

{omittedLabel(t, decision.omittedCandidates)}

+ ) : null} ); } -/** Look up a recent live routing decision by request id or decision id. */ -export default function RoutingDecisionLookup() { - const t = useTranslations("analytics") as Translator; - const [query, setQuery] = useState(""); +function omittedLabel(t: Translator, count: number): string { + const key = "routeDecisionOmittedCandidates"; + if (typeof t.has === "function" && t.has(key)) return t(key, { count }); + return `${count} lower-ranked candidates are not listed.`; +} + +function foundMessage(t: Translator, decision: RoutingDecision | null): string { + const selected = decision?.selected; + if (!selected) { + return label( + t, + "routeDecisionLookupFoundNoSelection", + "Decision found. No candidate was selected." + ); + } + const key = `${selected.providerId}/${selected.modelId}`; + if (typeof t.has === "function" && t.has("routeDecisionLookupFound")) { + return t("routeDecisionLookupFound", { selected: key }); + } + return `Decision found: ${key} was chosen.`; +} + +/** The text the live region announces for a lookup status. */ +function statusMessage( + t: Translator, + status: LookupStatus, + decision: RoutingDecision | null +): string | null { + switch (status) { + case "loading": + return label(t, "routeDecisionLookupLoading", "Looking up…"); + case "found": + return foundMessage(t, decision); + case "notFound": + return label( + t, + "routeDecisionLookupNotFound", + "No decision with that id in the last 30 minutes. Decisions are recorded for auto-combo routing." + ); + case "unauthorized": + return label( + t, + "routeDecisionLookupUnauthorized", + "Your session has expired or does not allow this lookup. Sign in again with a management account." + ); + case "error": + return label(t, "routeDecisionLookupFailed", "The decision could not be loaded. Try again."); + default: + return null; + } +} + +function statusForResponse(response: Response): LookupStatus | null { + if (response.status === 404) return "notFound"; + if (response.status === 401 || response.status === 403) return "unauthorized"; + return response.ok ? null : "error"; +} + +/** Fetch state of the lookup; focus moves to the results region once a decision loads. */ +function useDecisionLookup() { const [decision, setDecision] = useState(null); const [status, setStatus] = useState("idle"); + const resultsRef = useRef(null); - async function lookup(event: FormEvent) { - event.preventDefault(); - const id = query.trim(); - if (!id) return; + useEffect(() => { + if (status === "found") resultsRef.current?.focus(); + }, [status, decision]); + + async function run(id: string) { setStatus("loading"); try { const response = await fetch(`/api/omniroute/route/decisions/${encodeURIComponent(id)}`, { cache: "no-store", }); - if (response.status === 404) { + const failure = statusForResponse(response); + if (failure) { setDecision(null); - setStatus("notFound"); + setStatus(failure); return; } - if (!response.ok) throw new Error(`HTTP ${response.status}`); const body = (await response.json()) as { decision: RoutingDecision }; setDecision(body.decision); - setStatus("idle"); + setStatus("found"); } catch { setDecision(null); setStatus("error"); } } + return { decision, status, resultsRef, run }; +} + +/** Look up a recent live routing decision by request id or decision id. */ +export default function RoutingDecisionLookup() { + const t = useTranslations("analytics") as Translator; + const locale = useLocale(); + const [query, setQuery] = useState(""); + const { decision, status, resultsRef, run } = useDecisionLookup(); + + function lookup(event: FormEvent) { + event.preventDefault(); + const id = query.trim(); + if (id) void run(id); + } + return (
- {status === "loading" ? label(t, "routeDecisionLookupLoading", "Looking up…") : null} - {status === "notFound" - ? label( - t, - "routeDecisionLookupNotFound", - "No decision with that id in the last 30 minutes. Decisions are recorded for auto-combo routing." - ) - : null} - {status === "error" - ? label(t, "routeDecisionLookupFailed", "The decision could not be loaded. Try again.") - : null} + {statusMessage(t, status, decision)}
- {decision ? : null} + {decision ? ( +
+ +
+ ) : null}
); } diff --git a/src/app/api/omniroute/route/preview/route.ts b/src/app/api/omniroute/route/preview/route.ts index d3bab334cd3..78e1cdd204a 100644 --- a/src/app/api/omniroute/route/preview/route.ts +++ b/src/app/api/omniroute/route/preview/route.ts @@ -102,6 +102,26 @@ const autoRequestSchema = z.object({ type AutoPreviewBody = z.infer; +const MAX_REPORTED_ISSUES = 5; + +/** + * A readable 400 message: each problem as `field: message` (at most five), e.g. + * `Invalid route preview request: candidates: Too small: expected array to have >=1 items`. + */ +function invalidBodyMessage(error: z.ZodError): string { + const issues = error.issues.slice(0, MAX_REPORTED_ISSUES).map((issue) => { + const field = issue.path.map(String).join("."); + return field ? `${field}: ${issue.message}` : issue.message; + }); + const more = error.issues.length - issues.length; + const suffix = more > 0 ? ` (and ${more} more)` : ""; + return `Invalid route preview request: ${issues.join("; ")}${suffix}`; +} + +function badRequest(message: string): Response { + return NextResponse.json({ error: message }, { status: 400 }); +} + function previewRequestId(body: AutoPreviewBody, request: Request): string { if (body.request?.requestId) return body.request.requestId; const header = request.headers.get("x-request-id"); @@ -178,15 +198,18 @@ export async function POST(request: Request): Promise { const authError = await requireManagementAuth(request); if (authError) return authError; const raw: unknown = await request.json().catch(() => null); + if (raw === null || typeof raw !== "object" || Array.isArray(raw)) { + return badRequest("Invalid route preview request: the body must be a JSON object"); + } - if (raw !== null && typeof raw === "object" && "engine" in raw) { + if ("engine" in raw) { const auto = autoRequestSchema.safeParse(raw); - if (!auto.success) return NextResponse.json({ error: auto.error.message }, { status: 400 }); + if (!auto.success) return badRequest(invalidBodyMessage(auto.error)); return previewAuto(auto.data, request); } const parsed = requestSchema.safeParse(raw); - if (!parsed.success) return NextResponse.json({ error: parsed.error.message }, { status: 400 }); + if (!parsed.success) return badRequest(invalidBodyMessage(parsed.error)); const result = rankCandidates(parsed.data.candidates); return NextResponse.json({ diff --git a/src/i18n/messages/en.json b/src/i18n/messages/en.json index b0b6e14bb9c..3c56d06208d 100644 --- a/src/i18n/messages/en.json +++ b/src/i18n/messages/en.json @@ -2261,9 +2261,37 @@ "routeDecisionLookupLoading": "Looking up…", "routeDecisionLookupNotFound": "No decision with that id in the last 30 minutes. Decisions are recorded for auto-combo routing.", "routeDecisionLookupFailed": "The decision could not be loaded. Try again.", + "routeDecisionLookupFound": "Decision found: {selected} was chosen.", + "routeDecisionLookupFoundNoSelection": "Decision found. No candidate was selected.", + "routeDecisionLookupUnauthorized": "Your session has expired or does not allow this lookup. Sign in again with a management account.", + "routeDecisionResultsLabel": "Routing decision details", "routeDecisionPolicyVersion": "Policy", "routeDecisionSelected": "Selected", "routeDecisionNoneSelected": "No candidate was selected.", + "routeDecisionOmittedCandidates": "{count} lower-ranked candidates are not listed.", + "routeDecisionBadgeSelected": "selected", + "routeDecisionEligible": "eligible", + "routeDecisionExcluded": "excluded", + "routeDecisionLive": "live", + "routeDecisionPreview": "preview", + "routeDecisionQuotaAvailable": "available", + "routeDecisionQuotaLow": "low", + "routeDecisionQuotaExhausted": "exhausted", + "routeDecisionQuotaUnknown": "unknown", + "routeDecisionCircuitClosed": "closed", + "routeDecisionCircuitOpen": "open", + "routeDecisionCircuitHalfOpen": "half-open", + "routeDecisionModeDeterministic": "deterministic", + "routeDecisionModeRotation": "rotation", + "routeDecisionModeExploration": "exploration", + "routeDecisionReasonNotInCandidatePool": "not in candidate pool", + "routeDecisionReasonModelNotFound": "model not found", + "routeDecisionReasonCapabilityMissing": "capability missing", + "routeDecisionReasonQuotaExhausted": "quota exhausted", + "routeDecisionReasonCircuitOpen": "circuit open", + "routeDecisionReasonSelfHealingExcluded": "excluded by self-healing", + "routeDecisionReasonCostOverBudget": "over cost budget", + "routeDecisionReasonLatencyOverBudget": "over latency budget", "routeDecisionCandidate": "Candidate", "routeDecisionEligibility": "Eligibility", "routeDecisionQuota": "Quota", diff --git a/src/i18n/messages/pt-BR.json b/src/i18n/messages/pt-BR.json index ecc94fa8442..f55129b8012 100644 --- a/src/i18n/messages/pt-BR.json +++ b/src/i18n/messages/pt-BR.json @@ -2261,9 +2261,37 @@ "routeDecisionLookupLoading": "Consultando…", "routeDecisionLookupNotFound": "Nenhuma decisão com esse ID nos últimos 30 minutos. Decisões são registradas para o roteamento de combos automáticos.", "routeDecisionLookupFailed": "Não foi possível carregar a decisão. Tente novamente.", + "routeDecisionLookupFound": "Decisão encontrada: {selected} foi escolhido.", + "routeDecisionLookupFoundNoSelection": "Decisão encontrada. Nenhum candidato foi selecionado.", + "routeDecisionLookupUnauthorized": "Sua sessão expirou ou não permite esta consulta. Entre novamente com uma conta de gerenciamento.", + "routeDecisionResultsLabel": "Detalhes da decisão de roteamento", "routeDecisionPolicyVersion": "Política", "routeDecisionSelected": "Selecionado", "routeDecisionNoneSelected": "Nenhum candidato foi selecionado.", + "routeDecisionOmittedCandidates": "{count} candidatos de menor pontuação não estão listados.", + "routeDecisionBadgeSelected": "selecionado", + "routeDecisionEligible": "elegível", + "routeDecisionExcluded": "excluído", + "routeDecisionLive": "ao vivo", + "routeDecisionPreview": "prévia", + "routeDecisionQuotaAvailable": "disponível", + "routeDecisionQuotaLow": "baixa", + "routeDecisionQuotaExhausted": "esgotada", + "routeDecisionQuotaUnknown": "desconhecida", + "routeDecisionCircuitClosed": "fechado", + "routeDecisionCircuitOpen": "aberto", + "routeDecisionCircuitHalfOpen": "meio aberto", + "routeDecisionModeDeterministic": "determinístico", + "routeDecisionModeRotation": "rotação", + "routeDecisionModeExploration": "exploração", + "routeDecisionReasonNotInCandidatePool": "fora do pool de candidatos", + "routeDecisionReasonModelNotFound": "modelo não encontrado", + "routeDecisionReasonCapabilityMissing": "capacidade ausente", + "routeDecisionReasonQuotaExhausted": "cota esgotada", + "routeDecisionReasonCircuitOpen": "circuito aberto", + "routeDecisionReasonSelfHealingExcluded": "excluído pela autorrecuperação", + "routeDecisionReasonCostOverBudget": "acima do orçamento de custo", + "routeDecisionReasonLatencyOverBudget": "acima do orçamento de latência", "routeDecisionCandidate": "Candidato", "routeDecisionEligibility": "Elegibilidade", "routeDecisionQuota": "Cota", diff --git a/src/i18n/messages/vi.json b/src/i18n/messages/vi.json index f1f43a83ed0..c331ea6769b 100644 --- a/src/i18n/messages/vi.json +++ b/src/i18n/messages/vi.json @@ -2261,9 +2261,37 @@ "routeDecisionLookupLoading": "Đang tra cứu…", "routeDecisionLookupNotFound": "Không có quyết định nào với mã này trong 30 phút gần nhất. Quyết định được ghi lại cho định tuyến combo tự động.", "routeDecisionLookupFailed": "Không thể tải quyết định. Hãy thử lại.", + "routeDecisionLookupFound": "Đã tìm thấy quyết định: {selected} được chọn.", + "routeDecisionLookupFoundNoSelection": "Đã tìm thấy quyết định. Không có ứng viên nào được chọn.", + "routeDecisionLookupUnauthorized": "Phiên của bạn đã hết hạn hoặc không được phép tra cứu. Hãy đăng nhập lại bằng tài khoản quản trị.", + "routeDecisionResultsLabel": "Chi tiết quyết định định tuyến", "routeDecisionPolicyVersion": "Chính sách", "routeDecisionSelected": "Đã chọn", "routeDecisionNoneSelected": "Không có ứng viên nào được chọn.", + "routeDecisionOmittedCandidates": "{count} ứng viên xếp hạng thấp hơn không được liệt kê.", + "routeDecisionBadgeSelected": "đã chọn", + "routeDecisionEligible": "đủ điều kiện", + "routeDecisionExcluded": "bị loại", + "routeDecisionLive": "trực tiếp", + "routeDecisionPreview": "xem trước", + "routeDecisionQuotaAvailable": "khả dụng", + "routeDecisionQuotaLow": "thấp", + "routeDecisionQuotaExhausted": "đã cạn", + "routeDecisionQuotaUnknown": "không rõ", + "routeDecisionCircuitClosed": "đóng", + "routeDecisionCircuitOpen": "mở", + "routeDecisionCircuitHalfOpen": "nửa mở", + "routeDecisionModeDeterministic": "xác định", + "routeDecisionModeRotation": "luân phiên", + "routeDecisionModeExploration": "khám phá", + "routeDecisionReasonNotInCandidatePool": "không có trong nhóm ứng viên", + "routeDecisionReasonModelNotFound": "không tìm thấy mô hình", + "routeDecisionReasonCapabilityMissing": "thiếu khả năng", + "routeDecisionReasonQuotaExhausted": "hết hạn mức", + "routeDecisionReasonCircuitOpen": "mạch đang mở", + "routeDecisionReasonSelfHealingExcluded": "bị loại bởi cơ chế tự phục hồi", + "routeDecisionReasonCostOverBudget": "vượt ngân sách chi phí", + "routeDecisionReasonLatencyOverBudget": "vượt ngân sách độ trễ", "routeDecisionCandidate": "Ứng viên", "routeDecisionEligibility": "Điều kiện", "routeDecisionQuota": "Hạn mức", diff --git a/src/lib/monitoring/metricsExposition.ts b/src/lib/monitoring/metricsExposition.ts index cdd9438b6fb..8f981cd49a1 100644 --- a/src/lib/monitoring/metricsExposition.ts +++ b/src/lib/monitoring/metricsExposition.ts @@ -17,7 +17,7 @@ import { type GaugeFamily, } from "@omniroute/open-sse/services/routing/prometheusText.ts"; import { getCachedSettings } from "@/lib/db/readCache"; -import { getAllCircuitBreakerStatuses } from "@/shared/utils/circuitBreaker"; +import { getAllCircuitBreakerSnapshots } from "@/shared/utils/circuitBreaker"; import { evaluateSlo, type BreakerHistoryInput, type SloReport } from "./sloEvaluator"; import { resolveSloSettings } from "./sloSettings"; @@ -66,7 +66,7 @@ async function readQuotaMonitor(): Promise { export async function collectMetricsSnapshot(now = Date.now()): Promise { const settings = resolveSloSettings((await getCachedSettings()).slo); - const breakers = getAllCircuitBreakerStatuses(); + const breakers = getAllCircuitBreakerSnapshots(); const window = routingMetrics.window(settings.windowMinutes * 60_000); const { byState, openProviders } = summarizeBreakers(breakers); return { diff --git a/src/lib/monitoring/sloAlerts.ts b/src/lib/monitoring/sloAlerts.ts index 88012aa8193..b2b3456a6a9 100644 --- a/src/lib/monitoring/sloAlerts.ts +++ b/src/lib/monitoring/sloAlerts.ts @@ -7,6 +7,9 @@ * - `provider.circuit_open` a provider circuit breaker is newly OPEN (once per open cycle) * * `insufficient_data` keeps the previous state, so low traffic never flaps. + * Turning alerts off forgets that state, so re-enabling them never replays a + * transition that happened while they were off. Ticks never overlap: a tick + * that finds the previous one still dispatching (slow webhooks) is skipped. * Payloads carry only objective names, numbers and sanitized provider labels — * never prompts, responses, API keys, connection ids or account ids. * Off by default; enabled with settings `slo.alertsEnabled = true`. @@ -16,9 +19,10 @@ import { routingMetrics } from "@omniroute/open-sse/services/routing/metricsSink import { sanitizeLabelValue } from "@omniroute/open-sse/services/routing/metricLabels.ts"; import { getCachedSettings } from "@/lib/db/readCache"; import { dispatchEvent } from "@/lib/webhookDispatcher"; -import { getAllCircuitBreakerStatuses } from "@/shared/utils/circuitBreaker"; +import { getAllCircuitBreakerSnapshots } from "@/shared/utils/circuitBreaker"; +import type { RoutingMetricsWindow } from "@omniroute/open-sse/services/routing/metricsSink.ts"; import { evaluateSlo, type BreakerHistoryInput, type SloReport } from "./sloEvaluator"; -import { resolveSloSettings } from "./sloSettings"; +import { resolveSloSettings, type SloSettings } from "./sloSettings"; const SLO_ALERT_INTERVAL_MS = 60_000; @@ -71,6 +75,12 @@ export class SloAlertTracker { this.openProviders = nowOpen; return alerts; } + + /** Forget every remembered breach and open provider. */ + reset(): void { + this.breached.clear(); + this.openProviders = new Set(); + } } function objectivePayload( @@ -88,30 +98,66 @@ function objectivePayload( }; } -const tracker = new SloAlertTracker(); -let loopTimer: ReturnType | null = null; +export interface SloAlertRunnerDeps { + loadSettings(): Promise; + readBreakers(): BreakerHistoryInput[]; + readWindow(windowMs: number): RoutingMetricsWindow; + dispatch(event: SloAlert["event"], data: SloAlert["data"]): Promise; + now(): number; +} -async function runSloAlertTick(): Promise { - const settings = resolveSloSettings((await getCachedSettings()).slo); - if (!settings.alertsEnabled) return; - const breakers = getAllCircuitBreakerStatuses(); - const now = Date.now(); - const report = evaluateSlo( - settings, - routingMetrics.window(settings.windowMinutes * 60_000), - breakers, - now - ); - for (const alert of tracker.transitions(report, breakers)) { - await dispatchEvent(alert.event, alert.data).catch(() => undefined); +/** One SLO alert evaluation per `tick()`, never two at once. */ +export class SloAlertRunner { + private readonly tracker = new SloAlertTracker(); + private running = false; + + constructor(private readonly deps: SloAlertRunnerDeps) {} + + /** Evaluate and dispatch; returns false when skipped because a tick is still running. */ + async tick(): Promise { + if (this.running) return false; + this.running = true; + try { + await this.evaluate(); + } finally { + this.running = false; + } + return true; + } + + private async evaluate(): Promise { + const settings = await this.deps.loadSettings(); + if (!settings.alertsEnabled) { + this.tracker.reset(); + return; + } + const breakers = this.deps.readBreakers(); + const report = evaluateSlo( + settings, + this.deps.readWindow(settings.windowMinutes * 60_000), + breakers, + this.deps.now() + ); + for (const alert of this.tracker.transitions(report, breakers)) { + await this.deps.dispatch(alert.event, alert.data).catch(() => undefined); + } } } +const runner = new SloAlertRunner({ + loadSettings: async () => resolveSloSettings((await getCachedSettings()).slo), + readBreakers: getAllCircuitBreakerSnapshots, + readWindow: (windowMs) => routingMetrics.window(windowMs), + dispatch: dispatchEvent, + now: Date.now, +}); +let loopTimer: ReturnType | null = null; + /** Start the periodic SLO evaluation. Idempotent; the timer never keeps the process alive. */ export function startSloAlertLoop(): void { if (loopTimer) return; loopTimer = setInterval(() => { - runSloAlertTick().catch(() => undefined); + runner.tick().catch(() => undefined); }, SLO_ALERT_INTERVAL_MS); loopTimer.unref?.(); } diff --git a/src/lib/monitoring/sloEvaluator.ts b/src/lib/monitoring/sloEvaluator.ts index 0cc4c87de0c..edb4a0219ee 100644 --- a/src/lib/monitoring/sloEvaluator.ts +++ b/src/lib/monitoring/sloEvaluator.ts @@ -12,10 +12,18 @@ * and still succeeded ≥ failoverSuccessRateMin * - provider_recovery worst circuit OPEN→CLOSED duration (ms) * in the window, ongoing opens included ≤ providerRecoveryMaxMs + * while the breaker still fails in the window * * `failed` excludes cancelled and guardrail_blocked outcomes (client/policy * decisions, not service failures). An objective with fewer than `minSamples` * samples reports `insufficient_data` and never breaches. + * + * An OPEN/HALF_OPEN breaker only leaves that state when traffic probes it. A + * provider that stops receiving traffic (disabled, removed from combos, idle) + * would otherwise stay "recovering" forever, so an ongoing open episode counts + * only while the breaker has a failure (or an OPEN transition) inside the + * window. An idle open breaker adds no sample; it is still reported by the + * circuit-breaker gauges. */ import type { RoutingMetricsWindow } from "@omniroute/open-sse/services/routing/metricsSink.ts"; @@ -58,6 +66,8 @@ export interface BreakerHistoryInput { state: string; failureCount?: number; retryAfterMs?: number; + /** Epoch ms of the latest failure recorded by the breaker, when known. */ + lastFailureTime?: number | null; transitionHistory: ReadonlyArray<{ to: string; timestamp: number }>; } @@ -81,7 +91,18 @@ function ratio(numerator: number, denominator: number): number | null { return denominator > 0 ? numerator / denominator : null; } -/** Recovery episodes (ms) for one breaker that ended — or are still open — inside the window. */ +/** True when the breaker failed, or opened, inside the window. */ +function failedInWindow(breaker: BreakerHistoryInput, windowStartMs: number): boolean { + if ((breaker.lastFailureTime ?? 0) >= windowStartMs) return true; + return breaker.transitionHistory.some( + (transition) => transition.to === "OPEN" && transition.timestamp >= windowStartMs + ); +} + +/** + * Recovery episodes (ms) for one breaker: the ones that ended inside the window, plus the ongoing + * one while the breaker keeps failing inside the window. + */ function recoveryEpisodes( breaker: BreakerHistoryInput, windowStartMs: number, @@ -99,7 +120,9 @@ function recoveryEpisodes( } } const stillRecovering = breaker.state === "OPEN" || breaker.state === "HALF_OPEN"; - if (openedAt !== null && stillRecovering) episodes.push(Math.max(0, nowMs - openedAt)); + if (openedAt !== null && stillRecovering && failedInWindow(breaker, windowStartMs)) { + episodes.push(Math.max(0, nowMs - openedAt)); + } return episodes; } diff --git a/src/shared/contracts/routing.ts b/src/shared/contracts/routing.ts index ac3022dc76a..7af8a669b2a 100644 --- a/src/shared/contracts/routing.ts +++ b/src/shared/contracts/routing.ts @@ -35,7 +35,13 @@ export type RoutingExclusionReason = */ export type RoutingQuotaState = "available" | "low" | "exhausted" | "unknown"; -/** Optional per-request limits that the decision and every failover attempt must respect. */ +/** + * Optional per-request limits. Today they are read by route previews (candidates over a limit are + * excluded as `cost_over_budget` / `latency_over_budget`) and by the `planNextAttempt()` / + * `checkFailoverBudget()` library in `open-sse/services/routing/attemptPolicy.ts`. Live traffic has + * no per-request budget input: the live auto combo enforces only the combo's own `budgetCap` (see + * docs/routing/ROUTING_CONTRACT.md, "Guarantees"). + */ export interface RoutingBudget { /** Maximum estimated cost of the request in USD, summed over its attempts. */ maxCost?: number; @@ -94,6 +100,11 @@ export interface RoutingDecision { selectionMode?: RoutingSelectionMode; /** Router strategy that chose the candidate ("rules" is the scoring engine). */ strategy?: string; + /** + * Candidates left out of `candidates` when the decision was retained in its compact form (live + * decisions keep a bounded number of candidates). Absent when every candidate is listed. + */ + omittedCandidates?: number; } /** Result class of one provider attempt. */ diff --git a/src/shared/middleware/withRoutingRequestContext.ts b/src/shared/middleware/withRoutingRequestContext.ts index c99e704f153..312eca1c544 100644 --- a/src/shared/middleware/withRoutingRequestContext.ts +++ b/src/shared/middleware/withRoutingRequestContext.ts @@ -3,9 +3,12 @@ * * The authz pipeline stamps `x-request-id` on every forwarded request and on the response. Running * the handler inside that id's async context lets the router record its decision under the same - * id the client receives, so a request can be explained later by its request id. Non-streaming - * responses also carry the decision id and policy version; streaming responses are sent before - * routing finishes, so their decision is found through the lookup endpoint instead. + * id the client receives, so a request can be explained later by its request id. The response + * also carries the decision id and policy version whenever a decision was recorded before the + * handler returned its Response. That is the normal case for streaming responses too: the target + * is chosen before the stream starts. The headers are missing when no decision was recorded (a + * combo strategy other than `auto`, a direct model) or when the Response headers are immutable; + * the lookup endpoint by request id works in every case while the decision is retained. * * Only opaque identifiers leave through these headers: never provider credentials, prompts or * account ids. diff --git a/src/shared/utils/circuitBreaker.ts b/src/shared/utils/circuitBreaker.ts index 312b48d188a..95226e686a6 100644 --- a/src/shared/utils/circuitBreaker.ts +++ b/src/shared/utils/circuitBreaker.ts @@ -414,12 +414,25 @@ export class CircuitBreaker { getStatus(): CircuitBreakerStatus { this._refreshOpenState(); + return this._statusAs(this.state); + } + + /** + * Status without side effects: the state is `peekState()`, so an elapsed OPEN breaker reads as + * HALF_OPEN but nothing transitions, persists or consumes a probe. For metrics and SLO scrapes. + */ + peekStatus(): CircuitBreakerStatus { + return this._statusAs(this.peekState()); + } + + _statusAs(state: CircuitState): CircuitBreakerStatus { + const closed = state === STATE.CLOSED || state === STATE.DEGRADED; return { name: this.name, - state: this.state, + state, failureCount: this.failureCount, lastFailureTime: this.lastFailureTime, - retryAfterMs: this.getRetryAfterMs(), + retryAfterMs: closed ? 0 : this._timeUntilReset(), lastFailureKind: this.lastFailureKind, openCycleCount: this.openCycleCount, kindFailureCounts: { ...this.kindFailureCounts }, @@ -752,6 +765,25 @@ export function getAllCircuitBreakerStatuses() { return Array.from(registry.values()).map((cb) => cb.getStatus()); } +/** + * Read-only view of every breaker for metrics and SLO scrapes: registered breakers plus persisted + * ones not loaded in this process, via `peekStatus()`. Unlike `getAllCircuitBreakerStatuses()` it + * never transitions an elapsed OPEN breaker, writes to the database or registers a breaker. + */ +export function getAllCircuitBreakerSnapshots(): CircuitBreakerStatus[] { + const snapshots = Array.from(registry.values()).map((cb) => cb.peekStatus()); + try { + for (const persisted of loadAllCircuitBreakerStates()) { + if (!registry.has(persisted.name)) { + snapshots.push(new CircuitBreaker(persisted.name).peekStatus()); + } + } + } catch { + // Use registry only + } + return snapshots; +} + export function resetAllCircuitBreakers() { for (const cb of registry.values()) { cb.reset(); diff --git a/stryker.conf.json b/stryker.conf.json index 38366b5d44c..735b564ba25 100644 --- a/stryker.conf.json +++ b/stryker.conf.json @@ -286,6 +286,7 @@ "tests/unit/mcp-connect-scope.test.ts", "tests/unit/memory-embedding-remote.test.ts", "tests/unit/memory-embedding-transformers.test.ts", + "tests/unit/metrics-scrape-breaker-readonly.test.ts", "tests/unit/microsoft-designer-web-runtime-block.test.ts", "tests/unit/middleware-header-strip-5849.test.ts", "tests/unit/middleware-hooks-error-sanitization.test.ts", diff --git a/tests/unit/auto-explicit-strategy-budget-order.test.ts b/tests/unit/auto-explicit-strategy-budget-order.test.ts new file mode 100644 index 00000000000..6911aadd88e --- /dev/null +++ b/tests/unit/auto-explicit-strategy-budget-order.test.ts @@ -0,0 +1,137 @@ +/** + * An explicit auto router strategy (cost, latency, lkgp, ...) keeps its pick as the first target + * attempted, as in v3.8.53: the combo `budgetCap` orders the failover chain only on the scoring + * engine ("rules") path. Otherwise the recorded decision and the "Auto selection" log line would + * name a target that is not the one tried first. + */ +import { test, after } from "node:test"; +import assert from "node:assert/strict"; + +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; + +const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-explicit-budget-")); +process.env.DATA_DIR = TEST_DATA_DIR; + +const { resolveAutoStrategyOrder } = + await import("@omniroute/open-sse/services/combo/resolveAutoStrategy.ts"); +const { getRoutingDecision, resetRoutingDecisionStore } = + await import("@omniroute/open-sse/services/routing/decisionStore.ts"); +const { withRequestId } = await import("@/shared/utils/requestId.ts"); +const { resetDbInstance } = await import("@/lib/db/core.ts"); + +after(() => { + resetDbInstance(); + fs.rmSync(TEST_DATA_DIR, { recursive: true, force: true }); +}); + +const target = (provider: string, modelStr: string): never => + ({ + kind: "model", + stepId: "s1", + executionKey: `${provider}>${modelStr}`, + modelStr, + provider, + providerId: null, + connectionId: null, + weight: 1, + label: null, + }) as never; + +const candidate = (provider: string, model: string, overrides: Record = {}) => ({ + kind: "model", + stepId: "s1", + executionKey: `${provider}>${model}`, + modelStr: model, + provider, + model, + quotaRemaining: 100, + quotaTotal: 100, + circuitBreakerState: "CLOSED", + latencyStdDev: 10, + errorRate: 0, + ...overrides, +}); + +// "fast-model" is the latency pick but costs 0.05 USD per estimated request, over the 0.001 cap; +// "slow-model" is in budget. +const candidates = () => + [ + candidate("anthropic", "fast-model", { costPer1MTokens: 50, p95LatencyMs: 10 }), + candidate("openai", "slow-model", { costPer1MTokens: 0.01, p95LatencyMs: 5000 }), + ] as never; + +function capturingLog() { + const entries: string[] = []; + const push = (_tag: unknown, msg: unknown) => entries.push(String(msg)); + return { entries, info: push, warn: push, error: push, debug: push }; +} + +async function resolve(comboName: string, routerStrategy: string, requestId: string) { + const log = capturingLog(); + const result = await withRequestId( + new Request("http://localhost/v1/chat/completions", { headers: { "x-request-id": requestId } }), + () => + resolveAutoStrategyOrder({ + orderedTargets: [target("anthropic", "fast-model"), target("openai", "slow-model")], + body: { messages: [{ role: "user", content: "hi" }] }, + combo: { + id: comboName, + name: comboName, + autoConfig: { + routerStrategy, + candidatePool: ["anthropic", "openai"], + explorationRate: 0, + budgetCap: 0.001, + budgetFallback: "strict", + }, + }, + settings: null, + config: {}, + relayOptions: null, + resilienceSettings: { quotaPreflight: { enabled: false } }, + log, + buildAutoCandidates: (async () => candidates()) as never, + } as never) + ); + assert.ok("orderedTargets" in result, "expected an ordering result, not an earlyResponse"); + const selection = log.entries.find((entry) => entry.startsWith("Auto selection:")) ?? ""; + return { orderedTargets: result.orderedTargets, selection }; +} + +test("an explicit strategy's over-budget pick is still attempted first", async () => { + resetRoutingDecisionStore(); + const { orderedTargets, selection } = await resolve( + "explicit-budget-latency", + "latency", + "req-explicit-budget" + ); + assert.equal(orderedTargets[0].provider, "anthropic"); + assert.equal(orderedTargets[0].modelStr, "fast-model"); + assert.match(selection, /^Auto selection: fast-model .*strategy=latency/); + assert.equal(orderedTargets.length, 2, "the explicit path does not drop targets by budget"); + const decision = getRoutingDecision("req-explicit-budget"); + assert.equal(decision?.strategy, "latency"); + assert.equal( + decision?.selected?.modelId, + orderedTargets[0].modelStr, + "the recorded selection is the first target attempted" + ); +}); + +test("the rules path keeps its failover chain inside the budget cap", async () => { + resetRoutingDecisionStore(); + const { orderedTargets, selection } = await resolve( + "explicit-budget-rules", + "rules", + "req-rules-budget" + ); + assert.deepEqual( + orderedTargets.map((t: { modelStr: string }) => t.modelStr), + ["slow-model"], + "strict drops the over-budget target from the rules failover chain" + ); + assert.match(selection, /^Auto selection: slow-model /); + assert.equal(getRoutingDecision("req-rules-budget")?.selected?.modelId, "slow-model"); +}); diff --git a/tests/unit/auto-routing-decision-recording.test.ts b/tests/unit/auto-routing-decision-recording.test.ts index 0e1522ba8f8..39213799d9c 100644 --- a/tests/unit/auto-routing-decision-recording.test.ts +++ b/tests/unit/auto-routing-decision-recording.test.ts @@ -136,6 +136,75 @@ test("an explicit router strategy is recorded with its own pick and no live side assert.equal(getRoutingDecision("req-live-cost")?.decisionId, decision.decisionId); }); +test("an explicit strategy pick without a connection id is still reported as selected", async () => { + const candidates = [ + candidate("alpha", { connectionId: "conn-a1", costPer1MTokens: 1 }), + candidate("gamma", { connectionId: "conn-g1", costPer1MTokens: 5 }), + ]; + const decision = await inRequest("req-live-noconn", () => + recordExplicitStrategyDecision({ + config: config("live-strategy-noconn", { routerStrategy: "cost" }), + candidates, + routableCandidates: candidates, + taskType: "default", + body: { messages: [] }, + selection: { strategy: "cost", provider: "gamma", model: "gamma-model" }, + }) + ); + + assert.equal(decision.selected?.providerId, "gamma"); + assert.equal(decision.selected?.eligible, true); +}); + +test("an explicit strategy pick the explanation scoring refused is still selected", async () => { + // Every candidate is over the strict cap, so the scoring pass run for the record throws + // BudgetExceededError; explicit strategies ignore budgetCap and the request was served. + const candidates = [ + candidate("alpha", { costPer1MTokens: 20 }), + candidate("pricey", { costPer1MTokens: 50 }), + ]; + const decision = await inRequest("req-live-over-budget", () => + recordExplicitStrategyDecision({ + config: config("live-strategy-budget", { + routerStrategy: "latency", + budgetCap: 0.001, + budgetFallback: "strict", + }), + candidates, + routableCandidates: candidates, + taskType: "default", + body: { messages: [] }, + selection: { strategy: "latency", provider: "pricey", model: "pricey-model" }, + }) + ); + + assert.equal(decision.selected?.providerId, "pricey"); + assert.equal(decision.selected?.eligible, true); + assert.deepEqual(decision.selected?.exclusionReasons, []); + assert.equal( + decision.candidates.find((c) => c.providerId === "pricey")?.eligible, + true, + "the listed candidate agrees with the selection" + ); +}); + +test("an explicit strategy pick blocked by a quota cutoff is not reported as selected", async () => { + const blocked = candidate("beta", { quotaCutoffBlocked: true }); + const alpha = candidate("alpha"); + const decision = await inRequest("req-live-blocked-pick", () => + recordExplicitStrategyDecision({ + config: config("live-strategy-blocked", { routerStrategy: "cost" }), + candidates: [alpha, blocked], + routableCandidates: [alpha], + taskType: "default", + body: { messages: [] }, + selection: { strategy: "cost", provider: "beta", model: "beta-model" }, + }) + ); + + assert.equal(decision.selected, undefined); +}); + test("failover order respects the cost budget", () => { const targets = [{ executionKey: "a" }, { executionKey: "b" }, { executionKey: "c" }]; const candidates = [ diff --git a/tests/unit/metrics-scrape-breaker-readonly.test.ts b/tests/unit/metrics-scrape-breaker-readonly.test.ts new file mode 100644 index 00000000000..46c247bb59e --- /dev/null +++ b/tests/unit/metrics-scrape-breaker-readonly.test.ts @@ -0,0 +1,60 @@ +/** + * Scraping metrics must not change routing state: `/api/metrics` and the SLO alert timer read + * circuit breakers through `getAllCircuitBreakerSnapshots()` (`peekStatus()`), so an OPEN breaker + * whose cooldown elapsed is reported as HALF_OPEN without transitioning, persisting or consuming + * its half-open probe. + */ +import test from "node:test"; +import assert from "node:assert/strict"; +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; + +const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-scrape-breaker-")); +process.env.DATA_DIR = TEST_DATA_DIR; + +const core = await import("../../src/lib/db/core.ts"); +const breakers = await import("../../src/shared/utils/circuitBreaker.ts"); +const { collectMetricsSnapshot } = await import("../../src/lib/monitoring/metricsExposition.ts"); + +test.after(() => { + breakers.resetAllCircuitBreakers(); + core.resetDbInstance(); + fs.rmSync(TEST_DATA_DIR, { recursive: true, force: true, maxRetries: 5, retryDelay: 100 }); +}); + +test("scraping does not transition an OPEN breaker whose cooldown elapsed", async () => { + let clock = Date.now(); + const breaker = breakers.getCircuitBreaker("scrape-readonly-provider", { + failureThreshold: 1, + resetTimeout: 1000, + now: () => clock, + }); + await assert.rejects( + breaker.execute(async () => { + throw new Error("upstream down"); + }) + ); + assert.equal(breaker.state, "OPEN"); + clock += 5000; + const historyBefore = breaker.transitionHistory.length; + + const first = await collectMetricsSnapshot(); + const second = await collectMetricsSnapshot(); + + assert.equal(breaker.state, "OPEN", "the scrape did not move the breaker to HALF_OPEN"); + assert.equal(breaker.transitionHistory.length, historyBefore, "no transition was recorded"); + assert.equal(breaker.halfOpenAllowed, 0, "no half-open probe was granted by the scrape"); + for (const snapshot of [first, second]) { + assert.ok(snapshot.breakersByState.HALF_OPEN >= 1, "the effective state is reported"); + assert.equal(snapshot.openProviders.includes("scrape-readonly-provider"), false); + } + + const snapshot = breakers + .getAllCircuitBreakerSnapshots() + .find((status) => status.name === "scrape-readonly-provider"); + assert.equal(snapshot?.state, "HALF_OPEN"); + assert.equal(snapshot?.retryAfterMs, 0); + assert.equal(breaker.canExecute(), true, "live routing still gets its probe afterwards"); + assert.equal(breaker.state, "HALF_OPEN"); +}); diff --git a/tests/unit/omniroute-route-preview.test.ts b/tests/unit/omniroute-route-preview.test.ts index 2b2f3368c7c..81de68a4a1c 100644 --- a/tests/unit/omniroute-route-preview.test.ts +++ b/tests/unit/omniroute-route-preview.test.ts @@ -77,3 +77,42 @@ test("an invalid preview body is still rejected with 400", async () => { ); assert.equal(res.status, 400); }); + +async function previewError(body: string) { + const res = await POST( + new Request("http://localhost/api/omniroute/route/preview", { + method: "POST", + headers: { "content-type": "application/json" }, + body, + }) + ); + assert.equal(res.status, 400); + const json = (await res.json()) as { error?: unknown }; + assert.equal(typeof json.error, "string"); + return json.error as string; +} + +test("a 400 names the invalid fields in a readable string, not a JSON dump", async () => { + const missing = await previewError("{}"); + assert.match(missing, /^Invalid route preview request: candidates: /); + assert.equal(missing.includes("\n"), false); + assert.equal(missing.includes('"path"'), false); + assert.throws(() => JSON.parse(missing), SyntaxError, "not a serialized issue list"); + + const auto = await previewError( + JSON.stringify({ engine: "auto", candidates: [{ provider: "alpha", model: "m" }] }) + ); + assert.match(auto, /candidates\.0\.costPer1MTokens: /); + assert.match(auto, /candidates\.0\.p95LatencyMs: /); +}); + +test("a body that is not a JSON object gets a clear 400", async () => { + assert.equal( + await previewError("not json"), + "Invalid route preview request: the body must be a JSON object" + ); + assert.equal( + await previewError("[1, 2]"), + "Invalid route preview request: the body must be a JSON object" + ); +}); diff --git a/tests/unit/routing-decision-store.test.ts b/tests/unit/routing-decision-store.test.ts index 79286f34c17..f0dcd73b9e6 100644 --- a/tests/unit/routing-decision-store.test.ts +++ b/tests/unit/routing-decision-store.test.ts @@ -3,6 +3,10 @@ import assert from "node:assert/strict"; import { getRoutingDecision, + getRoutingDecisionStoreStats, + MAX_CANDIDATES_WITH_FACTORS, + MAX_STORE_BYTES, + MAX_STORED_CANDIDATES, recordRoutingDecision, resetRoutingDecisionStore, } from "../../open-sse/services/routing/decisionStore.ts"; @@ -47,3 +51,93 @@ test("the store keeps at most 2000 decisions, evicting the oldest", () => { assert.equal(getRoutingDecision("req-0", 1000), null); assert.ok(getRoutingDecision("rd_2000", 1000)); }); + +const FACTOR_NAMES = [ + "quota", + "health", + "cost", + "latency", + "task_fit", + "stability", + "tier", + "cache_affinity", + "context", + "throughput", + "freshness", + "region", + "priority", + "reliability", + "error_rate", + "jitter", + "session", +]; + +function wideCandidate(index: number) { + return { + providerId: `provider-${index % 40}`, + modelId: `catalog-model-${index}`, + score: 1 - index / 1000, + factors: FACTOR_NAMES.map((name) => ({ name, value: 0.5, weight: 0.1, contribution: 0.05 })), + eligible: index % 3 !== 0, + exclusionReasons: index % 3 === 0 ? ["circuit_open" as const] : [], + quota: "available" as const, + circuit: "closed" as const, + estimatedCostUsd: 0.001, + estimatedLatencyMs: 400, + }; +} + +function wideDecision(decisionId: string, requestId: string, candidateCount: number) { + const candidates = Array.from({ length: candidateCount }, (_, i) => wideCandidate(i)); + return { ...decision(decisionId, requestId), candidates, selected: candidates[250] }; +} + +test("a wide decision is stored compact: bounded candidates, factors only where they explain", () => { + recordRoutingDecision(wideDecision("rd_wide", "req-wide", 300), 1000); + const stored = getRoutingDecision("req-wide", 1000); + assert.ok(stored); + assert.equal(stored.candidates.length, MAX_STORED_CANDIDATES); + assert.equal(stored.omittedCandidates, 300 - MAX_STORED_CANDIDATES); + assert.equal(stored.selected?.modelId, "catalog-model-250"); + assert.equal(stored.selected?.factors.length, FACTOR_NAMES.length); + assert.ok( + stored.candidates.some((c) => c.modelId === "catalog-model-250"), + "the selected candidate stays listed even when it ranked below the kept ones" + ); + const withFactors = stored.candidates.filter((c) => c.factors.length > 0); + assert.ok(withFactors.length <= MAX_CANDIDATES_WITH_FACTORS + 1, `${withFactors.length}`); + assert.equal(stored.candidates[0].factors.length, FACTOR_NAMES.length); + assert.deepEqual(stored.candidates[MAX_CANDIDATES_WITH_FACTORS].factors, []); + assert.deepEqual( + stored.candidates[0].exclusionReasons, + ["circuit_open"], + "exclusion reasons survive compaction" + ); +}); + +test("2000 decisions with 300 candidates each stay inside the store byte budget", () => { + for (let i = 0; i < 2000; i += 1) { + recordRoutingDecision(wideDecision(`rd_w${i}`, `req-w${i}`, 300), 1000); + } + const stats = getRoutingDecisionStoreStats(); + assert.ok(stats.bytes <= MAX_STORE_BYTES, `estimated ${stats.bytes} bytes`); + assert.ok(stats.entries > 0 && stats.entries <= 2000); + const latest = getRoutingDecision("req-w1999", 1000); + assert.ok(latest, "the newest decision is retained"); + const serialized = JSON.stringify(latest).length; + assert.ok(serialized < 64 * 1024, `one stored decision serializes to ${serialized} bytes`); + assert.ok( + serialized * stats.entries < 2 * MAX_STORE_BYTES, + `retained decisions serialize to ${serialized * stats.entries} bytes` + ); + assert.equal(getRoutingDecision("req-w0", 1000), null, "the oldest were evicted by the budget"); +}); + +test("a small decision is stored unchanged", () => { + const small = wideDecision("rd_small", "req-small", 5); + recordRoutingDecision(small, 1000); + const stored = getRoutingDecision("rd_small", 1000); + assert.equal(stored?.omittedCandidates, undefined); + assert.equal(stored?.candidates.length, 5); + assert.ok(stored?.candidates.every((c) => c.factors.length === FACTOR_NAMES.length)); +}); diff --git a/tests/unit/routing-metrics-sink.test.ts b/tests/unit/routing-metrics-sink.test.ts index 39955f35297..b28e1d608e0 100644 --- a/tests/unit/routing-metrics-sink.test.ts +++ b/tests/unit/routing-metrics-sink.test.ts @@ -157,6 +157,43 @@ test("BoundedLabelSet keeps first N values and collapses the rest", () => { assert.equal(set.size, 2); }); +test("junk model ids from failed requests do not crowd real models out of the label cap", () => { + const registry = new RoutingMetricsRegistry(); + for (let i = 0; i < 500; i++) { + registry.record(event({ model: `junk-${i}`, outcome: "error", status: 404 })); + } + for (let i = 0; i < 500; i++) { + registry.record(event({ model: `sk-proj-${"x".repeat(12)}${i}` })); + } + registry.record(event({ model: "claude-opus-4-7" })); + registry.record(event({ model: "claude-opus-4-7", outcome: "error", status: 500 })); + + const models = new Set( + familySeries(registry, "omniroute_model_requests_total").map((s) => s.labels.model) + ); + assert.ok(models.has("claude-opus-4-7"), "a real model keeps its own label"); + assert.ok(models.has(OTHER_LABEL_VALUE), "failed junk ids collapse into other"); + assert.ok(models.has("redacted")); + assert.equal( + [...models].some((model) => model.startsWith("junk-")), + false, + "a failed request never admits a new model label" + ); + assert.equal(registry.totals().requests, 1002); +}); + +test("BoundedLabelSet: fixed values take no slot and admit=false never adds", () => { + const set = new BoundedLabelSet(1); + assert.equal(set.resolve("sk-proj-abcdefghijklmnop"), "redacted"); + assert.equal(set.resolve(""), "unknown"); + assert.equal(set.size, 0); + assert.equal(set.resolve("junk", false), OTHER_LABEL_VALUE); + assert.equal(set.size, 0); + assert.equal(set.resolve("real"), "real"); + assert.equal(set.resolve("real", false), "real", "a tracked value still resolves"); + assert.equal(set.resolve("sk-proj-abcdefghijklmnop"), "redacted", "fixed values bypass the cap"); +}); + test("histogram quantiles interpolate inside the matching bucket", () => { const hist = new FixedHistogram([100, 200, 400]); for (let i = 0; i < 90; i++) hist.observe(50); diff --git a/tests/unit/slo-evaluator.test.ts b/tests/unit/slo-evaluator.test.ts index 125534bc644..b79324db8df 100644 --- a/tests/unit/slo-evaluator.test.ts +++ b/tests/unit/slo-evaluator.test.ts @@ -16,7 +16,7 @@ import { evaluateSlo, type BreakerHistoryInput, } from "../../src/lib/monitoring/sloEvaluator.ts"; -import { SloAlertTracker } from "../../src/lib/monitoring/sloAlerts.ts"; +import { SloAlertRunner, SloAlertTracker } from "../../src/lib/monitoring/sloAlerts.ts"; import { sloSettingsSchema } from "../../src/shared/validation/schemas/slo.ts"; import { updateSettingsSchema } from "../../src/shared/validation/settingsSchemas.ts"; import { WEBHOOK_EVENT_VALUES } from "../../src/lib/webhooks/eventDescriptions.ts"; @@ -179,6 +179,125 @@ test("provider recovery time comes from OPEN→CLOSED transitions, including ong assert.equal(report.provider_recovery.provider, "openai"); }); +test("an idle breaker left open or half-open does not breach provider recovery forever", () => { + const windowStart = NOW - 15 * 60_000; + const idleHalfOpen: BreakerHistoryInput = { + name: "retired-provider", + state: "HALF_OPEN", + lastFailureTime: NOW - 2 * 60 * 60_000, + transitionHistory: [ + { to: "OPEN", timestamp: NOW - 2 * 60 * 60_000 }, + { to: "HALF_OPEN", timestamp: NOW - 2 * 60 * 60_000 + 30_000 }, + ], + }; + const idleOpen: BreakerHistoryInput = { + name: "idle-open", + state: "OPEN", + lastFailureTime: NOW - 60 * 60_000, + transitionHistory: [{ to: "OPEN", timestamp: NOW - 60 * 60_000 }], + }; + const idle = computeProviderRecovery([idleHalfOpen, idleOpen], windowStart, NOW); + assert.equal(idle.samples, 0); + assert.equal(idle.worstMs, null); + + const report = evaluateSlo( + resolveSloSettings({ providerRecoveryMaxMs: 120_000 }), + windowWith(0, 0), + [idleHalfOpen, idleOpen], + NOW + ); + assert.equal(byKey(report).provider_recovery.status, "insufficient_data"); + assert.notEqual(report.status, "breached"); + + // Still failing inside the window: the whole ongoing episode counts and breaches. + const stillFailing: BreakerHistoryInput = { + ...idleHalfOpen, + name: "still-failing", + lastFailureTime: NOW - 60_000, + }; + const failing = computeProviderRecovery([stillFailing], windowStart, NOW); + assert.equal(failing.samples, 1); + assert.equal(failing.worstMs, 2 * 60 * 60_000); + assert.equal(failing.provider, "still-failing"); +}); + +function runnerDeps(overrides: Partial[0]> = {}) { + const dispatched: string[] = []; + let enabled = true; + let window = windowWith(50, 50); + const deps = { + loadSettings: async () => resolveSloSettings({ minSamples: 10, alertsEnabled: enabled }), + readBreakers: () => [], + readWindow: () => window, + dispatch: async (event: string) => { + dispatched.push(event); + }, + now: () => NOW, + ...overrides, + }; + return { + deps, + dispatched, + setEnabled: (value: boolean) => (enabled = value), + setWindow: (value: ReturnType) => (window = value), + }; +} + +test("SloAlertRunner forgets alert state while alerts are off", async () => { + const harness = runnerDeps(); + const runner = new SloAlertRunner(harness.deps); + + await runner.tick(); + assert.ok(harness.dispatched.includes("slo.breached")); + harness.dispatched.length = 0; + + harness.setEnabled(false); + harness.setWindow(windowWith(100, 0)); + await runner.tick(); + assert.deepEqual(harness.dispatched, [], "nothing is sent while alerts are off"); + + harness.setEnabled(true); + await runner.tick(); + assert.deepEqual( + harness.dispatched, + [], + "re-enabling does not replay a recovery that happened while alerts were off" + ); + + harness.setWindow(windowWith(50, 50)); + await runner.tick(); + assert.ok(harness.dispatched.includes("slo.breached"), "a new breach alerts again"); +}); + +test("SloAlertRunner skips a tick while the previous one is still dispatching", async () => { + let release: () => void = () => undefined; + const gate = new Promise((resolve) => { + release = resolve; + }); + const events: string[] = []; + const harness = runnerDeps({ + dispatch: async (event: string) => { + events.push(event); + await gate; + }, + }); + const runner = new SloAlertRunner(harness.deps); + + const first = runner.tick(); + await new Promise((resolve) => setImmediate(resolve)); + assert.equal(await runner.tick(), false, "the overlapping tick is skipped"); + release(); + assert.equal(await first, true); + const breaches = events.filter((event) => event === "slo.breached").length; + assert.ok(breaches > 0); + assert.equal(await runner.tick(), true, "the next tick runs again"); + assert.equal( + events.filter((event) => event === "slo.breached").length, + breaches, + "each transition is sent once" + ); +}); + test("SloAlertTracker emits breach/recovery/circuit-open once per state change", () => { const settings = resolveSloSettings({ minSamples: 10 }); const tracker = new SloAlertTracker(); diff --git a/tests/unit/ui/routing-decision-lookup.test.tsx b/tests/unit/ui/routing-decision-lookup.test.tsx index dab3e6d3d1c..aab3f46110d 100644 --- a/tests/unit/ui/routing-decision-lookup.test.tsx +++ b/tests/unit/ui/routing-decision-lookup.test.tsx @@ -2,14 +2,24 @@ import { afterEach, describe, expect, it, vi } from "vitest"; import { cleanup, fireEvent, render, screen, waitFor } from "@testing-library/react"; +const intl = vi.hoisted(() => ({ + locale: "en", + messages: {} as Record, +})); + vi.mock("next-intl", () => ({ + useLocale: () => intl.locale, useTranslations: () => { - const t = (key: string) => key; - return Object.assign(t, { has: () => false }); + const t = (key: string, values?: Record) => + (intl.messages[key] ?? key).replace(/\{(\w+)\}/g, (_, name: string) => + String(values?.[name] ?? "") + ); + return Object.assign(t, { has: (key: string) => key in intl.messages }); }, })); import RoutingDecisionLookup from "../../../src/app/(dashboard)/dashboard/analytics/RoutingDecisionLookup"; +import ptBR from "../../../src/i18n/messages/pt-BR.json"; const decision = { decisionId: "rd_11111111-2222-3333-4444-555555555555", @@ -62,8 +72,17 @@ const decision = { afterEach(() => { cleanup(); vi.unstubAllGlobals(); + intl.locale = "en"; + intl.messages = {}; }); +async function lookUp(value = "req-ui-1") { + fireEvent.change(screen.getByLabelText(/Request id or decision id|ID da requisição/i), { + target: { value }, + }); + fireEvent.click(screen.getByRole("button", { name: /look up|consultar|buscar/i })); +} + describe("RoutingDecisionLookup", () => { it("looks up a decision by id and shows candidates with exclusion reasons", async () => { const fetchMock = vi.fn(async () => Response.json({ decision })); @@ -85,6 +104,86 @@ describe("RoutingDecisionLookup", () => { expect(screen.getByRole("table")).toBeTruthy(); }); + it("renders badges, enums and the timestamp in the active locale", async () => { + intl.locale = "pt-BR"; + intl.messages = (ptBR as { analytics: Record }).analytics; + vi.stubGlobal( + "fetch", + vi.fn(async () => Response.json({ decision })) + ); + render(); + await lookUp(); + + await waitFor(() => expect(screen.getByRole("table")).toBeTruthy()); + for (const english of ["selected", "eligible", "excluded", "live", "deterministic"]) { + expect(screen.queryByText(english)).toBeNull(); + } + expect(screen.getByText("selecionado")).toBeTruthy(); + expect(screen.getByText("excluído")).toBeTruthy(); + expect(screen.getByText("circuito aberto")).toBeTruthy(); + expect(screen.getByText("ao vivo")).toBeTruthy(); + const time = document.querySelector("time"); + expect(time?.getAttribute("dateTime")).toBe(decision.generatedAt); + expect(time?.textContent).not.toBe(decision.generatedAt); + }); + + it("announces a found decision and moves focus to it", async () => { + vi.stubGlobal( + "fetch", + vi.fn(async () => Response.json({ decision })) + ); + render(); + await lookUp(); + + await waitFor(() => + expect(screen.getByText("Decision found: alpha/alpha-model was chosen.")).toBeTruthy() + ); + const results = screen.getByLabelText("Routing decision details"); + expect(document.activeElement).toBe(results); + }); + + for (const status of [401, 403]) { + it(`tells the user to sign in again on ${status}`, async () => { + vi.stubGlobal( + "fetch", + vi.fn(async () => Response.json({ error: "x" }, { status })) + ); + render(); + await lookUp(); + + await waitFor(() => expect(screen.getByText(/Sign in again/)).toBeTruthy()); + expect(screen.queryByText(/Try again/)).toBeNull(); + expect(screen.queryByRole("table")).toBeNull(); + }); + } + + it("keeps the generic retry message for server errors", async () => { + vi.stubGlobal( + "fetch", + vi.fn(async () => Response.json({ error: "x" }, { status: 500 })) + ); + render(); + await lookUp(); + + await waitFor(() => expect(screen.getByText(/Try again/)).toBeTruthy()); + }); + + it("says how many candidates a compact decision left out", async () => { + vi.stubGlobal( + "fetch", + vi.fn(async () => Response.json({ decision: { ...decision, omittedCandidates: 260 } })) + ); + render(); + + fireEvent.change(screen.getByLabelText("Request id or decision id"), { + target: { value: "req-ui-1" }, + }); + fireEvent.click(screen.getByRole("button", { name: /look up/i })); + + await waitFor(() => expect(screen.getByRole("table")).toBeTruthy()); + expect(screen.getByText("260 lower-ranked candidates are not listed.")).toBeTruthy(); + }); + it("announces a not-found id without showing stale details", async () => { vi.stubGlobal( "fetch",