Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 18 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -792,10 +792,10 @@ When you no longer need OmniRoute, we provide two quick scripts for a clean remo

For most deployments, you only need:

| Variable | Default | Purpose |
| ------------------------ | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| `REQUEST_TIMEOUT_MS` | `600000` | Shared baseline for upstream fetch, hidden Undici timeouts, TLS fingerprint requests, and API bridge request/proxy timeouts |
| `STREAM_IDLE_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` | Maximum gap between streaming chunks before OmniRoute aborts the SSE stream |
| Variable | Default | Purpose |
| ------------------------ | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `REQUEST_TIMEOUT_MS` | `600000` | Shared baseline for upstream response-start timeout, hidden Undici timeouts, TLS fingerprint requests, and API bridge request/proxy timeouts |
| `STREAM_IDLE_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` | Maximum gap between streaming chunks before OmniRoute aborts the SSE stream |

Backward compatibility is preserved: existing `FETCH_TIMEOUT_MS`, `API_BRIDGE_PROXY_TIMEOUT_MS`, and other per-layer timeout vars still work and override the shared baseline.

Expand All @@ -808,20 +808,22 @@ only forwards client-provided `cache_control` markers. If the request does not i

Advanced overrides are available if you need finer control:

| Variable | Default | Purpose |
| ---------------------------------------- | ------------------------------------------ | -------------------------------------------------------------------- |
| `FETCH_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` | Total upstream request timeout used by the main fetch abort signal |
| `FETCH_HEADERS_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Undici time limit for receiving upstream response headers |
| `FETCH_BODY_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Undici time limit between upstream body chunks (`0` disables it) |
| `FETCH_CONNECT_TIMEOUT_MS` | `30000` | Undici TCP connect timeout |
| `FETCH_KEEPALIVE_TIMEOUT_MS` | `4000` | Undici idle keep-alive socket timeout |
| `TLS_CLIENT_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Timeout for TLS fingerprint requests made through `wreq-js` |
| `API_BRIDGE_PROXY_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` or `30000` | Timeout for `/v1` proxy forwarding from API port to dashboard port |
| `API_BRIDGE_SERVER_REQUEST_TIMEOUT_MS` | `max(API_BRIDGE_PROXY_TIMEOUT_MS, 300000)` | Incoming request timeout on the API bridge server |
| `API_BRIDGE_SERVER_HEADERS_TIMEOUT_MS` | `60000` | Incoming header timeout on the API bridge server |
| `API_BRIDGE_SERVER_KEEPALIVE_TIMEOUT_MS` | `5000` | Keep-alive timeout on the API bridge server |
| Variable | Default | Purpose |
| -------------------------------------- | ------------------------------------------ | ------------------------------------------------------------------ |
| `FETCH_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` | Upstream response-start timeout used until response headers arrive |
| `FETCH_HEADERS_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Undici time limit for receiving upstream response headers |
| `FETCH_BODY_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Undici time limit between upstream body chunks (`0` disables it) |
| `FETCH_CONNECT_TIMEOUT_MS` | `30000` | Undici TCP connect timeout |
| `FETCH_KEEPALIVE_TIMEOUT_MS` | `4000` | Undici idle keep-alive socket timeout |
| `TLS_CLIENT_TIMEOUT_MS` | inherits `FETCH_TIMEOUT_MS` | Timeout for TLS fingerprint requests made through `wreq-js` |
| `API_BRIDGE_PROXY_TIMEOUT_MS` | inherits `REQUEST_TIMEOUT_MS` or `30000` | Timeout for `/v1` proxy forwarding from API port to dashboard port |
| `API_BRIDGE_SERVER_REQUEST_TIMEOUT_MS` | `max(API_BRIDGE_PROXY_TIMEOUT_MS, 300000)` | Incoming request timeout on the API bridge server |
| `API_BRIDGE_SERVER_HEADERS_TIMEOUT_MS` | `60000` | Incoming header timeout on the API bridge server |
| `API_BRIDGE_SERVER_KEEPALIVE_TIMEOUT_MS` | `5000` | Keep-alive timeout on the API bridge server |
| `API_BRIDGE_SERVER_SOCKET_TIMEOUT_MS` | `0` | Socket inactivity timeout on the API bridge server (`0` disables it) |

For streaming requests, `FETCH_TIMEOUT_MS` only covers connection setup / waiting for the first upstream response. Once the stream is active, OmniRoute will only abort on an actual stall (`STREAM_IDLE_TIMEOUT_MS`) or Undici body inactivity (`FETCH_BODY_TIMEOUT_MS`).

If you run OmniRoute behind Nginx, Caddy, Cloudflare, or another reverse proxy, make sure the proxy
timeouts are also higher than your OmniRoute stream/fetch timeouts.

Expand Down
4 changes: 3 additions & 1 deletion open-sse/config/constants.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,9 @@ const upstreamTimeouts = getUpstreamTimeoutConfig(process.env, (message) => {
console.warn(`[open-sse] ${message}`);
});

// Timeout for non-streaming fetch requests (ms). Prevents stalled connections.
// Timeout for receiving the initial upstream response (ms).
// After headers arrive, active SSE streams are governed by STREAM_IDLE_TIMEOUT_MS
// and Undici's bodyTimeout instead of this one-shot startup timer.
export const FETCH_TIMEOUT_MS = upstreamTimeouts.fetchTimeoutMs;

// Idle timeout for SSE streams (ms). Closes stream if no data for this duration.
Expand Down
61 changes: 46 additions & 15 deletions open-sse/executors/base.ts
Original file line number Diff line number Diff line change
Expand Up @@ -122,19 +122,23 @@ export function applyConfiguredUserAgent(
export function mergeAbortSignals(primary: AbortSignal, secondary: AbortSignal): AbortSignal {
const controller = new AbortController();

const abortBoth = () => {
const abortFrom = (source: AbortSignal) => {
if (!controller.signal.aborted) {
controller.abort();
controller.abort(source.reason);
}
};

if (primary.aborted || secondary.aborted) {
abortBoth();
if (primary.aborted) {
abortFrom(primary);
return controller.signal;
}
if (secondary.aborted) {
abortFrom(secondary);
return controller.signal;
}

primary.addEventListener("abort", abortBoth, { once: true });
secondary.addEventListener("abort", abortBoth, { once: true });
primary.addEventListener("abort", () => abortFrom(primary), { once: true });
secondary.addEventListener("abort", () => abortFrom(secondary), { once: true });
Comment on lines +140 to +141

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This implementation of mergeAbortSignals adds listeners to the primary signal (usually the client request signal) but never removes them if the operation completes without an abort. In scenarios with multiple fallbacks or retries, each attempt will add a new listener to the same primary signal. This can lead to a MaxListenersExceededWarning in Node.js (default limit is 10) and a small memory leak for the duration of the request. Since Node 18 is supported, AbortSignal.any is not available, but you should consider a mechanism to remove these listeners once the fetch call or the resulting stream is finished.

return controller.signal;
}

Expand Down Expand Up @@ -252,6 +256,9 @@ export class BaseExecutor {

// Intra-URL retry config: retry same URL before falling back to next node
static readonly RETRY_CONFIG = { maxAttempts: 2, delayMs: 2000 };
// Timeout for receiving the initial upstream response headers. Once the response
// starts streaming, STREAM_IDLE_TIMEOUT_MS / Undici bodyTimeout handle stalls.
static FETCH_START_TIMEOUT_MS = FETCH_TIMEOUT_MS;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The use of a static property for FETCH_START_TIMEOUT_MS is problematic for two reasons:

  1. Concurrency Risk: In an asynchronous environment, concurrent requests attempting to set different timeouts would overwrite this shared value, leading to race conditions.
  2. Ignored Settings: The execute method (line 326) relies on this static value, which defaults to the global FETCH_TIMEOUT_MS. It does not appear to receive or respect the per-combo timeoutMs setting that was increased to 20 minutes in this PR's schema and UI changes.

To fix this, timeoutMs should be added to the ExecuteInput type and passed through from the combo configuration to the execute method.


// Override in subclass for provider-specific refresh
async refreshCredentials(credentials: ProviderCredentials, log: ExecutorLog | null) {
Expand Down Expand Up @@ -404,13 +411,26 @@ export class BaseExecutor {
const transformedBody = await this.transformRequest(model, body, stream, activeCredentials);

try {
// Apply timeout to all requests. Non-streaming requests need this to prevent
// stalled connections. Streaming requests also need it for the initial fetch() call
// to prevent hanging on unresponsive providers (e.g. 300s TCP default timeout — #769).
// Stream idle detection (STREAM_IDLE_TIMEOUT_MS) handles stalls after data starts flowing.
const timeoutMs = this.getTimeoutMs();
const timeoutSignal = AbortSignal.timeout(timeoutMs);
const combinedSignal = signal ? mergeAbortSignals(signal, timeoutSignal) : timeoutSignal;
// Only enforce the timeout while waiting for the initial fetch() response.
// Once headers arrive, active streams must not be cut off by total elapsed time;
// post-start stalls are handled separately by STREAM_IDLE_TIMEOUT_MS / bodyTimeout.
const fetchStartTimeoutMs = this.getTimeoutMs();
const timeoutController = fetchStartTimeoutMs > 0 ? new AbortController() : null;
let timeoutId: ReturnType<typeof setTimeout> | null = null;
if (timeoutController) {
timeoutId = setTimeout(() => {
const timeoutError = new Error(
`Fetch timeout after ${fetchStartTimeoutMs}ms on ${url}`
);
timeoutError.name = "TimeoutError";
timeoutController.abort(timeoutError);
}, fetchStartTimeoutMs);

Copilot AI Apr 14, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The startup timeout uses setTimeout(...) but the returned timer is not unref()'d. With the default (600s), an in-flight request timer can keep the Node process alive longer than expected during shutdown. Consider calling timeoutId.unref?.() after creating the timer (similar to open-sse/services/quotaMonitor.ts:244) or using an approach that doesn't keep the event loop alive.

Suggested change
}, fetchStartTimeoutMs);
}, fetchStartTimeoutMs);
timeoutId.unref?.();

Copilot uses AI. Check for mistakes.
}
const timeoutSignal = timeoutController?.signal ?? null;
const combinedSignal =
signal && timeoutSignal
? mergeAbortSignals(signal, timeoutSignal)
: signal || timeoutSignal;

// Apply CLI fingerprint ordering if enabled for this provider
let finalHeaders = headers;
Expand Down Expand Up @@ -438,7 +458,15 @@ export class BaseExecutor {
};
if (combinedSignal) fetchOptions.signal = combinedSignal;

const response = await fetch(url, fetchOptions);
let response;
try {
response = await fetch(url, fetchOptions);
} finally {
if (timeoutId) {
clearTimeout(timeoutId);
timeoutId = null;
}
}

// Intra-URL retry: if 429 and we haven't exhausted per-URL retries, wait and retry the same URL
if (
Expand Down Expand Up @@ -467,7 +495,10 @@ export class BaseExecutor {
// Distinguish timeout errors from other abort errors
const err = error instanceof Error ? error : new Error(String(error));
if (err.name === "TimeoutError") {
log?.warn?.("TIMEOUT", `Fetch timeout after ${this.getTimeoutMs()}ms on ${url}`);
log?.warn?.(
"TIMEOUT",
`Fetch timeout after ${this.getTimeoutMs()}ms on ${url}`
);
Comment on lines 497 to +501

Copilot AI Apr 14, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Timeout logging uses BaseExecutor.FETCH_START_TIMEOUT_MS in the warning message, but the actual timeout duration used for this request is read into fetchStartTimeoutMs earlier. If FETCH_START_TIMEOUT_MS is changed at runtime (e.g., in tests or future dynamic config), the log can report the wrong value. Consider capturing the effective timeout outside the try/catch and using that captured value in the log message.

Copilot uses AI. Check for mistakes.
}
lastError = err;
if (urlIndex + 1 < fallbackCount) {
Expand Down
14 changes: 8 additions & 6 deletions open-sse/handlers/chatCore.ts
Original file line number Diff line number Diff line change
Expand Up @@ -968,10 +968,10 @@ export async function handleChatCore({
let translatedBody = body;
const isClaudePassthrough = sourceFormat === FORMATS.CLAUDE && targetFormat === FORMATS.CLAUDE;
const isClaudeCodeCompatible = isClaudeCodeCompatibleProvider(provider);
// Respect the client's explicit non-streaming intent for CC-compatible providers.
// Most upstreams can answer JSON directly; the SSE->JSON fallback remains as a
// compatibility path when an upstream still responds with event-stream.
const upstreamStream = stream;
// CC-compatible providers are most reliable when OmniRoute always requests SSE
// upstream. If the client asked for JSON, chatCore will still collect the SSE
// response and return a non-streaming payload after the stream finishes.
const upstreamStream = isClaudeCodeCompatible ? true : stream;
let ccSessionId: string | null = null;

// Determine if we should preserve client-side cache_control headers
Expand All @@ -998,8 +998,10 @@ export async function handleChatCore({
} else if (isClaudeCodeCompatible) {
let normalizedForCc = { ...body };

// Claude Code-compatible providers expect Anthropic Messages-shaped payloads,
// but we extract only role/text/max_tokens/effort from an OpenAI-like view first.
// CC-compatible relays are optimized for gateway compatibility, not for
// lossless request preservation. Normalize through an OpenAI-like view,
// then rebuild a Claude Code-shaped payload that is more likely to pass
// upstream client fingerprint checks than a field-for-field passthrough.
if (sourceFormat !== FORMATS.OPENAI) {
const normalizeToolCallId = getModelNormalizeToolCallId(
provider || "",
Expand Down
13 changes: 13 additions & 0 deletions open-sse/services/claudeCodeCompatible.ts
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,16 @@ import {
} from "./claudeCodeConstraints.ts";
import { obfuscateInBody } from "./claudeCodeObfuscation.ts";

/**
* `anthropic-compatible-cc-*` targets Anthropic relay gateways that only accept
* traffic which looks like the official Claude Code client, often because those
* gateways resell the same models at materially lower prices than the direct API.
*
* This bridge is intentionally compatibility-first, not lossless. We normalize
* requests into the smallest Claude Code-shaped surface that consistently passes
* provider-side client checks, instead of trying to preserve every original
* field one-to-one.
*/
export const CLAUDE_CODE_COMPATIBLE_PREFIX = "anthropic-compatible-cc-";
export const CLAUDE_CODE_COMPATIBLE_DEFAULT_CHAT_PATH = "/v1/messages?beta=true";
export const CLAUDE_CODE_COMPATIBLE_DEFAULT_MODELS_PATH = "/models";
Expand Down Expand Up @@ -111,6 +121,9 @@ export function buildClaudeCodeCompatibleHeaders(
stream = false,
sessionId?: string | null
): Record<string, string> {
// These headers intentionally mirror Claude Code's wire image closely.
// For CC-compatible relays, passing the upstream's client-gating checks is
// more important than forwarding arbitrary caller-specific header shapes.
return {
"Content-Type": "application/json",
Accept: stream ? "text/event-stream" : "application/json",
Expand Down
1 change: 0 additions & 1 deletion src/app/(dashboard)/dashboard/combos/page.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -3195,7 +3195,6 @@ function ComboFormModal({ isOpen, combo, onClose, onSave, activeProviders }) {
<input
type="number"
min="1000"
max="600000"
step="1000"
value={config.timeoutMs ?? ""}
placeholder="120000"
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ export default function ComboDefaultsTab() {
const numericSettings = [
{ key: "maxRetries", label: t("maxRetriesLabel"), min: 0, max: 5 },
{ key: "retryDelayMs", label: t("retryDelayLabel"), min: 500, max: 10000, step: 500 },
{ key: "timeoutMs", label: t("timeoutLabel"), min: 5000, max: 300000, step: 5000 },
{ key: "timeoutMs", label: t("timeoutLabel"), min: 5000, step: 5000 },
{ key: "maxComboDepth", label: t("maxNestingDepth"), min: 1, max: 10 },
];

Expand Down Expand Up @@ -439,7 +439,6 @@ export default function ComboDefaultsTab() {
<Input
type="number"
min="5000"
max="300000"
step="5000"
value={config.timeoutMs ?? 120000}
onChange={(e) =>
Expand Down
2 changes: 1 addition & 1 deletion src/shared/validation/schemas.ts
Original file line number Diff line number Diff line change
Expand Up @@ -226,7 +226,7 @@ const comboRuntimeConfigSchema = z
strategy: comboStrategySchema.optional(),
maxRetries: z.coerce.number().int().min(0).max(10).optional(),
retryDelayMs: z.coerce.number().int().min(0).max(60000).optional(),
timeoutMs: z.coerce.number().int().min(1000).max(600000).optional(),
timeoutMs: z.coerce.number().int().min(1000).optional(),
concurrencyPerModel: z.coerce.number().int().min(1).max(20).optional(),
queueTimeoutMs: z.coerce.number().int().min(1000).max(120000).optional(),
healthCheckEnabled: z.boolean().optional(),
Expand Down
6 changes: 3 additions & 3 deletions tests/unit/cc-compatible-provider.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -460,7 +460,7 @@ test("validateProviderApiKey uses CC skeleton request after /models fallback", a
assert.equal(calls[1].headers.Accept, "text/event-stream");
});

test("handleChatCore respects non-streaming upstream requests for CC compatible providers", async () => {
test("handleChatCore forces SSE upstream for CC compatible providers while returning JSON to non-stream clients", async () => {
const calls = [];
globalThis.fetch = async (url, init = {}) => {
calls.push({
Expand Down Expand Up @@ -535,8 +535,8 @@ test("handleChatCore respects non-streaming upstream requests for CC compatible

assert.equal(result.success, true);
assert.equal(calls.length, 1);
assert.equal(calls[0].headers.Accept, "application/json");
assert.equal(calls[0].body.stream, undefined);
assert.equal(calls[0].headers.Accept, "text/event-stream");
assert.equal(calls[0].body.stream, true);
assert.equal(JSON.stringify(calls[0].body).includes('"cache_control"'), false);

const payload = await result.response.json();
Expand Down
4 changes: 2 additions & 2 deletions tests/unit/chatcore-translation-paths.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -457,8 +457,8 @@ test("chatCore builds Claude Code-compatible upstream requests for CC providers"
});

assert.equal(result.success, true);
assert.equal(call.headers.Accept ?? call.headers.accept, "application/json");
assert.equal(call.body.stream, undefined);
assert.equal(call.headers.Accept ?? call.headers.accept, "text/event-stream");
assert.equal(call.body.stream, true);
assert.equal(call.body.context_management.edits[0].type, "clear_thinking_20251015");
assert.equal(typeof call.body.metadata.user_id, "string");
assert.equal(call.body.messages[0].role, "user");
Expand Down
16 changes: 16 additions & 0 deletions tests/unit/combo-config.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,22 @@ test("resolveComboConfig ignores null and undefined overrides", () => {
assert.equal(result.strategy, "priority");
});

test("updateComboDefaultsSchema accepts arbitrarily large timeout defaults and provider overrides", () => {
const parsed = updateComboDefaultsSchema.parse({
comboDefaults: {
timeoutMs: 3600000,
},
providerOverrides: {
anthropic: {
timeoutMs: 5400000,
},
},
});

assert.equal(parsed.comboDefaults.timeoutMs, 3600000);
assert.equal(parsed.providerOverrides.anthropic.timeoutMs, 5400000);
});

test("resolveComboConfig preserves explicit empty handoffProviders overrides", () => {
const result = resolveComboConfig(
{
Expand Down
Loading