Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 33 additions & 1 deletion open-sse/config/providerErrorRules.ts
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,35 @@ function buildOpencodeRules(): ProviderErrorRule[] {
];
}

// ─── Cloudflare Workers AI ────────────────────────────────────────────────
// Free tier = 10,000 Neurons/day, shared across the WHOLE account
// (docs/reference/FREE_TIERS.md; official: developers.cloudflare.com/
// workers-ai/platform/errors/). The exhaustion body doesn't match any
// QUOTA_PATTERNS keyword (src/shared/utils/classify429.ts) so it falls
// through to rate_limit and gets retried every ~60s against a budget that
// only resets at UTC midnight.
//
// Scope note: `scope: "connection"` (not "provider") for the same reason as
// Opencode above — the neuron budget is per-account, and a single OmniRoute
// connection maps to one Cloudflare account. Multiple connections under the
// same provider name would mean multiple accounts, each with its own budget.
function buildCloudflareAiRules(): ProviderErrorRule[] {
return [
{
id: "cloudflare-ai-daily-neuron-allocation",
match: ({ status, body }) => {
if (status !== 429) return null;
const text = JSON.stringify(body ?? "").toLowerCase();
if (!text.includes("daily free allocation")) return null;
// No cooldownMs: recordModelLockoutFailure already sets
// quota_exhausted without one to "next UTC midnight" — exactly this
// budget's real reset semantics.
return { reason: "quota_exhausted", scope: "connection" };
},
},
];
}
Comment on lines +119 to +134

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

There are two critical issues with the current implementation that will prevent the daily neuron exhaustion from being correctly cooled down in production:

1. applyErrorState does not pass structuredError or headers to checkFallbackError

In open-sse/services/accountFallback.ts, applyErrorState is responsible for applying the error state and cooldown to the connection/account. It calls checkFallbackError with only 5 arguments:

const fallbackDecision = checkFallbackError(status, errorText, backoffLevel, null, provider);

Because headers and structuredError are omitted, they default to null/undefined. Consequently, getProviderErrorRuleMatch is called with body = null and headers = null. Thus, this cloudflare-ai-daily-neuron-allocation rule (which matches on body) will never match during the connection cooldown phase, and the connection will fall back to a standard 5-second rate-limit cooldown.

2. Missing cooldownMs in the rule

Even if the rule did match, it does not return a cooldownMs. While recordModelLockoutFailure has fallback logic to set quota_exhausted to "next UTC midnight", applyErrorState (which handles connection-level lockouts) does not. It relies entirely on the cooldownMs returned by checkFallbackError. If no cooldownMs is specified by the provider rule, it falls back to the default scaled backoff cooldown (e.g., 5 seconds).

Recommended Fixes:

  1. Return cooldownMs directly from the rule: Update the rule to calculate and return the cooldownMs (time until next UTC midnight) directly (as suggested below).
  2. Update isDailyQuotaExhausted in accountFallback.ts: Since accountFallback.ts is not in the diff, please also update isDailyQuotaExhausted in that file to include "daily free allocation". This ensures that checkFallbackError natively recognizes the daily quota exhaustion and applies the correct midnight cooldown even when body is not passed:
export function isDailyQuotaExhausted(errorText: string): boolean {
  if (!errorText) return false;
  const lower = errorText.toLowerCase();
  return (
    lower.includes("today's quota") ||
    lower.includes("daily quota") ||
    lower.includes("try again tomorrow") ||
    lower.includes("daily free allocation")
  );
}
Suggested change
function buildCloudflareAiRules(): ProviderErrorRule[] {
return [
{
id: "cloudflare-ai-daily-neuron-allocation",
match: ({ status, body }) => {
if (status !== 429) return null;
const text = JSON.stringify(body ?? "").toLowerCase();
if (!text.includes("daily free allocation")) return null;
// No cooldownMs: recordModelLockoutFailure already sets
// quota_exhausted without one to "next UTC midnight" — exactly this
// budget's real reset semantics.
return { reason: "quota_exhausted", scope: "connection" };
},
},
];
}
function buildCloudflareAiRules(): ProviderErrorRule[] {
return [
{
id: "cloudflare-ai-daily-neuron-allocation",
match: ({ status, body }) => {
if (status !== 429) return null;
const text = JSON.stringify(body ?? "").toLowerCase();
if (!text.includes("daily free allocation")) return null;
// Calculate ms until next UTC midnight to precisely match Cloudflare's reset window
const now = Date.now();
const nextMidnight = new Date(now);
nextMidnight.setUTCHours(24, 0, 0, 0);
const cooldownMs = nextMidnight.getTime() - now;
return { reason: "quota_exhausted", scope: "connection", cooldownMs };
},
},
];
}


// ─── Minimax ────────────────────────────────────────────────────────────────
// Minimax returns per-model quota info via custom headers. The body is generic
// "rate limit exceeded" so we MUST read the headers. Other models on the same
Expand Down Expand Up @@ -141,6 +170,7 @@ export const providerRuleRegistry = new Map<string, ProviderErrorRule[]>([
["opencode-cli", buildOpencodeRules()],
["minimax", buildMinimaxRules()],
["minimax-passthrough", buildMinimaxRules()],
["cloudflare-ai", buildCloudflareAiRules()],
]);

/**
Expand Down Expand Up @@ -194,7 +224,9 @@ export function getProviderErrorRuleMatch(
*/
export function parseResetCountdownMs(text: string): number | null {
if (typeof text !== "string" || text.length === 0) return null;
const match = text.match(/resets?\s+in\s+(\d+)\s+(day|days|hour|hours|minute|minutes|second|seconds)\b/);
const match = text.match(
/resets?\s+in\s+(\d+)\s+(day|days|hour|hours|minute|minutes|second|seconds)\b/
);
if (!match) return null;
const n = Number(match[1]);
if (!Number.isFinite(n) || n <= 0) return null;
Expand Down
9 changes: 9 additions & 0 deletions src/shared/utils/classify429.ts
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,15 @@ const QUOTA_PATTERNS: ReadonlyArray<RegExp> = [
/individual quota reached/i,
/enable overages/i,
/INSUFFICIENT_G1_CREDITS_BALANCE/i,

// Cloudflare Workers AI daily neuron budget exhaustion ("you have used up
// your daily free allocation of 10,000 neurons, please upgrade to
// Cloudflare's Workers Paid plan..."). No "quota"/"limit"/"exceed"/"credit"
// substring, so none of the patterns above match it. This is the primary
// provider-specific rule in providerErrorRules.ts (scope: "connection");
// this entry is defense-in-depth for classify429FromError callers that
// bypass provider rule matching.
/daily free allocation/i,
];

/**
Expand Down
44 changes: 44 additions & 0 deletions tests/unit/provider-error-rules.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,50 @@ test("S2b: provider error rules match canonical-cased plain header records", asy
assert.equal(minimaxMatch.scope, "model");
});

test("S2c: Cloudflare Workers AI daily neuron exhaustion → QUOTA_EXHAUSTED with connection scope", async () => {
// Cloudflare's free tier is a 10,000-Neurons/day budget shared across the
// WHOLE account. The exhaustion body doesn't contain "quota"/"limit"/
// "exceed"/"credit" so it would otherwise fall through to the default
// RATE_LIMIT_EXCEEDED and get retried every ~60s against a budget that
// only resets at UTC midnight.
const { getProviderErrorRuleMatch } = await import("../../open-sse/config/providerErrorRules.ts");

const match = getProviderErrorRuleMatch(
"cloudflare-ai",
429,
{},
{
errors: [
{
message:
"AiError: AiError: you have used up your daily free allocation of 10,000 neurons, please upgrade to Cloudflare's Workers Paid plan if you would like to continue usage.",
code: 4006,
},
],
success: false,
}
);
assert.ok(match, "cloudflare-ai must have a rule matching the daily neuron allocation body");
assert.equal(match.reason, "quota_exhausted");
assert.equal(
match.scope,
"connection",
"Cloudflare's neuron budget is account-wide, so the lock must scope to the connection, not a single model"
);

// A 429 without the exhaustion wording (a real transient rate-limit) must
// NOT match — this rule is specific to the daily-allocation body.
const noMatch = getProviderErrorRuleMatch(
"cloudflare-ai",
429,
{},
{
errors: [{ message: "Too many requests, please retry shortly.", code: 3040 }],
}
);
assert.equal(noMatch, null, "a generic 429 without the exhaustion wording must not match");
});

test("S3: Regression — provider with no rules falls back to global ERROR_RULES unchanged", () => {
// A provider not in the registry (e.g. "unknown-vendor") must NOT cause
// classifyError to crash or return a different result. It must behave
Expand Down