Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -2131,6 +2131,12 @@ APP_LOG_TO_FILE=true
# Used by: open-sse/services/rateLimitManager.ts
# RATE_LIMIT_MAX_WAIT_MS=15000

# Limiter-managed execution backstop (Bottleneck `expiration`): bounds a job's
# post-dispatch execution, never queue wait. Must stay ABOVE upstream
# fetch-start timeouts on non-incremental gateways. Default: 600000 (10 min)
# Used by: open-sse/services/rateLimitManager.ts
# RATE_LIMIT_EXECUTION_MAX_WAIT_MS=600000

# Rate limit queue admission cap: reject with 429 queue_full once this many requests
# are already queued (0 = disabled/unbounded, the default). Used by: open-sse/services/rateLimitManager.ts
# RATE_LIMIT_MAX_QUEUE_DEPTH=0
Expand Down
10 changes: 9 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -976,7 +976,11 @@ jobs:
# 10min was sized before #7114 added the lcov reporter (Codecov/Sonar need it);
# merging 8 shard JSONs + text+json+lcov now takes ~10-12min — three consecutive
# release-tip runs died at exactly 10m as job-timeout "cancelled" (2026-07-15/16).
timeout-minutes: 20
# 30, not 20 (2026-08-29): the informational Codecov upload below hung for the rest of
# the budget on two consecutive main runs (33207760653, 33215115341); the job ended
# `cancelled` and dragged the whole run's conclusion to `cancelled` although every
# blocking job was green. The upload step now has its own ceiling; this is headroom.
timeout-minutes: 30
needs: test-unit
if: ${{ !cancelled() && needs.test-unit.result == 'success' && !contains(github.event.pull_request.labels.*.name, 'hotfix') }}
env:
Expand Down Expand Up @@ -1055,6 +1059,10 @@ jobs:
# (if-no-files-found: warn) — Sonar consumes the same file.
- name: Upload coverage to Codecov (informational)
if: always()
# Informational means informational: its own ceiling and continue-on-error, so a
# stalled upload can neither eat the job's budget nor turn a green job cancelled.
timeout-minutes: 5
continue-on-error: true
uses: codecov/codecov-action@fb8b3582c8e4def4969c97caa2f19720cb33a72f # v7.0.0
with:
files: coverage/lcov.info
Expand Down
22 changes: 22 additions & 0 deletions .github/workflows/electron-release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,11 @@ on:
description: "Release version (e.g., v1.6.8)"
required: true
type: string
publish_npm:
description: "Also run the npm publish leg (turn off when re-attaching desktop assets to a release whose npm package already shipped)"
required: false
default: true
type: boolean

# Least-privilege default: read-only at the top level; each job grants the writes it
# needs (build/release upload assets, publish-npm forwards npm provenance / packages
Expand Down Expand Up @@ -76,6 +81,9 @@ jobs:
- uses: actions/checkout@v7
with:
persist-credentials: false
# workflow_dispatch: build the tag being (re)built, not the dispatching branch. On a
# tag push this resolves to the same commit.
ref: ${{ needs.validate.outputs.version }}
- name: Setup Node
uses: actions/setup-node@v7
with:
Expand Down Expand Up @@ -161,6 +169,9 @@ jobs:
- uses: actions/checkout@v7
with:
persist-credentials: false
# workflow_dispatch: build the tag being (re)built, not the dispatching branch. On a
# tag push this resolves to the same commit.
ref: ${{ needs.validate.outputs.version }}
- name: Setup Node
uses: actions/setup-node@v7
with:
Expand Down Expand Up @@ -347,6 +358,8 @@ jobs:
with:
persist-credentials: false
fetch-depth: 0
# Source archives + SBOM come from the tag being released, not the dispatching branch.
ref: ${{ needs.validate.outputs.version }}

# `merge-multiple` is deliberately OFF. It resolves same-name collisions by ARRIVAL
# ORDER, and the two macOS jobs each emit their own `latest-mac.yml` listing only their
Expand Down Expand Up @@ -462,11 +475,20 @@ jobs:
publish-npm:
name: Publish to npm
needs: [validate, release]
# A re-dispatch that only re-attaches desktop assets must not publish the npm package again.
if: ${{ github.event_name != 'workflow_dispatch' || inputs.publish_npm }}
permissions:
# Must be `write`, not `read`: this job calls the reusable npm-publish.yml whose
# `publish` job needs `contents: write` (gh release upload — attach the SBOM, #3874).
# A reusable workflow's job cannot request more permission than the caller grants,
# so a `read` here makes GitHub reject the run at startup (startup_failure).
#
# `actions: read` for the same reason: the called `publish` job downloads the next-build
# artefact and requests it. v3.8.50 (run 33005490476) died at startup with "The nested
# job 'publish' is requesting 'actions: read', but is only allowed 'actions: none'" — and
# because `release` lives in this same workflow, the tag shipped with ZERO assets. Keep
# this block a superset of every job's permissions in npm-publish.yml.
actions: read
contents: write
id-token: write # npm provenance (forwarded to the reusable workflow)
packages: write # publish to npm.pkg.github.com
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(resilience):** decouple the limiter-managed execution backstop from the queue-wait budget — new `requestQueue.executionMaxWaitMs` (env `RATE_LIMIT_EXECUTION_MAX_WAIT_MS`, default 600000 = 10 min) now feeds Bottleneck's post-dispatch `expiration`, while `requestQueue.maxWaitMs` keeps its documented queue-wait semantics. Previously the queue-wait budget doubled as the execution expiration, so legitimate long-running calls on non-incremental gateways (whole generation buffered before the first upstream byte, e.g. Console Go / Command Code tiers serving GLM models) were killed mid-flight at the queue budget with a false 504 `RATE_LIMIT_EXECUTION_TIMEOUT` — the local limiter undercut the provider-aware upstream fetch-start timeouts. The surfaced 504 message now names `requestQueue.executionMaxWaitMs`; the error keeps the #4165 guarantees (disclaims an upstream timeout, preserves the Bottleneck error as `cause`, branded code + trusted provenance, classified request-scoped so combo falls back). A real queue-wait bound (the `Promise.race` around `limiter.schedule()` sketched in #9533) remains future work. (#12025)
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(diagnostics):** preserve the error field (truncated to 4KB with a `[truncated: …]` suffix) in every call-log artifact size-limit fallback stage. Previously the minimal fallback replaced the error with `[omitted: call log artifact size limit exceeded]`, so an oversized artifact row showed nothing about WHY the request failed — e.g. 91 of 847 opencode-go 504 rows on one production instance were undiagnosable from the dashboard. Oversized request/response bodies are still omitted exactly as before; the error cap is independent of the payload sizes that tripped the fallback. (#12026)
1 change: 1 addition & 0 deletions changelog.d/fixes/v3850-electron-release-assets.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- Electron release workflow: the `publish-npm` job now grants `actions: read` to the reusable `npm-publish.yml` it calls (its `publish` job requests it), which is what made GitHub refuse the whole v3.8.50 run at startup and ship the release with zero desktop assets; a `workflow_dispatch` now builds the requested tag instead of the dispatching branch and can skip the npm leg (`publish_npm=false`) when only re-attaching assets
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- `Coverage` job on `ci.yml`: the informational Codecov upload gets its own 5-minute ceiling and `continue-on-error`, and the job budget grows from 20 to 30 minutes (the 8-shard c8 merge alone takes ~10) — a stalled upload no longer ends the job `cancelled` and drags a fully green `main` run's conclusion down with it
23 changes: 18 additions & 5 deletions open-sse/services/rateLimitManager.ts
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,16 @@ export function resolveRequestQueueMaxWaitMs(
return resolveOverride(override, legacyDefault);
}

/**
* Limiter-managed execution backstop (Bottleneck `expiration`). Starts only
* after a job leaves QUEUED; bounds execution, never queue wait. Kept strictly
* separate from the queue-wait budget (`maxWaitMs`) so the backstop cannot
* undercut upstream fetch-start timeouts on non-incremental gateways.
*/
export function resolveExecutionMaxWaitMs(): number {
return currentRequestQueueSettings.executionMaxWaitMs;
}

function buildLimiterDefaults() {
// 0 or missing values mean "infinite" / no rate limit applies. This treats
// the global request-queue settings the same way per-connection overrides
Expand Down Expand Up @@ -562,10 +572,13 @@ export async function withRateLimit(provider, connectionId, model, fn, signal =
await awaitProviderDefaultSlot(provider, connectionId, signal, maxWaitMs);

const limiter = getLimiter(provider, connectionId, model);
// Bottleneck's `expiration` starts only after a job leaves QUEUED. The
// legacy maxWaitMs setting therefore bounds limiter-managed execution; it
// is not a queue-wait deadline.
const executionExpirationMs = maxWaitMs;
// Bottleneck's `expiration` starts only after a job leaves QUEUED, so it
// bounds limiter-managed execution — not queue wait. It is therefore fed by
// the dedicated execution backstop (`requestQueue.executionMaxWaitMs`),
// never by the queue-wait budget: non-incremental gateways legitimately run
// for minutes before first bytes, and an expiration at the queue budget
// killed them mid-flight (false 504s on opencode-go/glm-5.3-flash).
const executionExpirationMs = resolveExecutionMaxWaitMs();
const scheduleOpts =
executionExpirationMs && executionExpirationMs > 0 ? { expiration: executionExpirationMs } : {};

Expand Down Expand Up @@ -641,7 +654,7 @@ export async function withRateLimit(provider, connectionId, model, fn, signal =
throw markLocalRateLimitError(
new Error(
`Request exceeded OmniRoute's local rate-limit execution expiration ` +
`(legacy resilienceSettings.requestQueue.maxWaitMs=${executionExpirationMs}ms) for ` +
`(resilienceSettings.requestQueue.executionMaxWaitMs=${executionExpirationMs}ms) for ` +
`${model ? `${provider}/${model}` : provider}. Bottleneck applies this deadline only ` +
`after dispatch; it does not bound queue wait and is not an upstream-generated timeout.`,
{ cause: err }
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ type RequestQueueSettings = {
minTimeBetweenRequestsMs: number;
concurrentRequests: number;
maxWaitMs: number;
executionMaxWaitMs: number;
};

type ConnectionCooldownProfileSettings = {
Expand Down Expand Up @@ -254,6 +255,13 @@ function RequestQueueCard({
suffix="ms"
onChange={(maxWaitMs) => setDraft((prev) => ({ ...prev, maxWaitMs }))}
/>
<NumberField
label={t("resilienceMaxExecutionWait")}
value={draft.executionMaxWaitMs}
min={1}
suffix="ms"
onChange={(executionMaxWaitMs) => setDraft((prev) => ({ ...prev, executionMaxWaitMs }))}
/>
</>
) : (
<>
Expand Down Expand Up @@ -289,6 +297,12 @@ function RequestQueueCard({
{formatMs(value.maxWaitMs)}
</div>
</div>
<div className="rounded-xl border border-border bg-bg-subtle p-4">
<div className="text-xs text-text-muted">{t("resilienceMaxExecutionWait")}</div>
<div className="mt-1 text-sm font-semibold text-text-main">
{formatMs(value.executionMaxWaitMs)}
</div>
</div>
</>
)}
</div>
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/ar.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "تتحكم هذه الطبقة في الصف والسرعة فقط. لا يقوم بتخزين فترات التهدئة أو قواطع الدائرة المفتوحة.",
"resilienceAutoEnableApiKeyProvidersDesc": "لتمكين حماية قائمة الانتظار بشكل افتراضي لاتصالات مفتاح API النشطة.",
"resilienceMaxQueueWait": "الحد الأقصى لوقت الانتظار في قائمة الانتظار",
"resilienceMaxExecutionWait": "مهلة التنفيذ (حد التنفيذ الاحتياطي لتحديد المعدل)",
"resilienceConnectionCooldownScope": "اتصال فردي",
"resilienceConnectionCooldownTrigger": "عندما يُرجع الاتصال فشلًا عابرًا في المنبع",
"resilienceConnectionCooldownEffect": "يتخطى هذا الاتصال مؤقتًا ويزيد من التراجع في حالة الفشل المتكرر",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/az.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "This layer only controls queueing and pacing. It does not store cooldowns or open circuit breakers.",
"resilienceAutoEnableApiKeyProvidersDesc": "Enables queue protection by default for active API key connections.",
"resilienceMaxQueueWait": "Maximum queue wait time",
"resilienceMaxExecutionWait": "İcra vaxt limiti (sürət limiti ehtiyat həddi)",
"resilienceConnectionCooldownScope": "Individual connection",
"resilienceConnectionCooldownTrigger": "When a connection returns a transient upstream failure",
"resilienceConnectionCooldownEffect": "Temporarily skips that connection and increases backoff for repeated failures",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/bg.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "Този слой контролира само опашката и темпото. Той не съхранява охлаждания или отворени прекъсвачи.",
"resilienceAutoEnableApiKeyProvidersDesc": "Активира защита на опашката по подразбиране за активни API ключ връзки.",
"resilienceMaxQueueWait": "Максимално време за изчакване на опашка",
"resilienceMaxExecutionWait": "Таймаут на изпълнението (резервен лимит на скоростта)",
"resilienceConnectionCooldownScope": "Индивидуална връзка",
"resilienceConnectionCooldownTrigger": "Когато връзката върне преходна грешка нагоре по веригата",
"resilienceConnectionCooldownEffect": "Временно пропуска тази връзка и увеличава забавянето при повтарящи се повреди",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/bn.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "এই স্তরটি শুধুমাত্র সারিবদ্ধ এবং পেসিং নিয়ন্ত্রণ করে। এটি কুলডাউন বা খোলা সার্কিট ব্রেকার সংরক্ষণ করে না।",
"resilienceAutoEnableApiKeyProvidersDesc": "সক্রিয় API কী সংযোগের জন্য ডিফল্টরূপে সারি সুরক্ষা সক্ষম করে৷",
"resilienceMaxQueueWait": "সর্বোচ্চ সারি অপেক্ষার সময়",
"resilienceMaxExecutionWait": "এক্সিকিউশন টাইমআউট (রেট-লিমিট ব্যাকস্টপ)",
"resilienceConnectionCooldownScope": "স্বতন্ত্র সংযোগ",
"resilienceConnectionCooldownTrigger": "যখন একটি সংযোগ একটি ক্ষণস্থায়ী আপস্ট্রিম ব্যর্থতা প্রদান করে",
"resilienceConnectionCooldownEffect": "সাময়িকভাবে সেই সংযোগটি এড়িয়ে যায় এবং বারবার ব্যর্থতার জন্য ব্যাকঅফ বাড়ায়",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/cs.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "Tato vrstva řídí pouze řazení a rychlost zobrazování. Neukládá cooldowny ani přerušené jističe.",
"resilienceAutoEnableApiKeyProvidersDesc": "Ve výchozím nastavení povoluje ochranu fronty pro aktivní připojení klíče API.",
"resilienceMaxQueueWait": "Maximální doba čekání ve frontě",
"resilienceMaxExecutionWait": "Časový limit provádění (pojistný limit rychlosti)",
"resilienceConnectionCooldownScope": "Individuální připojení",
"resilienceConnectionCooldownTrigger": "Když připojení vrátí přechodné selhání proti proudu",
"resilienceConnectionCooldownEffect": "Dočasně toto připojení vynechá a zvýší backoff pro opakované selhání",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/da.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "Dette lag styrer kun kø og pacing. Den opbevarer ikke nedkøling eller åbne afbrydere.",
"resilienceAutoEnableApiKeyProvidersDesc": "Aktiverer købeskyttelse som standard for aktive API-nøgleforbindelser.",
"resilienceMaxQueueWait": "Maksimal ventetid i kø",
"resilienceMaxExecutionWait": "Eksekveringstimeout (rate-limit-sikkerhed)",
"resilienceConnectionCooldownScope": "Individuel tilslutning",
"resilienceConnectionCooldownTrigger": "Når en forbindelse returnerer en forbigående opstrømsfejl",
"resilienceConnectionCooldownEffect": "Springer midlertidigt den forbindelse over og øger backoff for gentagne fejl",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/de.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "Diese Ebene steuert nur Warteschlange und Taktung. Sie speichert keine Cooldowns und öffnet keine Circuit Breaker.",
"resilienceAutoEnableApiKeyProvidersDesc": "Aktiviert den Queue-Schutz standardmäßig für aktive API-Key-Verbindungen.",
"resilienceMaxQueueWait": "Maximale Wartezeit in der Queue",
"resilienceMaxExecutionWait": "Ausführungs-Timeout (Rate-Limit-Rückfallebene)",
"resilienceConnectionCooldownScope": "Einzelne Verbindung",
"resilienceConnectionCooldownTrigger": "Wenn eine Verbindung einen vorübergehenden Upstream-Fehler zurückgibt",
"resilienceConnectionCooldownEffect": "Überspringt diese Verbindung vorübergehend und erhöht den Backoff bei wiederholten Fehlern",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/en.json
Original file line number Diff line number Diff line change
Expand Up @@ -8016,6 +8016,7 @@
"resilienceRequestQueueDesc": "This layer only controls queueing and pacing. It does not store cooldowns or open circuit breakers.",
"resilienceAutoEnableApiKeyProvidersDesc": "Enables queue protection by default for active API key connections.",
"resilienceMaxQueueWait": "Maximum queue wait time",
"resilienceMaxExecutionWait": "Execution timeout (rate-limit backstop)",
"resilienceConnectionCooldownScope": "Individual connection",
"resilienceConnectionCooldownTrigger": "When a connection returns a transient upstream failure",
"resilienceConnectionCooldownEffect": "Temporarily skips that connection and increases backoff for repeated failures",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/es.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "Esta capa solo controla las colas y el ritmo. No almacena tiempos de reutilización ni disyuntores abiertos.",
"resilienceAutoEnableApiKeyProvidersDesc": "Habilita la protección de colas de forma predeterminada para conexiones de clave API activas.",
"resilienceMaxQueueWait": "Tiempo máximo de espera en cola",
"resilienceMaxExecutionWait": "Tiempo de espera de ejecución (límite de respaldo)",
"resilienceConnectionCooldownScope": "Conexión individual",
"resilienceConnectionCooldownTrigger": "Cuando una conexión devuelve un error ascendente transitorio",
"resilienceConnectionCooldownEffect": "Omite temporalmente esa conexión y aumenta la interrupción en caso de fallas repetidas",
Expand Down
1 change: 1 addition & 0 deletions src/i18n/messages/fa.json
Original file line number Diff line number Diff line change
Expand Up @@ -8004,6 +8004,7 @@
"resilienceRequestQueueDesc": "این لایه فقط صف و سرعت را کنترل می کند. خنک کننده ها یا کلیدهای مدار باز را ذخیره نمی کند.",
"resilienceAutoEnableApiKeyProvidersDesc": "حفاظت از صف را به طور پیش فرض برای اتصالات کلید API فعال فعال می کند.",
"resilienceMaxQueueWait": "حداکثر زمان انتظار صف",
"resilienceMaxExecutionWait": "زمان انتظار اجرا (حد پشتیبان نرخ)",
"resilienceConnectionCooldownScope": "ارتباط فردی",
"resilienceConnectionCooldownTrigger": "هنگامی که یک اتصال یک شکست گذرا در بالادست را برمی گرداند",
"resilienceConnectionCooldownEffect": "به طور موقت از آن اتصال پرش می شود و برای خرابی های مکرر، عقب نشینی را افزایش می دهد",
Expand Down
Loading