fix(resilience): stop hammering permanently-moved endpoints and billing-suspended accounts - #11774
Merged
diegosouzapw merged 249 commits intoAug 29, 2026
Conversation
…ouzapw#189, diegosouzapw#190) Bumps: nanoid ^3.3.17 (was transitive, now overridden), dompurify ^3.4.13 (with monaco-editor scoped override). Closes Dependabot diegosouzapw#189, diegosouzapw#190. Remaining diegosouzapw#182-diegosouzapw#188 (js-yaml + mermaid) already closed by diegosouzapw#9651 merge — awaiting Dependabot re-scan. npm audit → 0 vulnerabilities.
…egosouzapw#190 Closes Dependabot diegosouzapw#189 (dompurify 3.4.13) and diegosouzapw#190 (nanoid 3.3.17). npm audit → 0.
_tasks is a SEPARATE nested git repo (gitignored). The pattern _tasks/ (trailing slash) ignores only a directory, not a SYMLINK named _tasks. A self-referential _tasks symlink can slip in via git add -A and, once pulled, checkout materializes it over the real _tasks repo (destroying plans/specs/hands-off). Anchored /_tasks ignores the symlink too, preventing re-capture.
…pw#10026) Mirror the request-time exclusion rule (provider_specific_data.excludedModels) in the unified catalog builder: a model is hidden when its provider has connections but none of them is eligible for it. Applied across the PROVIDER_MODELS, synced, custom, alias-backed, and managed-fallback loops so ghost models no longer appear as available. Co-authored-by: ritheshcn25 <ritheshcn25@users.noreply.github.com>
…osouzapw#10055) * fix(models): memoize getModelsDevPricing for /v1/models catalog resolveCatalogPricing called getModelsDevPricing once per model while building GET /v1/models. Each call re-scanned models_dev_pricing and JSON.parsed every row (~10k SQL scans + multi-GB parse work), pegging the event loop so even /healthz timed out (diegosouzapw#9685, diegosouzapw#10052). Memoize the parsed map until saveModelsDevPricing / clearModelsDevPricing and add a unit test for invalidation. Signed-off-by: Ravi Tharuma <RaviTharuma@users.noreply.github.com> * fix(db): invalidate modelsDevPricing cache on DB reset (diegosouzapw#10055) Copilot review fixes: 1. Register invalidateModelsDevPricingCache() with DB state reset system so resetDbInstance() clears the process-local memo, preventing stale pricing data from surviving across DB reset/restore operations. 2. Add test assertion verifying DB reset bypasses the memo (Copilot diegosouzapw#10055). The process-local memo at modelsDevSync.ts:204 caches getModelsDevPricing() results until saveModelsDevPricing()/clearModelsDevPricing() to avoid re-scanning all pricing rows on every /v1/models request. Without this hook, backup restore and test DB resets would serve stale cached data from the previous connection. Tests: npm run test:unit:serial -- tests/unit/modelsDevSync-extended.test.ts --------- Signed-off-by: Ravi Tharuma <RaviTharuma@users.noreply.github.com> Co-authored-by: Ravi Tharuma <RaviTharuma@users.noreply.github.com> Co-authored-by: Cursor Agent <cursoragent@cursor.com>
⭐5 — Fornecedores locais compartilhados (ollama-local, LM Studio, vLLM) declaram passthroughModels:true no registry, mas hasPerModelQuota() não consultava o registry compartilhado — fallha de modelo faltante virava cooldown de conexão inteira. Broadens a classificação de model-lockout. TDD + 78/291 testes + typecheck + lint verdes. Fecha diegosouzapw#11071.
Landed with the design call resolved per the owner's pick — **option 1**: the synced store is now endpoint-agnostic (persistDiscoveredModels and managedModelImport no longer drop non-chat models at write time), and chat selectability moved to read time (auto-pool expansion in autoStrategy applies filterChatSelectableModels; the models-route projection already had its chatOnly filter). Your discovery test now passes end-to-end (3/3): /api/show capabilities persist per connection and image/embedding requests route through the advertising host. Reconciliation notes: conflicted areas merged onto the current tip (adobe discovery import, requestedModel preflight signature, resolvedProvider fast-path coexists with the synced-route override — explicit resolution wins); carried base-red drains (diegosouzapw#10055 memoization, diegosouzapw#11071 test variants) dropped as already-landed; the managed-model-import exclusion test was propagated to the new contract (image/video models persist; the read filter still hides them from chat pickers — pinned by a new assertion). Full battery: 205/206 focused (the one red is a confirmed periodic-timer timing flake on the loaded devbox — 20/20 isolated), autoCombo vitest 30/30, combo suites 46/46, gates + typecheck clean. Thank you @yourspraveen — the capability probe + routing design was right; it just needed the store contract opened up. Fixes diegosouzapw#11087.
…11300) (diegosouzapw#11309) Merging --admin: only fails are ESLint warnings ratchet drift (inherited) and dast-smoke (advisory, isRequired:null). Zero overlap with this PR's file scope (src/app/api/v1/models/catalog.ts).
…ota_exhausted errors (diegosouzapw#11277) (diegosouzapw#11310) Merging --admin: only fails are ESLint warnings ratchet drift (inherited base-red) and dast-smoke (advisory, isRequired:null). Zero overlap with this PR's scope (src/lib/usage/providerLimits.ts).
…stream ids (diegosouzapw#11326) Merging --admin with red discrimination (merge-gates §4). Fails: ESLint warnings ratchet drift (inherited base-red), Unit Tests shards containing stream-timing.test.ts (CPU-contention timing flake, assert.ok(total >= 15)ms — unrelated to this PR's scope, open-sse/handlers/imageGeneration.ts), and dast-smoke (advisory, isRequired:null).
…rtener (diegosouzapw#11329) Validated on a 17-PR combined board: TSX parses clean, eslint clean. Adapta tutorial CTA href now points at the branded shortener (link.omniroute.online/adapta) while keeping the visible link text as the real domain. Completes diegosouzapw#11196's shortener rollout.
…ream (diegosouzapw#11328) Validated on a 17-PR combined board: upstream-headers-proxy-auth within the board's 287/287, typecheck:core clean, gates within baseline. proxy-authorization and proxy-authenticate join the FORBIDDEN denylist — forwarding proxy-authorization to a model provider would hand that provider the operator's own proxy credential. Thank you @ntdat812!
… 3 DB-state tests (diegosouzapw#11327) Validated on a 17-PR combined board: capture-critical-db-state 7/7 (all three previously-skipped tests now run) within the board's 287/287, typecheck:core clean. Fixes the racy DATA_DIR-after-dynamic-import isolation and removes a duplicate type declaration. Thank you @pacocartones!
…iegosouzapw#11325) Validated on a 17-PR combined board: i18n-placeholder-parity within the board's 287/287, typecheck:core clean. Restores 3 dropped placeholders in pt.json (the visible one: the cache tile's subtitle was repeating its own label instead of showing the total) and adds a 42-locale placeholder-set gate so this class of drift can't recur silently. Thank you @ntdat812!
…diegosouzapw#11322) Validated on a 17-PR combined board: typecheck:core clean, gates within baseline. Restores 3 missing pt-BR CLI keys (setup.opencode, serve.tls_cert, serve.tls_key) — parity restored, 823/823. Thank you @pacocartones!
…i.yml gates (diegosouzapw#11321) Validated on a 17-PR combined board: validate-release-green within the board's 287/287, typecheck:core clean. Two accuracy bugs in the release-green verdict tool: an unanchored regex blamed a passing test line (matching a filename containing 'fail'), and 6 gates were double-recorded as both hard-failure and drift due to an id-format mismatch (ci.yml script name vs curated id). Found while reading the diegosouzapw#9985 verdict — good catch.
…11320) Validated on a 17-PR combined board: token-health-check + token-health-no-refresh-token-expired-5326 + token-refresh-service within the board's 287/287, typecheck:core clean. GitHub access-token-only connections are now actively verified on each due health interval (via the existing Copilot token exchange); the parent credential is marked expired only on a confirmed 401, never on 403/429/5xx/network failures; response bodies and transport messages no longer enter token-refresh logs. Closes diegosouzapw#10352. Thank you @RaviTharuma!
…ouzapw#11319) Validated on a 17-PR combined board: upstream-proxy-host-spelling 8/8 within the board's 287/287, typecheck:core clean. Routes src/lib/db/upstreamProxy.ts through the shared outbound-guard helpers instead of a private dotted-quad regex copy that had drifted since diegosouzapw#10843 — closes the IPv4-mapped IPv6, ULA, link-local and CGNAT bypasses while preserving the deliberate loopback allow (CLIProxyAPI on localhost:8317). Multicast widened from /224\. to the full 224.0.0.0/4, called out explicitly. Thank you @ntdat812!
…souzapw#11318) Validated on a 17-PR combined board: compression-worker + colocate-standalone-esm-scope within the board's 287/287, typecheck:core clean, env-doc-sync clean. Offloads eligible sync compression engines into a bounded worker_threads pool with a strict serializable DTO boundary and fail-open on spawn/worker/timeout failure. Closes diegosouzapw#11023. Thank you @RaviTharuma!
Validated on a 17-PR combined board: gemini-tts + vertex-media + audio-speech-handler (41/41) within the board's 287/287, typecheck:core clean. Registers public google/gemini-*-tts speech models and translates OpenAI-compatible /v1/audio/speech to the AI Studio generateContent audio contract, reusing the Vertex inline-audio/PCM/WAV conversion path. Batch TTS only, Gemini Live is out of scope. Thank you @RaviTharuma!
…osouzapw#9985) (diegosouzapw#11317) Validated on the resolved merge against the current release tip: pack-artifact-policy + cli-mcp-call-commands + cli-resilience-commands + cli-skills-commands + model-hide-multikey-11300 39/39, typecheck:core clean, eslint clean. Resolved a pt-BR.json wording conflict against diegosouzapw#11322 (kept the tip's wording, semantically identical). Drains the real lint-fallout from the wave that was blocking the release-green verdict — dead code + newly-enforced React-Compiler hook rules. Thank you @jonlwheat2-gif!
…zapw#11314) Validated on a 17-PR combined board: cliproxy-accounts + cliproxy-tab + cliproxy-account-health + cliproxy-resolve-spawn-args-6877 (16/16) within the board's 287/287, typecheck:core clean, env-doc-sync clean. Exposes a sanitized read-only CLIProxyAPI account health view (5s-bounded client, explicit allowlist excluding names/paths/emails/tokens/status messages) through a management-authenticated API + dashboard card. Closes diegosouzapw#6342. Thank you @RaviTharuma!
…uzapw#11312) Validated on a 17-PR combined board: elevenlabs-native-routes + hard-session-lease-bypass-inventory (9/9) within the board's 287/287, typecheck:core clean. Native ElevenLabs compatibility routes (voices, TTS, STT) reusing the stored credential via quota-preflight, sent only as xi-api-key; client authorization headers never forwarded. Closes diegosouzapw#10556. Thank you @RaviTharuma!
…uzapw#11311) Validated on a 17-PR combined board: group-model-pattern-regex-escape within the board's 287/287, typecheck:core clean. matchesModelPattern() only substituted * before compiling to RegExp — every other metacharacter kept its regex meaning, so a malformed group pattern (unbalanced parens/brackets) threw uncaught and broke EVERY request for keys in that group, not just the malformed rule (isModelAllowedForKey has no try/catch and runs on the chat completion path and the /v1/models catalog). Thank you @ntdat812!
) Validated on a 17-PR combined board: models-catalog-combo-metadata + ollama-cloud-reasoning-effort-tiers-10788 within the board's 287/287, typecheck:core clean, check:open-sse-typecheck clean, vitest 405/405. Publishes Ollama Cloud's native none/low/medium/high/max effort vocabulary for reasoning-capable passthrough/tagged models with no exact registry declaration, adds none to DeepSeek V4/GLM 5.x, and preserves narrower exact-model vocabularies (GPT-OSS) via intersection. Refs diegosouzapw#10788. Thank you @ekinnee!
* fix(deps): keep unused pnpm peers out of production * docs(changelog): link dependency policy fix to PR 11342
…sponses (diegosouzapw#11826) (diegosouzapw#11916) stripStore() now forces store=false for stateless OpenAI-compatible Responses-API targets unless the connection explicitly opts in via providerSpecificData.openaiStoreEnabled, instead of only handling the openai/agentrouter cases — a client-supplied store value previously passed through untouched to backends that don't actually persist responses server-side. Closes diegosouzapw#11826. Thanks!
…ouzapw#11918/diegosouzapw#11916 (diegosouzapw#11938) Adds the 3 missing changelog fragments.
…elease/v3.8.51 Fifteen unit files were red on this branch's PRs; running them on the pre-sync tip (d5dfcff) and on the synced one showed thirteen already failed before the sync — the cycle's own drift — and exactly two regressed: - open-sse/services/tokenExtractionConfig.ts: git kept BOTH sides' identical volcengine-console config (23 entries instead of 22). The duplicate is gone. - src/lib/usage/providerLimits.ts: the sync took release/v3.8.50's cooldown release helper, which is looser than this branch's diegosouzapw#11277 contract (it frees an extra_usage block when the policy is off and a window with no reset evidence). tests/unit/provider-limits-recovery.test.ts pins the contract; the pre-sync call site is restored and the unused helper and its imports dropped. 20/20 again, siblings unchanged.
…pfs (diegosouzapw#11896) * fix(ci): keep the next-build artefact on disk, not on the runner's tmpfs On the .113 pool /tmp is a 12 GB tmpfs — it is RAM. The 1.3 GB next-build artefact was parked there four times over: the Build job tar'd it to /tmp/e2e-build.tar.gz (6 min), three E2E jobs downloaded it to /tmp/ and extracted from there, and npm-publish.yml pulled it with gh run download into /tmp/next-build. Measured on the v3.8.50 publish runs: that download step took 27 min (9th attempt) and 32 min (10th) — 42% of a 76-minute job — while the very same bytes upload from disk in 2 min and the box pulls from GitHub at 7.3 MB/s (1.3 GB ≈ 3 min). Network was never the bottleneck; a tmpfs at 75% under memory pressure was. Every site now uses $RUNNER_TEMP / ${{ runner.temp }}: per-runner, on disk (_work/_temp under the runner dir on the pool, /home/runner/work/_temp on hosted images), and cleaned by the runner between jobs. It also removes a latent race: e2e-build.tar.gz is a FIXED name under a /tmp shared by every runner on the box, so two E2E shards on different runners could overwrite each other's download mid-extraction. RUNNER_TEMP is per runner. The supply-chain guard in tests/unit/npm-publish-artifact-provenance.test.ts pins the candidate-run selection and the --name, not the directory; it stays green. check:workflows --ratchet: zizmor unchanged at the baseline. * fix(ci): download the next-build artefact to a workspace-relative dir (pwsh has no $RUNNER_TEMP) The Electron Package Smoke matrix runs on windows-latest, whose default shell is pwsh: $RUNNER_TEMP is empty there (pwsh spells it $env:RUNNER_TEMP), so the first cut's tar -xzf "$RUNNER_TEMP/e2e-build.tar.gz" tried to open '/e2e-build.tar.gz' and failed. A path relative to the workspace works in bash and pwsh alike, and hosted workspaces are ephemeral. The producer (Build, Linux, bash) and npm-publish keep $RUNNER_TEMP.
…iegosouzapw#11932) The .113 box (31 GB) holds one next-build (14–16 GB RSS) comfortably and two at the edge; on 2026-08-28 the kernel killed main's build twice while PR builds ran beside it. Labels are the runner-side cap: only omniroute-113-5 and omniroute-113-6 carry omni-build (added through the runners API, no re-registration), and every job that runs a next build — ci.yml build, npm-publish.yml publish, both nightly-release-green validations — now asks for that label. A third heavy job queues on GitHub instead of racing for memory. The six other runners keep omni-release and no longer take builds. Pairs with the heavy-build-* concurrency lanes (diegosouzapw#11901); documented in docs/ops/RUNNER_BOX.md.
…-main-into-3851-20260828b
…ote-only replace CodeQL js/incomplete-sanitization (diegosouzapw#888): the hand-rolled replace only escaped double quotes, so a backslash in the fixture would have produced a malformed YAML scalar. JSON.stringify covers every escape the double-quoted YAML scalar needs.
…diegosouzapw#11931) * feat(ci): publish to npm through Trusted Publishing (OIDC) by default npm rejects provenance from self-hosted runners and is retiring tokens that bypass 2FA; v3.8.49 answered with staged publishing (WS1.3) so a leaked token could never publish alone — at the price of a manual `npm stage approve` per release. Trusted Publishing gives the same guarantee with no token at all: the github-hosted stage-npm job exchanges GitHub's id-token for a credential scoped to that run, provenance included, and the flow is automatic again as it was up to v3.8.48. publish_mode gains `auto` (the default, also the path for the release event); `staged` now runs only when asked for; `direct` stays as the emergency token fallback. Until the owner registers the Trusted Publisher on npmjs.com (diegosouzapw/OmniRoute, workflow npm-publish.yml) the automatic step fails with ENEEDAUTH and either other mode can be dispatched — documented in docs/ops/RELEASE_CHECKLIST.md. * docs(release): date the checklist for the Trusted Publishing change and drop the env-var claim check-deprecated-versions flags a touched doc whose header still says 2026-06-28 / v3.8.40; the fabricated-docs gate read the backticked NPM_TOKEN as an environment variable the code never reads (it is a repository secret).
diegosouzapw#11919 and diegosouzapw#11876 shipped on release/v3.8.51 (diegosouzapw#11944) Eleven PRs landed on release/v3.8.51 while the branch carried fifteen base reds, and nine more red tests hid among them. None is a defect in the shipped code; each test still encoded the contract that the merged PR deliberately replaced: - openai-to-claude finish deferral (dd35750, diegosouzapw#11933): a finish chunk that carries no usage is now held until the end-of-stream flush that production performs (open-sse/utils/stream.ts flush -> translateResponse(..., null, state)). The drivers in stream-markdown-token-boundary, translator-tool-call-shim and gemini-malformed-function-call-finish-reason-2462 fed the finish chunk and asserted the terminal events immediately; they now mirror the flush. Assertions unchanged. - authoritative live catalog (3d2832b, diegosouzapw#11919 fixes diegosouzapw#11829): a synced catalog replaces the static registry, so model-lifecycle-integration no longer expects the static-only gpt-5.6-sol row to survive a sync. The diegosouzapw#8627 contract the file guards (stale chat rows suppressed, typed media retained) is untouched. - provider asset provenance (diegosouzapw#11876): the unit shards check out with depth 1. The fixture pinned a historical commit as auditedCommit (absent on a shallow clone), the "binds auditedCommit" case relied on the repository root commit (the grafted HEAD on a shallow clone, which matches the physical snapshot), and the real-manifest case needs the audited commit fetched. The fixture now audits HEAD, the mismatch case builds a dangling empty-tree commit (no ref written), and the real-manifest case skips only on a shallow checkout that lacks the commit - the gate itself keeps running on both fetch-depth-0 rails, which the next test asserts. All five files pass locally (30, 11, 38, 3 and 18 tests); lint with the frozen suppressions is clean.
…was born with (diegosouzapw#11940) * fix(release): drain the twelve reds every PR against release/v3.8.51 was born with Measured on the cycle tip: fifteen unit files were red on every PR. Two came from the v3.8.50 sync-back (fixed in diegosouzapw#11929); the other thirteen predate it and are the branch's own drift. This sweep clears all of them but the ESLint debt (diegosouzapw#11924), each with the smallest change that keeps the guard honest: - .env.example + ENVIRONMENT.md: NEXT_PUBLIC_SW_BUILD_ID / OMNIROUTE_SW_BUILD_ID / SOURCE_VERSION (diegosouzapw#11779 service-worker cache busting) documented — the env/docs contract gate was failing on every PR. - stryker.conf.json: the six tests the mutation gate found covering mutated modules (four retirement runtime-block suites, combo connection-aware expansion, tunnel error sanitization) registered in tap.testFiles. - dependency-allowlist: eslint-plugin-react-hooks 7.0.1 approved; its findings are tracked in diegosouzapw#11924. - i18n: the six combo.sort.* strings (d5dfcff) translated for vi (strict parity) and pt-BR. - docs/providers/CHATGPT_WEB.md: the retirement test is migration-168, not 163. - g4f gateways: authHint now says member key, which the discontinued-providers guard asserts. - tests realigned to the catalog the branch actually ships: qwen-web (diegosouzapw#11713) and chatgpt-web (diegosouzapw#11720) are retired, so web-session-contract and token-health-check-webcookie use perplexity-web, grok-web and chatgpt-web-codex. - db-core-init: the two minimal legacy fixtures gained the columns migrations 164-168 UPDATE (error_code, last_error*, test_status) — they exist on every real legacy DB (base CREATE TABLE); the fixtures simply never declared them. - no-js-extension guard: a .js specifier whose target is a genuine JavaScript file (open-sse/lib/deepseek-pow-hash.js, shared with a worker) is not the diegosouzapw#10674 defect; the test now skips targets that exist as .js. All twelve files pass locally; docs-sync, docs-counts, env-doc-sync, the tap drift gate and the fabricated-docs gates are green on the tree. * test(release): move the deferred-finish translator test into a collected path tests/unit/translator/ is not one of the unit collectors (package.json test:unit, merge-train.sh, build-test-impact-map, check-test-discovery), so the suite that dd35750 added there never ran — check:test-discovery flagged it as a new orphan on every PR. Relocated next to its sibling openai-to-claude-trailing-usage-11817 under tests/unit/, where the root glob collects it (5/5 pass). * fix(dashboard): type the four sort-method sites diegosouzapw#11812 left red on the dashboard typecheck ratchet d5dfcff added the combo model sort and raised combos/page.tsx from 23 to 27 scoped TypeScript errors (TS2339 +1, TS2345 +2, TS2322 +1), which fails check:dashboard-typecheck on every PR against release/v3.8.51: - initialSortMethod: sanitizeComboRuntimeConfig() is untyped, so config.modelSort is unknown; narrow it before reading .method (normalizeSortMethod takes unknown anyway). - handleAddModels: the batch path passes ComboBuilderDraftModelStep[] to the ComboStep[] sort helpers without the cast handleSortChange already uses; mirror it. - ComboSortSelect expects a translate-with-fallback (k, f) => string, but received next-intl's Translator whose second argument is a values object. Pass the page's getI18nOrFallback adapter instead of the raw translator — that is also what makes the `has()` check and the fallback text actually work at runtime. Baseline untouched (no widening). Scoped tsc: 0 new/regressed errors.
…ote-only replace (diegosouzapw#11942) CodeQL js/incomplete-sanitization (alert diegosouzapw#888 on diegosouzapw#11929): the hand-rolled replace only escaped double quotes, so a backslash in the fixture would have produced a malformed YAML scalar. JSON.stringify covers every escape the double-quoted YAML scalar needs. Test-only change (7/7 pass).
…finish test diegosouzapw#11940 moved tests/unit/translator/openai-to-claude-trailing-usage.test.ts one level up so a collector would run it, but kept the ../../../ import path from the old directory, so the file failed to load and painted Unit Tests fast-path (3/4) red on every PR since a94fe23. The path now matches its new location (5/5 pass).
…from a clean-room run (diegosouzapw#11955) * chore(quality): re-freeze the ESLint suppressions on release/v3.8.51 from a clean-room run `No new ESLint warnings` failed on every PR against release/v3.8.51 with exit 2: "There are suppressions left that do not occur anymore". Measured in a depth-1 clone with `npm ci` from the branch's own lockfile and the job's exact command (`npm run lint:json -- --max-warnings 0`): 56 errors — 55 `no-explicit-any` in six files that landed while the base was red (diegosouzapw#11843 isFree tests: 22; b710214 socks connect timeout: 33) plus one `no-unused-vars` — and stale entries for files that no longer violate. The devbox figure previously quoted in diegosouzapw#11924 (280, with 224 react-hooks/*) does not reproduce on the lockfile install and is withdrawn. - config/quality/eslint-suppressions.json: `--prune-suppressions` (two stale file entries removed) and the 55 pre-existing `any` frozen at their exact counts — the file is a ratchet, counts only go down; the debt stays tracked in diegosouzapw#11924. - open-sse/services/adobeFireflyCatalog.ts: remove `GPT_SIZE_MAP`, a constant the f3d9279 split left behind with no reader (the real violation, fixed not frozen). Verification in the clean room after both changes, same command as CI: exit 0, 0 errors, 0 warnings (1238 files / 5487 suppressions). * chore(quality): tighten openapiCoverage.pct to the measured 39 (require-tighten) With ESLint back to 0/0 on this PR, the job's next step (check-quality-ratchet --require-tighten) started failing: openapiCoverage.pct improved from 38.4 to 39 (delta 0.6 > slack 0.5) and the baseline must be tightened in the same PR. 39 is the value CI collect-metrics measured on run 33213844112 and a clean-room checkout of 777d9d1 reproduces it; the cycle's new routes landed documented in docs/openapi.yaml. Only this metric moves; annotation follows the file's convention.
…int caches (diegosouzapw#11963) - quality.yml fast-unit: timeout-minutes: 30. A shard finishes in ~10 min; without a ceiling a hung test process holds the PR for GitHub's 6 h default. On 2026-08-28 shard 1/4 sat 64 min without a line of output — twice at the same spot, a timing race that vanished on the third run — while the other three shards were long green. A fast red plus a re-run beats a silent multi-hour hold. - quality.yml lint-guard + the earlier ESLint cache block: drop the `restore-keys: eslint-<os>-` fallback (diegosouzapw#11600, P-II.1 of the v3.8.50 postmortem). The key already hashes the lint config, the suppressions file and the lockfile; the fallback restored a cache built under a DIFFERENT configuration and its stale per-file verdicts are how 215 pre-existing errors stayed invisible for a cycle. Exact key or a cold full lint — never a partial cache from another configuration. check:workflows --ratchet unchanged (194/194); check-workflows suite 32/32.
…osouzapw#11964) Records the v3.8.50 → v3.8.51 precedent in the single source of truth: a main → release/vX+1 sync PR lands by fast-forward push so main stays an ancestor of the release branch (squash re-conflicts the next sync-back on every file main touched — 551 conflicts this cycle before the two-step merge), plus the two post-landing checks (ancestry assert; ratchet files carried main's freezes). The full procedure lives in the generate-release Phase 5 skill.
…apw#11946, option 3) (diegosouzapw#11962) The hosted 7 GB runner cannot build release/v3.8.51 in any profile: `Build App` (build.yml, push on every branch, full `build:release`) died in 19 of the last 30 runs — the branch tip included — with "The runner has received a shutdown signal" ~8 min into `next build`, swapfile and all; the advisory quality.yml build failed on 8/8 recent fork PRs with the same recipe; and `DAST smoke (PR)`'s backend-only build died ~7 min in before the server even started, hidden as a permanently red continue-on-error check. Together they painted every PR into release/** red with zero signal and, on build.yml, produced an artefact nothing downloads. - build.yml: workflow_dispatch only. The bundle is validated where a build fits — ci.yml `Build` on the self-hosted omni-build pool after every merge to main, and nightly-release-green.yml on the same pool for release/**. - dast-smoke.yml: pull_request into main only (plus workflow_dispatch to smoke a release branch by hand); main's tree still builds on the hosted runner in ~5.5 min. - quality.yml: the fork-only rationale of `Build (advisory)` updated to say why own-origin PRs no longer get a hosted build either. Behaviour unchanged. check:workflows --ratchet: 194 zizmor findings, baseline 194. check-workflows and backend-only-smoke-workflows suites pass. Trade-off stated in the PR: own-origin PRs into release/** lose a pre-merge build that was not succeeding anyway; the nightly rail files a base-red issue within a day if a merge breaks the build.
…iegosouzapw#11965) (diegosouzapw#11967) Four nightly jobs run a backend-only `next build` on ubuntu-latest (7 GB): Schemathesis, promptfoo injection guard, garak probes and the axe a11y suite (self-building webServer). On release/v3.8.51 three of them died with the hosted VM shutdown signature and nobody saw it — nightlies have no audience — and the fourth passes by a margin of minutes. They now target [self-hosted, omni-light] (hosted fallback when USE_VPS_RUNNER is off), a new two-listener label on the .113 box for jobs that need ~6 GB, not the 14-16 GB of a full build; they run once a day in the 04:00-06:00 UTC window, when the box is idle. Fleet reshaped the same day and documented in docs/ops/RUNNER_BOX.md: 4 active OmniRoute listeners (omniroute-113-5/-6 omni-build, omniroute-113/-2 omni-light), omniroute-113-3/-4/-7/-8 disabled (systemctl enable --now brings one back), janitor ceiling MAX_ACTIVE_RUNNERS=4. The remaining headroom limit is the VM's 31 GB of RAM (2 heavy + 2 light ≈ 42 GB peak, inside the 16 GB swap); more RAM on the Proxmox VM is the lever that turns the label ceilings into 3 heavy + 2 light. check:workflows --ratchet unchanged (194/194); check-workflows and backend-only-smoke-workflows suites pass; docs-sync PASS.
…ard on ENOTEMPTY (diegosouzapw#11966) (diegosouzapw#11968) * test(infra): retry recursive temp-dir removal instead of failing a shard on ENOTEMPTY (diegosouzapw#11966) Two shards on release/v3.8.51 went red in one day with the same signature — "ENOTEMPTY, Directory not empty: /tmp/omniroute-<test>-XXXXXX" — from combo-same-provider-cascade (Unit Tests fast-path 4/4, on a PR that touches only .github/) and auth-policy-embeddings-webfetch-7785 (the 20k-test TIA step). Both pass alone and on re-run: the cleanup races something still writing into the directory (SQLite WAL/-shm checkpoint, a worker, the backup) and under a loaded hosted runner the window opens. 1154 test files do their own cleanup with fs.rmSync(dir, { recursive: true, force: true }); 57 already asked for retries. One-shot codemod (scripts/ad-hoc/codemod-rm-maxretries.mjs, kept for the record): every rm / rmSync / rmdirSync option object with `recursive: true` and no `maxRetries` gains `maxRetries: 5, retryDelay: 100` — Node itself then retries ENOTEMPTY/EBUSY/EPERM for up to ~0.5 s before giving up. 2243 call sites in 1292 files under tests/, the shared tests/_setup/isolateDataDir.ts exit hook included. Only the option object changes: no call site, assertion or import is touched. Validation: prettier and ESLint (with the frozen suppressions) clean on all 1292 files; a random 20-file sample runs green (quota-redis-store hangs identically on the untouched tree — it needs a Redis on localhost, an environment matter). The four unit shards on this PR are the full run. * fix(quality): let check-forgotten-sibling-tests read a 1,000-file diff The gate shells out to `git diff` through execFileSync with Node's default 1 MB maxBuffer; the 1,292-file codemod in this PR is the first diff large enough to overflow it, and the gate died with `spawnSync git ENOBUFS` before comparing anything. 64 MB is far above any real PR and costs nothing when unused.
…ob and the main run (diegosouzapw#11972) The job has timeout-minutes: 20; the c8 merge across 8 shards takes ~10 min and the Codecov upload (declared informational) then hung for the rest of the budget on two consecutive main runs (33207760653, 33215115341) — GitHub cancels the step, the job ends cancelled, and the run's conclusion turns cancelled although every blocking job was green. The upload step now has its own 5-minute ceiling and continue-on-error; the job budget is 30 min. check-workflows suite 32/32; zizmor ratchet unchanged.
…ead to the npm leg (release/v3.8.51 twin of diegosouzapw#11973) (diegosouzapw#11974) Same three changes as diegosouzapw#11973 on main, applied to this branch's newer copy of the workflow so the v3.8.51 tag does not repeat v3.8.50's zero-asset release: publish-npm grants actions:read (the called publish job requests it — a caller that grants less is refused at startup and the release job dies with it), a publish_npm dispatch input gates the npm leg, and web-build/build/release check out the tag named by the dispatch. actionlint clean; the five workflow-pinning suites pass.
…ouzapw#11924 (diegosouzapw#11975) Production (open-sse/utils/socksConnectorWithFamily.ts, 4 sites): every cast was redundant — undici's buildConnector.BuildOptions already has `timeout?: number | null`, socks' SocksClientOptions has `timeout?: number`, and Agent.Options' `connect` / `connectTimeout` narrow to the connector's parameter types on their own. Behaviour unchanged; check:open-sse-typecheck stays at the frozen 5. Tests (51 sites): the socks-timeout mocks now carry the real types — the patched SocksClient.createConnection is typed as the static it replaces, the fake buildConnector returns buildConnector.connector, the proxy is a SocksProxy, the dynamic import is typed as the module it loads; the e2e suite passes a SocksProxy and Agent.Options and no longer casts undici's fetch init (its RequestInit already has `dispatcher`); the isFree suites narrow getCustomModels()' JSON to a declared row shape, feed deliberately-wrong values through `unknown`, and stop casting for zod's safeParse, which takes unknown. The six files' suppression entries are removed: 1238 → 1232 files, 5487 → 5432 suppressed. ESLint without the suppressions file reports 0 problems on all six; with it, no stale entry is left. The five suites pass (4, 2, 5, 4, 4).
…koff (Gemini ban prevention) (diegosouzapw#11762) Root-caused via a real Gemini-ban incident log: deprecated-model 404/410s (e.g. gemini-2.5-flash "no longer available to new users") fell through checkFallbackError's generic transient-cooldown branch, so combo/auto-routing kept re-selecting a permanently dead model every cooldown window forever — the hammering that got the account flagged as abusive. Fix: `MODEL_PERMANENTLY_UNAVAILABLE_PATTERNS` + `isModelPermanentlyUnavailable()` classify these as a 24h lockout instead, surfaced via `quotaResetHintMs` so combo's per-request model-lockout honors it in full. Validated: 6/6 new tests + 133/133 existing accountFallback/error-classification tests, no regressions. Thanks for tracing this end-to-end with real production logs!
diegosouzapw
merged commit Aug 29, 2026
91f9a01
into
diegosouzapw:release/v3.8.51
14 of 15 checks passed
diegosouzapw
pushed a commit
that referenced
this pull request
Aug 29, 2026
… requests (#11781) Follow-up to #11762/#11774, same bug class in combo's own model-lockout wiring: GitHub rejects several models (gpt-5.4, gpt-5.3-codex, etc.) with a 400 that's permanently unavailable for this account's Copilot integration, but nothing recorded a cross-request lockout — combo's #5249 in-request advance guard is correct but doesn't persist, so the same doomed model gets retried from scratch on every new request, indefinitely. Fix: on a model-scoped 400 (`isModelScoped400`), call `lockModelIfPerModelQuota(provider, connectionId, rawModel, "model_capacity", 1h)`. GitHub already has per-model-quota enabled, so only the rejected model locks — siblings keep working. `isModelLocked()` is already checked pre-dispatch, so no other wiring needed. Validated: 3/3 new tests + fixed a pre-existing test-isolation gap in combo-model-scoped-400-advance.test.ts (shared model name across sub-tests without clearing lockout state). Thanks!
diegosouzapw
pushed a commit
that referenced
this pull request
Aug 29, 2026
…EADME (#11772) Finishes the Freepik → Magnific rebrand from #10594 across 40 locale files and 3 README feature-list bullets (README.md, docs/i18n/it, docs/i18n/tr) — legacy `freepik` alias intentionally left in code/tests/redirects for backward compatibility, and historical CHANGELOG entries left untouched as documented history. The README bullet had base-drifted since the PR branched (release tip's "What's New" changelog snippet had already dropped two providers mentioned nowhere else in the codebase, unrelated to this PR's scope) — resolved by keeping the tip's current bullet shape and applying only the Freepik→Magnific rename on top, in both the combined-worktree validation and the pushed branch. Validated: all 40 edited locale JSON files parse; re-verified after resync onto the updated tip (post #11762/#11774/#11781).
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…ng-suspended accounts (diegosouzapw#11774) Follow-up to diegosouzapw#11762, same bug class hitting freeaiapikey (410 permanently-moved endpoint) and fireworks (412 billing-suspension) — both fell through checkFallbackError's generic transient-cooldown branch and got retried every ~1 minute for a full day. Fix: `ENDPOINT_PERMANENTLY_MOVED_PATTERNS`/`isEndpointPermanentlyMoved()` → 24h lockout; `ACCOUNT_SUSPENDED_BILLING_PATTERNS`/`isAccountSuspendedForBilling()` → treated as credits-exhausted (1h cooldown), independent of status code so it also catches Fireworks' 412. diegosouzapw#11762 landed first and touched the same file — rebased/re-merged onto the updated tip (additive, no logic changes) and re-validated: 13/13 tests pass. Thanks for tracing this with real production logs again!
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
… requests (diegosouzapw#11781) Follow-up to diegosouzapw#11762/diegosouzapw#11774, same bug class in combo's own model-lockout wiring: GitHub rejects several models (gpt-5.4, gpt-5.3-codex, etc.) with a 400 that's permanently unavailable for this account's Copilot integration, but nothing recorded a cross-request lockout — combo's diegosouzapw#5249 in-request advance guard is correct but doesn't persist, so the same doomed model gets retried from scratch on every new request, indefinitely. Fix: on a model-scoped 400 (`isModelScoped400`), call `lockModelIfPerModelQuota(provider, connectionId, rawModel, "model_capacity", 1h)`. GitHub already has per-model-quota enabled, so only the rejected model locks — siblings keep working. `isModelLocked()` is already checked pre-dispatch, so no other wiring needed. Validated: 3/3 new tests + fixed a pre-existing test-isolation gap in combo-model-scoped-400-advance.test.ts (shared model name across sub-tests without clearing lockout state). Thanks!
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…EADME (diegosouzapw#11772) Finishes the Freepik → Magnific rebrand from diegosouzapw#10594 across 40 locale files and 3 README feature-list bullets (README.md, docs/i18n/it, docs/i18n/tr) — legacy `freepik` alias intentionally left in code/tests/redirects for backward compatibility, and historical CHANGELOG entries left untouched as documented history. The README bullet had base-drifted since the PR branched (release tip's "What's New" changelog snippet had already dropped two providers mentioned nowhere else in the codebase, unrelated to this PR's scope) — resolved by keeping the tip's current bullet shape and applying only the Freepik→Magnific rename on top, in both the combined-worktree validation and the pushed branch. Validated: all 40 edited locale JSON files parse; re-verified after resync onto the updated tip (post diegosouzapw#11762/diegosouzapw#11774/diegosouzapw#11781).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #11762 — same bug class, different providers
After #11762 (Gemini deprecated-model 404 lockout), I checked a fresh 1h request log and found the same underlying bug hitting other providers: a permanent failure signal isn't classified by
checkFallbackError, so it falls through to the generic transient-error branch (short, seconds-to-minutes cooldown) and gets retried roughly every cooldown window — forever, since the target can never succeed. Two concrete cases:1.
freeaiapikey— permanently moved API endpoint (410), retried every ~1 minute for the whole hourThis recurs at 08:52, 08:53, 08:56, 08:57, 09:01, 09:02, 09:05, 09:16, 09:19 — i.e. roughly every 1–3 minutes, non-stop, for both
claude-sonnet-4.6andclaude-opus-4.6. The base_url is stale and will never succeed until an operator updates it, but nothing incheckFallbackErrorrecognized "endpoint has moved / no longer works" as permanent, so it kept getting the same short cooldown as a normal transient error.2.
fireworks— account suspended for billing (412), also retried every ~1 minuteTwo problems compounded here:
412isn't incheckFallbackError'sretryableStatusesset at all, so it never got dedicated handling.ACCOUNT_DEACTIVATED_SIGNALS' fixed substring"your account has been suspended"— different phrasing, different provider.Result: the same suspended Fireworks account got hit dozens of times over the hour with a request that can only ever fail until the invoice is paid.
The fix
In
open-sse/services/accountFallback.ts(mirrors the pattern from #11762):ENDPOINT_PERMANENTLY_MOVED_PATTERNS+isEndpointPermanentlyMoved()— matches "endpoint has moved", "no longer works", "update your base_url". A 410/404 matching this now gets a 24h lockout (reason: "not_found", surfaced viaquotaResetHintMsso combo's per-request model-lockout honors it in full instead of the normal ~20min ceiling).ACCOUNT_SUSPENDED_BILLING_PATTERNS+isAccountSuspendedForBilling()— matches "suspended ... spending limit/billing/invoice/payment" in either word order, independent of status code (so it also catches Fireworks' un-mapped 412). Classified the same way as an existing credits-exhausted account (creditsExhausted: true, 1h cooldown) so it reuses the existingcredits_exhaustedterminal connection status and self-heals if the invoice gets paid.Testing
Added
tests/unit/permanent-failure-hammering-other-providers.test.ts(7 tests, all passing):isEndpointPermanentlyMovedmatches freeaiapikey's exact phrasing, doesn't match a generic"[410]: Gone".isAccountSuspendedForBillingmatches Fireworks' exact phrasing, doesn't match an unrelated suspension message.checkFallbackErrorreturns a 24h lockout for the moved-endpoint 410.checkFallbackErrorreturnscreditsExhausted: truefor the billing-suspended 412.Ran locally:
Also registered the new test in
stryker.conf.json'stap.testFiles(learned from CI on #11762 that this is required for themutation-test-coveragegate).Files changed
open-sse/services/accountFallback.ts— new pattern sets + classification branches.tests/unit/permanent-failure-hammering-other-providers.test.ts— new regression test.stryker.conf.json— register the new test file.