cherry-pick(pr-9714): feat(resilience): expose providerQuotaOverrides via /api/resilience - #9871
Merged
Merged
Conversation
diegosouzapw
added a commit
that referenced
this pull request
Aug 12, 2026
… A2A TaskManager (PRD RF1) (#8080) * fix(api): enforce model permissions on gateway mirrors (#9854) Co-authored-by: Xiangzhe <xiangzhedev@gmail.com> * cherry-pick(pr-9787): fix(sse): apply Azure param rules on azure-ai and clamp gpt-4o-mini output tokens (#9855) * fix(sse): apply Azure request-param rules on the azure-ai wire path Azure rejects several stock Chat Completions params on its newer deployments and returns HTTP 400 rather than ignoring them: max_tokens -> 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead. reasoning_effort -> Function tools with reasoning_effort are not supported. Those rules lived inline in AzureOpenAIExecutor, so they only covered the azure-openai provider. azure-ai (Azure AI Foundry) had no executor entry and fell through to the bare DefaultExecutor, so the SAME Azure deployment succeeded on one connection and 400'd on the other. Every agentic client sends tools on every turn, so azure-ai failed on the first request. Extract the rules to open-sse/executors/azureParamRules.ts, add an AzureAiExecutor that inherits DefaultExecutor's azure-ai URL/header/apiType handling unchanged and applies the shared rules, and register it for azure-ai. Also widen the deployment pattern to cover gpt-chat-latest: it is a moving alias that resolves to a GPT-5-era model and rejects max_tokens, but carries no version number for the token-boundary pattern to key on. Verified against the base regex - gpt-chat-latest did not match, which is exactly the observed 400. Regression guard: tests/unit/azure-param-rules.test.ts, including an assertion that getExecutor("azure-ai") no longer resolves to a bare DefaultExecutor. * fix(sse): clamp Azure gpt-4o-mini completion tokens to its 16384 ceiling Azure gpt-4o-mini deployments accept at most 16384 completion tokens and 400 on anything larger: max_tokens is too large: 32000. This model supports at most 16384 completion tokens, whereas you provided 32000. The 32000 is OmniRoute's own doing: adjustMaxTokens raises any smaller max_tokens to DEFAULT_MIN_TOKENS (32000) whenever tools are present, to avoid truncated tool arguments. That floor has no upper bound, so an agentic client asking for far less still trips the model ceiling on its first turn. Add scoped maxOutputCap rules in paramSupport.ts for both Azure wire paths. PROVIDER_MAX_TOKENS is the wrong lever here - it is provider-wide, and the same Azure resource also serves GPT-5 deployments with a much higher ceiling. Regression guard: tests/unit/azure-max-output-clamp.test.ts, which also pins that the clamp does not leak to gpt-5.1 or to gpt-4o-mini on other providers. --------- Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * maint: final follow-up cherry-pick #9783 (#9904) * fix(deps): bump transitive deps for 6 Dependabot + remaining audit vulns on main Same overrides as #9464 (ip-address, hono, fast-uri, socket.io-parser, undici) applied directly to main. Also covers brace-expansion (scoped), js-yaml v4 copies, and mermaid. npm audit: 6→0 vulnerabilities. Closes Dependabot #161-#166. * fix(deps): bump nanoid, dompurify for 2 new Dependabot alerts (#189, #190) Bumps: nanoid ^3.3.17 (was transitive, now overridden), dompurify ^3.4.13 (with monaco-editor scoped override). Closes Dependabot #189, #190. Remaining #182-#188 (js-yaml + mermaid) already closed by #9651 merge — awaiting Dependabot re-scan. npm audit → 0 vulnerabilities. * fix(repo): harden .gitignore to also ignore a _tasks symlink (/_tasks) _tasks is a SEPARATE nested git repo (gitignored). The pattern _tasks/ (trailing slash) ignores only a directory, not a SYMLINK named _tasks. A self-referential _tasks symlink can slip in via git add -A and, once pulled, checkout materializes it over the real _tasks repo (destroying plans/specs/hands-off). Anchored /_tasks ignores the symlink too, preventing re-capture. * fix(translator): keep Responses namespace identity across the hub-and-spoke pivot Step 1 of the pivot (openai-responses -> openai) flattens namespace sub-tools to a qualified wire name (#8295) and records the `{namespace, name}` pair on a non-enumerable `_toolNameMap`. Step 2 (openai -> target) returns a brand-new object, so the property was dropped for every non-OpenAI target. chatCore then handed `null` to the #7936 response seam and namespace sub-tool calls reached the client under their flattened name, which Codex rejects with `unsupported call: <name>` — the symptom #7936 was opened to fix. Copying `_toolNameMap` through is not viable: openai-to-claude and openai-to-gemini publish their own `Map<string, string>` alias map on that same property during step 2, so it carries two incompatible types. This adds a dedicated `_namespaceToolIdentityMap`, propagated by translateRequest across the pivot; chatCore prefers it and falls back to `_toolNameMap` for the non-pivot producers. Both keys are stripped from the cliproxyapi wire body. Fixes #9780 * fix(chat): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reduce combined file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reduce combined file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Co-authored-by: VXNCXNX <vincent@preuve.ai> * fix(sse): route claude/<provider>/<model> aliases for catalog-only providers (#9856) The /v1/models catalog mirrors `claude/<provider>/<model>` ids purely from the alias gate -- ccAliasPredicate.ts consults no provider registry. The request path additionally required the prefix to be an open-sse REGISTRY entry or an operator-defined custom node. Enterprise-cloud providers such as azure-ai / azure-openai live only in the provider catalog (src/shared/constants/providers/apikey/enterprise-cloud.ts). They route fine directly -- `azure-ai/Phi-4` returns 200 -- but have no open-sse registry entry, so the two sides disagreed: the catalog advertised `claude/azure-ai/<model>` while stripCcDiscoveryAlias refused to strip it. The unstripped id then fell through to normal resolution, which splits on the first / and parsed `claude` as the provider. Every Claude Code request for an Azure model was routed to the Claude provider instead: ROUTING: Provider: claude, Model: azure-ai/DeepSeek-V4-Flash Extract the predicate as `isRoutableProviderPrefix()` and widen it to the provider catalog (id + alias) alongside the open-sse registry, so the request path recognises exactly what the catalog can advertise. Regression guard: tests/unit/cc-discovery-alias-routable-prefix.test.ts pins azure-ai/azure-openai/azure as routable, keeps openai/anthropic routable, and keeps an unknown prefix non-routable. Verified failing before the widening. Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * fix(i18n): translate validation model keys in 34 locales (#9857) The provider-connection dialog (AddApiKeyModal / EditConnectionModal) rendered humanized key names instead of real copy for providers.validationModelId{Label,Placeholder,Hint} in 34 of 43 locales — the values read "Validation Model Id Label", "Validation Model Id Placeholder" and "Validation Model Id Hint" verbatim. Each translation follows the terminology and register already used by the neighbouring provider keys in its own file — e.g. de Anbieter/API-Schlüssel with formal Sie, fr fournisseur/clé API, ru провайдер/ключ API — and each locale's own "e.g." convention (z. B., 例:, напр., ör., cth., hal.). Source of truth is en.json, which labels the field "Validation Model" (no "ID"); a few older locales say "validation model ID" and were left untouched rather than propagating that divergence. Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * cherry-pick(pr-9770): chore(repo): ignore Electron build output unpacked into repo root (#9858) * chore(repo): ignore Electron build output unpacked into repo root electron-builder (squirrel-windows target) unpacks the packaged app -- the entire Chromium runtime, ~24k files -- directly into the repository root: OmniRoute.exe, chrome_*.pak, *.dll, locales/, resources/, icudtl.dat, snapshot blobs and the Chromium license files. None of it was covered by .gitignore, so `git add -A` would commit the whole runtime. Every rule is root-anchored (leading `/`) because a bare `locales/` or `resources/` would also swallow tracked sources -- notably the CLI translations in bin/cli/locales/*.json. Verified with `git check-ignore`: all artifact paths ignored, and bin/cli/locales/{en,de}.json remain tracked. * chore(electron): sync package-lock for windows installer deps Adds the lockfile entries for the Windows installer/signing toolchain that the electron build now pulls in: electron-builder-squirrel-windows, electron-winstaller and @electron/windows-sign (plus their transitive fs-extra/jsonfile/universalify/mkdirp pins), and bumps app-builder-lib and builder-util-runtime. Lockfile-only change; no source or runtime behaviour is affected. --------- Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * fix(skills): normalize web fetch credentials (#9859) Co-authored-by: backryun <bakryun0718@proton.me> * fix(types): narrow DeepSeek tool calls (#9860) Co-authored-by: backryun <bakryun0718@proton.me> * fix(perf): memoize synced pricing reads (#9861) Co-authored-by: chloeassistant <279834366+chloeassistant@users.noreply.github.com> * cherry-pick(pr-9744): test(integration): add general live-test tool for the real "default" combo + rootless wire capture (#9862) * test(integration): add general live-test tool for the real "default" combo Temporary WIP commit on this deferred branch — lands in its own separate PR once the bug-fix extraction batch is done (never bundled into a bug-fix PR). Unlike liveGeminiShared.ts (provisions its own narrow 2-model Gemini-only combo), this reads the REAL "default" combo currently configured on the target instance directly from the DB and exercises every provider/model step in it directly, bypassing combo routing, so live-test coverage always matches whatever is actually configured instead of a hardcoded snapshot. Live-verified against omniroute-beta (seeded with the real 18-model, 5-provider default combo): 14/18 models pass consistently across non-streaming + streaming Chat Completions and streaming Responses API. The 4 consistent failures are real external state (cerebras credits_exhausted, one deprecated openrouter free-tier model), not code regressions. (cherry picked from commit c40b13a48fd897259c56f5122e9e57a3dc7654ba) * test(integration): add rootless wire-capture correlation to the live-test tool Temporary WIP commit on this deferred branch — lands in the same final live-test-tool PR as the general default-combo suite, never bundled into a bug-fix PR. liveContainerHarness.ts spins up a dedicated, throwaway podman container (same runner-base image target as the operator's local dev/beta containers) so wire-capture tests are fully self-contained: builds the image if missing, starts the container with a persistent data dir, waits for health, seeds the real "default" combo + provider connections from the operator's local omniroute-dev instance (idempotent — only runs once per data dir), and provisions API keys via the running instance's own auth flow. wireCapture.ts captures the container's actual network traffic via `podman unshare nsenter --net=<container netns> -- tcpdump` — no root needed, verified working live (this generalizes the root-requiring `sudo nsenter -t $PID` command scripts/sre/tcp-close-analyzer.py already documented for the same rootless-Podman netns problem; that script's docstring now documents both). Capture and analysis needed two real fixes found only by running the pipeline live: `-U` (unbuffered tcpdump writes) plus a `pkill -f <pcap path>` fallback, since `podman unshare -> nsenter -> tcpdump` is a 3-level subprocess chain and SIGTERM to the top-level process doesn't reach the tcpdump grandchild, leaving an orphaned process and a truncated/unreadable pcap; and filtering on the container's internal listening port (20128) rather than the dynamically-assigned host port, since capture happens inside the container's own network namespace where only the internal port is meaningful. live-default-combo-wire-capture.test.ts (gated on RUN_LIVE_WIRE_CAPTURE=1) ties it together: sends a small representative sample of requests through the real default combo, then cross-checks each one's app-level JSON status against the actual HTTP status line observed on the wire via scripts/sre/tcp-close-analyzer.py's stream reassembly — catching bugs where the app layer claims success but the wire shows a truncated/reset stream, not just what liveDefaultComboShared.ts's existing breadth suite already covers. Live-verified end-to-end: 4/4 sampled requests correlated correctly across 8 captured TCP streams, container + capture process fully torn down afterward (verified no orphaned podman container or tcpdump process left running). sendModelRequest/filterActiveModelTargets (liveDefaultComboShared.ts) gain optional baseUrl/apiKey overrides, defaulting to the existing module-level omniroute-beta target, so the wire-capture suite can point the same request-sending logic at its own dedicated container instead. (cherry picked from commit 914a7e42cbe914f257db9f72eedc902ee1532083) --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * maint: follow-up cherry-pick fix-in-place #9741 (conflict-resolved fallback) (#9895) * fix(responses-api): sync reasoning-cache write index with the fixed read side The turn-index-hardcoding fix updated the reasoning-cache read side (translator/index.ts's main replay loop) to key lookups by the assistant message's real position in the messages array, but two other spots still used the old hardcoded convention: - chatCore.ts's write side (both the streaming and non-streaming completion paths) still cached every response under a hardcoded messageIndex: 0. - translator/index.ts's own plain-turn (non-tool-call) cache-key lookup ALSO still hardcoded messageIndex 0 at its call site — a second, previously undiscovered instance of the same class of bug, found while re-verifying this fix against the current upstream tip (the original fix only addressed the write side). Past the first assistant turn these conventions no longer matched, so DeepSeek/Xiaomi-mimo plain-turn reasoning replay silently missed the cache and fell back to the placeholder (or, once #9573 removed the placeholder fallback, to an absent field) in ordinary multi-turn conversations. Compute the write-side index from the incoming request's message count instead, and use the real loop-provided messageIndex on the read-side lookup, both matching the position the response occupies once the client appends it to history for the next turn. Note: this was originally part of a larger squashed fix (output_index collision prevention across reasoning/message/tool_call items, reasoning-content-alias generalization) that has since been superseded by upstream's own independent fix — translator/response/openai-responses.ts now has its own dense-output-index-sort + getReadableReasoningValue implementation (own comment: "mirrors upstream PR #721"). Only this narrower, still-genuinely-broken write/read index sync survives as a distinct bug. Test plan: - TDD: tests/unit/reasoning-cache.test.ts's new end-to-end "write side (chatCore's messageIndex) and read side (translateRequest) agree on the same key end-to-end" test, plus the pre-existing "should inject placeholder for a plain (non-tool-call) DeepSeek turn" and "should replay cached reasoning for a plain (non-tool-call) DeepSeek turn when available" tests — confirmed failing against the pre-fix code on a clean release/v3.8.50 checkout (both the hardcoded-0 write side AND the hardcoded-0 read-side lookup independently reproduce the mismatch), passing after both fixes - npm run typecheck:core — clean - npm run lint — clean - npm run check:file-size — clean (chatCore.ts rebaselined 5034->5042 for the messageIndex computation at both call sites; reasoning-cache.test.ts frozen at 1035, matching the original fix's own rebaseline) - 2 pre-existing, unrelated test failures in the same file ("should replace empty-string reasoning_content with NON_ANTHROPIC_THINKING_PLACEHOLDER on cache miss", "should inject placeholder for a plain (non-tool-call) DeepSeek turn missing reasoning_content") confirmed present on a completely clean, untouched release/v3.8.50 checkout — these test obsolete placeholder-injection behavior the code deliberately removed per #9573 (see the code's own comment); not touched by this PR * fix(chat): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reconcile file-size baseline Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * cherry-pick(pr-9738): feat(logging): make the chat-log truncation limit configurable, bumped default 128x (#9863) * feat(logging): make the chat-log truncation limit configurable, bumped default 128x The 8KB cap on logged request/response bodies (open-sse/handlers/chatCore/logTruncation.ts::truncateForLog()) was hardcoded — trivially exceeded by any real multi-turn agentic conversation, meaning the dashboard's "Full Conversation" panel could only ever show a placeholder instead of the actual messages for nearly every logged row of any conversation with real substance. - Added CHAT_LOG_MAX_BODY_KB env var (src/lib/logEnv.ts:: getChatLogMaxBodyBytes()), default 1024 KB (1MB) — a 128x bump from the old hardcoded 8KB — following the same configurable-limit pattern as the sibling CHAT_LOG_TEXT_LIMIT/CHAT_LOG_ARRAY_TAIL_ITEMS/etc. vars. - Documented in .env.example and docs/reference/ENVIRONMENT.md. estimateSizeFast() (open-sse/utils/estimateSize.ts) has been substantially rewritten upstream since this bug was first found (now an iterative Frame-based walker with a separate node-visit budget, not the simple stack loop originally patched) — re-implemented the fix against the current algorithm rather than porting the old diff: the byte early-exit was unconditionally the module-level ESTIMATE_SIZE_BYTE_LIMIT (256 KiB) with no way for a caller to raise it, so any caller comparing against a bigger configured threshold could never see a size above ~256 KiB — every payload between 256 KiB and the caller's real limit looked "under threshold" and truncation never fired, the opposite of intended. Added an optional byteLimit parameter (default unchanged at ESTIMATE_SIZE_BYTE_LIMIT, so isSmallEnoughForSemanticCache's existing behavior is untouched) threaded through both the byte-check early-exit and the node-budget-exhaustion fail-closed fallback, with truncateForLog() now passing its own configured getChatLogMaxBodyBytes() value through. * feat(dashboard): show conversation session tag in request detail metadata Adds a "Conversation" field to the request detail panel's metadata grid (after "Combo"), showing the request's conversation id (sessionTag) for quick reference/copy. --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * cherry-pick(pr-9735): feat(logging): bump CHAT_LOG_ARRAY_TAIL_ITEMS default 24 -> 128 (#9864) * feat(logging): bump CHAT_LOG_ARRAY_TAIL_ITEMS default 24 -> 128 Real agentic CLIs with many MCP servers routinely declare 40-50+ tools in a single request — a live OpenClaw session logged 47. The tail-24 default silently dropped the array's earlier entries behind an _omniroute_truncated_array marker, so investigating why a specific tool call (apply_patch) behaved oddly turned up nothing: its declared shape (function vs custom type) was unrecoverable from the call log across 40 recent requests, even though the calls themselves succeeded. Bumped the configurable default to comfortably cover real large tool lists with headroom. Updated .env.example and docs/reference/ ENVIRONMENT.md to match (env-doc-sync check passes). * test(logging): pin CHAT_LOG_ARRAY_TAIL_ITEMS default at 128 The bump commit had no dedicated test asserting the literal default value; the existing chatcore-log-truncation.test.ts derives its expectations from getChatLogArrayTailItems() itself, so it can't discriminate a regression back toward the old, too-small 24 default. --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * fix(logging): use configurable max-depth when bounding logged tool_calls (#9865) requestLogger.ts's cloneBoundedForLog had its own hardcoded depth cap of 6, independent of the existing configurable getChatLogMaxDepth(). A typical Chat Completions response body's responseBody.choices[0].message.tool_calls[0].function sits at exactly depth 6, so every logged tool call's function field (name+arguments) was silently replaced with the literal string "[MaxDepth]" before ever being stored — corrupting the data, not just how it renders. Bumped the shared default 6->20 and switched requestLogger.ts to read it instead of using its own literal. (cherry picked from commit a2df6cf289cbab7cd618b8e55272434812f7a4a7) Co-authored-by: Markus Hartung <mail@hartmark.se> * fix(executors): repair DuckDuckGo AI Chat challenge solver (418 ERR_CHALLENGE) (#9866) Every duckduckgo-web chat request failed with HTTP 418 ERR_CHALLENGE while duck.ai worked normally in a browser from the same IP. Ground truth was established by driving a real headful Chromium at duck.ai from that IP (it returned 200), so the environment was never the problem — the anti-abuse challenge solver was. Six independent defects were found; the first alone disabled the solver completely. 1. Module syntax inside the vm sandbox source. CHALLENGE_STUBS is executed with vm.runInContext, which compiles in SCRIPT mode. A refactor mass-added `export` to the five `function` declarations inside that template literal (they read as ordinary top-level TS functions), so every solve threw SyntaxError. The executor swallows solve failures and posts the raw unsolved challenge, which upstream answers with 418. 2. Double-escaped regex in a String.raw template. `\\s` in __parseCssDisplay reached the sandbox as a literal backslash, so the display regex never matched and a getComputedStyle probe silently read empty. 3. buildHtmlLookup undercounted descendants by one. `count` backs el.querySelectorAll('*').length; that returns DESCENDANTS and countHtmlElements already skips the #document-fragment root, so the `- 1` was wrong. Chromium reports 3 for '<li><div></li><li></div'; we reported 2, and a variant multiplies innerHTML.length by that count. 4. Browser-fidelity probes. Newer challenge variants assert JS/DOM invariants a flat stub cannot satisfy: real prototype chains (HTMLDivElement -> HTMLElement -> Element), NodeList identity, a live body.children HTMLCollection, native-code toString, and sloppy-mode `this === window`. Nine of thirteen failed. Notably Math must NOT be sealed — Chromium reports Object.isSealed(Math) === false, and sealing it made our vector differ by one. 5. The solved payload dropped meta.origin / meta.stack / meta.duration. The duck.ai bundle always sends all three; captured browser requests confirm it. Without them upstream returns 418 even when every client_hash is correct. 6. reasoningEffort is now mandatory on duckchat/v1/chat. An otherwise byte-identical payload returns 200 with the field and 400 ERR_BAD_REQUEST without it (A/B verified live, repeated). Also removes the throwaway "seed" chat POST that ran before every real request. It existed to coax a usable challenge out of the upstream while the solver was broken; it only doubled chat calls against an IP-rate-limited endpoint, showing up as spurious 429 ERR_RATE_LIMIT. Verification: the solver now reproduces real Chromium's probe vectors exactly for all 8 captured challenge variants, and the executor returns 200 end-to-end live (non-streaming, streaming, claude-haiku-4-5, and a math prompt returning "42"). Tests: tests/unit/duckduckgo-challenge-solver-regression.test.ts (32 tests) and tests/unit/duckduckgo-reasoning-effort-required.test.ts (5 tests), backed by tests/fixtures/duckduckgo/challenge-variants.json — real captured challenge programs plus the probe vectors a real browser produced for them, so the suite asserts against recorded browser behaviour rather than our own output. Each fix was confirmed to fail its test when individually reverted. Co-authored-by: Mynacol <git@mynacol.xyz> * cherry-pick(pr-9730): fix(compression): persist RTK renderer configuration (#9867) * fix(compression): persist RTK renderer configuration * docs(changelog): add fragment for #9730 Adds the changelog.d/fixes/9730-persist-rtk-renderers.md fragment required by check:changelog-integrity for the RTK enableRenderers persistence fix in PR #9730. --------- Co-authored-by: Isaac <isaaclyons98@gmail.com> * fix(dashboard): unregister leftover service workers in dev mode (#9868) A phone that previously loaded a production build on this origin (or an old dev build from before the registration was gated) kept an active service worker across dev restarts. It intercepted every navigation/asset fetch, occasionally serving a JS chunk that didn't match the running dev server, which tripped Next's dev-client chunk-mismatch auto-reload — visible as an unexplained, unstoppable refresh loop on that device only (confirmed via a clean private tab on the same phone/URL not looping). PwaRegister now actively unregisters any existing service worker registrations and clears their caches outside production, instead of just skipping a new registration. (cherry picked from commit 66a2515cbce7a6132639614d88d48349a83bdcde) Co-authored-by: Markus Hartung <mail@hartmark.se> * fix(combo): remove stray brace from #9630 error handling (#9894) Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com> * feat(oauth): add Openference OAuth and API key provider integration (#9869) Wire Openference as a first-party OAuth gateway (PKCE, rotating refresh) and an API-key catalog entry on api.openference.com, with live model discovery, connection testing, free-tier badges, and regression tests. Co-authored-by: Anh Tran <anhlead@outlook.com> * maint: follow-up cherry-pick fix-in-place #9719 (conflict-resolved fallback) (#9893) * fix(db): clear combo pins when connections are deleted * docs: add changelog entry for #9719 --------- Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com> * cherry-pick(pr-9718): feat(src): proxy-pool-toolbar-minor-improvements (#9870) * feat(proxy-pool): streamline pool actions * test(proxy-pool): cover toolbar layout * refactor(settings): extract proxy registry helpers Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(settings): reduce proxy registry component size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Agnes <linkscrazy2@gmail.com> * feat(resilience): expose providerQuotaOverrides via /api/resilience (#9871) Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9712 (conflict-resolved fallback) (#9892) * fix(build): colocateLlmlinguaOptionals skip-check treated a Next-traced stub as fully copied Debugging the omniroute-beta Docker rebuild: `npm run build` (and the Dockerfile's own post-build verification) failed with `Cannot find module '.../node_modules/@atjsh/llmlingua-2/dist/index.js'`. Root cause, reproduced directly (both against a live Docker builder image and in a unit test): Next.js's own standalone trace creates a stub directory for `@atjsh/llmlingua-2` containing only `package.json` — it references the package (a dynamically-imported optional dependency) but can't fully bundle it. colocateLlmlinguaOptionals's skip checks (both the closure-level early return and the per-package loop) only tested `existsSync(dest)`, so that stub was indistinguishable from "already fully co-located" — the function skipped copying the real `dist/` output entirely, silently shipping a package with a manifest but no code. Fix: check for the package's declared `main` entry file when it has one (the real-world case for every actual SLM optional). Packages with no `main` field fall back to comparing the destination's top-level entries against the source's — correct both for genuinely multi-file packages and for a metadata-only source (package.json is then its complete, faithfully- copied contents), which the existing idempotency test exercises. Covered by tests/unit/colocate-optionals.test.ts's new stub-reproduction case (fails against the pre-fix code, passes after — confirmed directly) plus the 6 pre-existing cases, all still green. (cherry picked from commit 359aba59c7b362a5efaa4f0cd4d482d48ee2df66) * fix(build): register onnxruntime-node's native bin/ as a standalone asset (#9687) Docker/standalone builds of the LLMLingua SLM compression tier failed at runtime with "Error: libonnxruntime.so.1: cannot open shared object file: No such file or directory" (open-sse/services/compression/engines/llmlingua's worker, via @huggingface/transformers -> onnxruntime-node). onnxruntime-node's dist/binding.js is a normal JS file Next.js's standalone trace bundles correctly, but binding.js dlopen()s a platform-specific native library shipped under bin/napi-v3/<platform>/<arch>/libonnxruntime.so.1 — a dynamic native load static file tracing can't see (same blind-spot class as the separate colocateLlmlinguaOptionals stub bug, just for a .so instead of a JS import, via NATIVE_ASSET_ENTRIES instead). That directory was simply never registered, unlike better-sqlite3's native binary, which already goes through the exact same mechanism correctly. Fix: add an entry for onnxruntime-node/bin, mirroring the existing better-sqlite3 entry. Confirmed against a real Docker build of the Dockerfile's own post-build verification step: this was the very next failure once the separate llmlingua-2 stub bug was fixed and the build progressed far enough to reach it. Covered by tests/unit/assemble-standalone-onnxruntime-native-asset.test.ts (fails against the pre-fix code on both assertions, passes after). (cherry picked from commit 8c98a59f26a27e844678e673b31c1c21aaf72b0e) --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * maint: follow-up cherry-pick fix-in-place #9707 (conflict-resolved fallback) (#9890) * fix(db): renumber ccr_blocks migration 134 -> 139 134 was taken by 134_proxy_logs_egress_ip, so two migrations shared the same numeric prefix and check-migration-numbering failed. Move ccr_blocks to the next free slot and add the retroactive isSchemaAlreadyApplied guard so a DB that already applied it under 134 skips the re-run. * fix(combo): restore missing preferAntigravityConnectionsWithStoredProject quotaStrategies imported the reset-aware pool filter from ../antigravityProjectPersistence.ts, a module that does not exist — the helper belongs in antigravityProjectPersist.ts and was never added there, breaking typecheck. Add the helper alongside the persist path, point the import at the real module, and cover the filter with unit tests. * chore: add Makefile wrapping the canonical npm scripts * fix(compression): remove duplicate Antigravity project helper The release branch already includes the generic project-aware connection selection helper. Keep that implementation and remove the duplicate introduced while cherry-picking #9707. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Matias Baglieri <168452313+matiasbaglieri@users.noreply.github.com> * cherry-pick(pr-9695): fix(docker): make the webpack build-arg escape hatch actually work (#9872) * build(docker): make the bundler build-arg actually take effect A bare ENV shadows a same-named ARG for the rest of the stage, so --build-arg OMNIROUTE_USE_TURBOPACK=0 was silently ignored and the webpack escape hatch the surrounding comment advertises only ever worked through -e at runtime, never at build time. That mattered because Turbopack compiles in native Rust memory living outside the V8 heap, so OMNIROUTE_BUILD_MEMORY_MB cannot bound it. A build host with a memory ceiling gets SIGKILLed by the cgroup OOM killer with no error text at all, which reads like a hung build rather than an out-of-memory one. * docs(docker): correct the builder stage facts and document its cost The stage table described a builder that no longer exists: it named node:24.15.0-trixie-slim where every stage now derives from node:26-trixie-slim, and said the stage runs `npm run build -- --webpack` where it runs plain `npm run build`, which is Turbopack by default. That second one is worse than stale. A reader who needs the webpack fallback would conclude the Docker build already uses it and never look for the switch. Adds a Build-time resources section covering the two build args, why the V8 heap arg cannot bound Turbopack, and measured ceilings for both bundlers. The runtime paragraphs that followed get their own heading so they no longer read as part of the build-time story. * docs(docker): correct the runtime heap defaults Same drift as the builder stage, in the paragraphs just below it. The image exports OMNIROUTE_MEMORY_MB=1024 and derives NODE_OPTIONS from it, but the guide reported 512 in three places, including the environment variable table. The "if unset, the launcher uses 512" line was misleading in both readings: the image always sets the variable so that branch cannot fire under Docker, and outside Docker the launcher calibrates from host RAM rather than using a flat 512. * docs(changelog): add fragment for #9695 --------- Co-authored-by: Minxi Hou <houminxi@gmail.com> * maint: follow-up cherry-pick fix-in-place #9693 (conflict-resolved fallback) (#9887) * fix(web-tools): anchor tool contract at prompt tail + user-turn reminder The <tool> contract from prepareToolMessages was prepended as the first system message. Web executors fold all system messages into one block, so with agentic clients whose system prompts exceed ~28K chars the contract sat at the head of a huge block and web models ignored it, refusing tool calls with "tool X is not in my tool set" (chatgpt-web, 0/3 at 30K chars). Two changes, both required in testing: - Dual placement: the full contract now rides as a trailing system message (folds to the tail of the system block) and a one-line reminder naming the tools is appended to the latest user message. - Rewording: the contract now frames injected tools as client tools invoked via a plain-text protocol, distinct from the model's native tool registry (web.run, python.exec, ...), and instructs the model to never claim they are unavailable. Without this the model resolved tool names against its native registry and refused even when it had seen the contract. Measured on cgpt-web gpt-5.5-thinking/gpt-5.6-thinking/o3: prepend 0/3 tool calls at 30K chars; dual placement 16/17 across 30K-250K system prompts, 30-tool sets, multi-turn tool history, streaming, and 3-way concurrency, with no spurious calls on no-tool prompts. Known limit: ~40K-char single user messages still flake (2/3) due to the upstream model's own injection heuristics. All prepareToolMessages consumers parse system messages position-independently and select the current user turn by role scan, so the trailing system message is shape-safe for every web executor. * test(web-tools): cover contract placement edge cases --------- Co-authored-by: Ryan Brosas <ryanjoserbrosas@gmail.com> * maint: follow-up cherry-pick fix-in-place #9631 (conflict-resolved fallback) (#9886) * feat(db): add a job registry for scheduled background work Background jobs each ship their own timer today, so there is no list of what is scheduled, no history of what ran, and no way to pause one without an environment variable and a restart. The registry gives them one home: a jobs table holding the schedule, a job_runs table holding the outcomes, and a loopback-only API to inspect and control both. Cron jobs read their expression through an optional cronGetter rather than the stored column, so an operator changing OMNIROUTE_WARMUP_CRON does not need the row rewritten. register() is an idempotent upsert that refreshes the schedule but never overwrites `enabled` or `created_at`, which is what lets a job be re-registered on every boot without discarding the operator's toggle. Run history is pruned per job rather than globally, and safeRun records a failure for a handler that throws as well as one that returns success:false, so a crashing job leaves a trail instead of a gap. The API is under /api/jobs and gated to loopback in the route guard. It can trigger a run and flip a job off, which is runtime administration and does not belong on a remotely reachable surface. Signed-off-by: Minxi Hou <houminxi@gmail.com> * feat(jobs): move the budget reset and token health check onto the registry Both jobs owned their own timer and started themselves as an import side effect, so nothing could report whether they were running, when they last ran, or why a run failed. They now register with the job registry and are started from it, which also means their schedule and run history are visible through /api/jobs. startAll() runs each interval job's first tick synchronously, so both entry points start the registry only after initializeCloudSync() has been awaited. The old wiring reached that ordering two different ways: the budget reset was started after the init call, and the health check's first sweep sat behind a 10s timer. Replacing both with one startAll() would otherwise have moved the two handlers in front of the initialisation they run against. Both entry points also register the same pair of jobs. Registering one and not the other is how a background job goes missing without anything failing. sweep() now returns how many connections it swept, so the health check can record a real records_affected the way the budget reset does. The migration documents that column as a per-job count, and hardcoding zero would have left one of the two jobs reporting a number the schema promises but the code never produces. A skipped or empty sweep reports zero. Every existing caller ignores the return value. The token health check keeps its own disable semantics: the handler still calls isHealthCheckDisabled() before sweeping, so OMNIROUTE_DISABLE_TOKEN_HEALTHCHECK, the production-build phase and the automated-test guard behave as before. Its registry adapter lives in src/lib/jobs/ next to the budget reset rather than in tokenHealthCheck.ts, which is already above its frozen size ceiling on the base branch and should not grow further. The adapter lets a failing sweep throw rather than reporting it itself, matching the budget reset: safeRun records a thrown error as a failure run with its message. The warmup job is seeded disabled. Its handler arrives with the warmup scheduler, and startAll() filters on enabled before it looks for a handler, so seeding it enabled here would warn about the missing handler on every boot. * fix: allowlist cron-parser dep and document OMNIROUTE_RUNNOW_TIMEOUT_MS env var Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> * fix: pass max reasoning effort through by default, add global model registry fallback (#8057) (#9883) Co-authored-by: Mo'men Qatr <momen.qatr04@eng-st.cu.edu.eg> * cherry-pick(pr-9605): ci(test): route orphaned Vitest tests through blocking CI (#9875) * ci(test): route orphaned Vitest tests through blocking CI * docs: fix advisory status in AGENTS.md and refresh baseline note * fix(changelog): fix fragment format for #9415 * fix(changelog): preserve upstream fragment format --------- Co-authored-by: MohitRawat017 <rawatmohit17906@gmail.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> * cherry-pick(pr-9601): feat(responses): add encrypted reasoning replay opt-in (#9876) * feat(codex): add encrypted reasoning replay opt-in * feat(responses): generalize encrypted reasoning replay * docs: clarify encrypted reasoning provider scope * fix(ui): group reasoning replay with connection controls * fix(logs): omit encrypted reasoning payloads * fix(chat): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reduce combined file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: jackjinke <jack.kejin@gmail.com> * cherry-pick(pr-9572): fix(providers): reject the dashboard password as a connection API key (#9877) * fix(providers): refuse to store the dashboard password as a connection API key A browser autofilled the management password into a connection's API-key field. The resulting credential authenticates against nothing, so every request routed through that connection came back 401, and because the field looks like any other password input the same autofill fired again while the connection was being repaired by hand. The refusal belongs on the write path rather than in the form. Twenty routes create or update connections and all of them funnel through createProviderConnection and updateProviderConnection, so one check there covers every entry point including a future one. The two other places that write api_key are left alone on purpose: one re-encrypts rows that already exist and the other is the one-time db.json import, and neither takes a value an operator just typed. Update checks the incoming value, never the merged one. A connection that already holds the password has to stay editable or the operator cannot repair the exact state this prevents, and re-checking the merged value would spend a bcrypt round on every unrelated field edit. Only a real match blocks the write. An unreadable settings row or a throwing bcrypt call logs and allows, because a guard against one specific mistake must not turn into a way to lock out every connection write. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(providers): compare the untrimmed credential, and cover the guard's branches The guard trimmed the incoming value before comparing it, which catches a paste carrying whitespace the password does not have. It missed the mirror case: neither the login route nor the set-password route trims, so a dashboard password may itself begin or end with a space, and an autofill reproducing it exactly was trimmed into a value that no longer matched the stored hash. The write then went through, which is the state this guard exists to prevent. Both forms are compared now, the second only when the first fails on a string that differs, so an ordinary key still costs a single bcrypt round. Two branches carried no coverage and both are load-bearing. The catch that logs and allows is the only path that lets a write through; a stored hash bcrypt cannot parse reaches it without needing a mock, since the shape check accepts an impossible cost factor that the comparison then rejects. The early return is what keeps a token renewal -- a write carrying tokens but no apiKey -- from paying for a settings read and a bcrypt round every time it fires, and the same unparseable hash makes that path observable, so an absent warning is proof the return happened. The narrower scope is deliberate and now says so in the code: the OAuth tokens arrive from a provider's token endpoint rather than from a form, so extending the comparison to them would charge every renewal for a field no autofill can reach. Signed-off-by: Minxi Hou <houminxi@gmail.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> * cherry-pick(pr-9569): fix(settings): use provider prefixes in model overrides (#9878) * fix(settings): use provider prefixes in model overrides * refactor(settings): extract pricing tab helpers Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Xiangzhe <xiangzhedev@gmail.com> * fix: address self-review findings (#9900) Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com> * cherry-pick(pr-9675): fix(providers): per-provider opt-out for anonymous no-auth fallback (opencode-go/zen 401s) (#9873) * fix(providers): add per-provider opt-out for anonymous no-auth fallback API-key providers with anonymousFallback: true (opencode-go, opencode-zen, pollinations, kilocode) receive a synthetic "noauth" connection whenever all real connections are terminal (credits_exhausted/banned/expired) or unavailable. The opencode upstream now rejects anonymous requests with 401 Missing API key, so the fallback adds a guaranteed-failing round trip and health/reconnect noise before the combo moves on. Add a noAuthFallbackDisabledProviders settings array (zod-validated, persisted via /api/settings, following the blockedProviders pattern). When a provider is listed, maybeSyntheticNoAuthFallback returns null for anonymousFallback-only providers, so exhausted providers are skipped immediately as allExpired/allRateLimited while real keyed connections keep working and recover automatically once quota state clears. True no-auth providers are unaffected; blockedProviders remains their disable mechanism. Default (absent/empty list) preserves current behavior. Provider detail pages for anonymousFallback providers gain an "Anonymous fallback" toggle (default ON) backed by the new setting. Refs #9674 * fix(auth): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Hermes Agent <hermes@hermes-chloe.hyades.io> * cherry-pick(pr-9634): fix(test): reconcile base-drifted test expectations on release/v3.8.50 (#9874) * fix(combo): restore routing module load * fix(db): resolve ccr migration version collision Renumber the CCR block-store migration from 134 to 139, reconcile databases that already applied the legacy slot, and add regression coverage for both upgrade paths. Co-Authored-By: GPT-5 <noreply@openai.com> * fix(changelog): format the aggregator balance fragment as a bullet The fragment landed with YAML frontmatter rather than the bullet the aggregator reads, so check:changelog-integrity exits 1 on every branch and takes the merge-integrity job down with it regardless of what the branch changed. Only the format changes. The entry text is the author's, unedited, and now carries the link to the pull request that shipped it. * fix(test): update expected auth/vision/provider schema for base-drifted expectations * fix(test): narrow this branch to the drifted test expectations Three other PRs already cover what this one was carrying. #9618 renumbers the colliding ccr_blocks migration, #9632 repairs the malformed aggregator changelog fragment, and #9676 restores the combo module load by implementing the selection helper the import was reaching for, rather than deleting the caller the way this branch did. Keeping any of it here would put two files back on the same migration slot and overwrite a better fix with a worse one. What survives is the part none of them touch. Once the combo barrel loads again, three assertions in the context-window filter suite start failing: they demand that catalog-too-small targets be dropped, while the file's own header and its four neighbouring tests say those targets stay available as runtime fallback. The unresolved import was masking them. A new case pins the output-token limit as a genuine hard requirement so the relaxation cannot drift further. The provider count assertion kept one literal at the old value after the rest of the file moved to 198, so the partition check failed on a sum that was correct. * chore(quality): re-time migrationRunner for the 139 guard on the new tip --------- Co-authored-by: alexey.nazarov@softmg.ru <alexey.nazarov@softmg.ru> Co-authored-by: GPT-5 <noreply@openai.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> * cherry-pick(pr-9556): fix(translator): preserve Kimi K3 Responses reasoning (#9879) * fix(translator): preserve Kimi K3 Responses reasoning * fix(translator): make K3 reasoning preservation model-driven * fix(translator): replay cached Kimi reasoning before fallback * fix(translator): keep authentic K3 reasoning through cleanup * refactor(reasoning): use replay policy for K3 --------- Co-authored-by: jackjinke <jack.kejin@gmail.com> * maint: follow-up cherry-pick fix-in-place #9510 (fallback resolution) (#9880) * feat(api): add GET /api/resilience/connections for per-account state The three temporary-failure mechanisms each have their own scope -- the provider circuit breaker covers a whole provider, connection cooldown covers one account, model lockout covers a provider/connection/model triple -- and until now nothing showed them side by side. Diagnosing "why is this key being skipped" meant reading three separate surfaces and correlating by hand, which is exactly what the docs' own debugging guidance asks an operator to do. The route returns all three keyed by connection, plus the breaker's transition history so a flapping provider is visible as a sequence rather than a single current state. getStatus() already assembled everything except that history; it now returns a copy of it and carries an explicit CircuitBreakerStatus type instead of an inferred one. Reading raw connection rows for this meant widening getRawProviderConnections' column projection, so the existing allowlist is exported and the route selects through it. A test asserts every column the route names is in that allowlist, which turns a future typo into a failure here rather than a silent empty field. Each of the three data sources is wrapped independently: one of them throwing degrades that section and sets meta.degraded rather than failing the whole response, since a partial view still answers most of the questions the page exists for. Loopback-gated. It spawns nothing, unlike every other entry on that list, but it exposes per-account operational state and the comment says so to keep it from being read as precedent for gating read-only routes generally. Tests are real isolated-DB integration tests rather than mocks -- ESM mocking is unavailable here (no mock.module, non-configurable exports) and the codebase already has the isolated-DB pattern, which exercises more than a mock would anyway. Signed-off-by: Minxi Hou <houminxi@gmail.com> * feat(dashboard): add the per-account resilience connections page Renders what the API added: every connection with its cooldown, its provider breaker, and its model lockouts in one table, with a detail view per connection and the breaker's transitions drawn as a timeline. The timeline is the part that is hard to get from the existing surfaces -- a breaker sitting at CLOSED right now looks healthy, and only the sequence shows it has opened four times in the last hour. Polls rather than streams. The state it displays changes on the order of seconds to minutes and the page is loopback-gated, so an SSE channel would buy nothing over an interval. ModelCooldownsCard had its own formatRemaining. The new table needs the same countdown format and two copies would drift, so it moves to shared/utils/formatRemaining.ts and both import it -- behaviour unchanged, the extracted version differs from the deleted one only in local variable names. DataTable's column and row interfaces are exported for the same reason: the new table types against them rather than restating their shape. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(i18n): translate new resilience-connections screen strings PR #9510 added the "Connection Resilience" dashboard screen but the sync-added i18n keys (sidebar.resilienceConnections/Subtitle and the full resilienceConnections namespace) were left as __MISSING__: in every non-English locale, dropping i18nUiCoverage.pct below the 99 ratchet baseline. Translate all ~78 new leaf strings into all 41 non-English locales. Pre-existing unrelated __MISSING__ debt (hermesRole*, apiProtocol*, grokAutoTopUp*, featureFlagExposeFunctionalGatewayMirrorsDescription) is left untouched — out of scope for this fix. Co-authored-by: HouMinXi <HouMinXi@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Co-authored-by: HouMinXi <HouMinXi@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9549 (conflict-resolved fallback) (#9881) * fix(adobe-firefly): open browser sign-in and resolve provider slug in /login POST /api/providers/[id]/login passed the connection DB id to inAppLoginService.startLogin, but that service looks up the provider by slug in TOKEN_EXTRACTION_CONFIGS. The lookup always missed and returned "No extraction config" without launching a browser — so the VibeProxy "Sign in" button for Adobe Firefly (and every other web-cookie provider) never opened a browser. Adobe Firefly additionally had no extraction config because its IMS JWT is never in cookies/localStorage — it only rides on the Authorization: Bearer header of firefly-3p.ff.adobe.io XHRs. - Resolve the provider slug from the connection row and pass the slug (not the DB id) to inAppLoginService.startLogin. - Add open-sse/services/adobeFireflyBrowserLogin.ts: a Playwright service that launches a visible browser at firefly.adobe.com and intercepts firefly-3p requests to capture the IMS JWT + sherlockToken cookie. Wire it into the /login route for the adobe-firefly slug. - Fix latent bug: updateProviderConnection reads camelCase keys (apiKey, providerSpecificData), so the previous snake_case call never persisted extracted credentials. * fix(adobe-firefly): open browser sign-in and resolve provider slug in /login POST /api/providers/[id]/login passed the connection DB id to inAppLoginService.startLogin, but TOKEN_EXTRACTION_CONFIGS is keyed by provider slug — so browser login never launched for web-cookie providers. Adobe Firefly also cannot use cookie extraction: the IMS JWT only appears on Authorization headers to firefly-3p.ff.adobe.io. Add a dedicated Playwright interceptor and persist credentials with camelCase keys that updateProviderConnection actually reads. * fix(adobe-firefly): use system Chrome/Edge CDP for browser sign-in Playwright is not available inside the pkg-packaged VibeProxyServices.exe, so import('playwright') always failed with 'Playwright not installed' and never opened a window. Launch Chrome/Edge with --remote-debugging-port and capture the firefly-3p Authorization Bearer via pure CDP WebSocket instead. * fix(adobe-firefly): live x-arp-session-id / Arkose wire (stop 408 under load) Browser generate-async requires x-arp-session-id as base64({sid,ark,ftr}) with a real Arkose blob (sherlockToken). JWT alone frequently returns colligo HTTP 408 system under load while credits still work. - Match live ftr magic __UDF43-m4_31ck + Arkose pk in synthetic ARP fallback - Ranked extract of sherlockToken / x-arp from Cookie, HAR, fetch() paste, and space-joined JWT+ARP (PasswordBox newline collapse) - Reuse one ARP for storage upload + generate-async - Clearer 408 errors when browser ARP is missing vs stale - Unit suite 42/42 * fix(adobe-firefly): durable session ARP rebuild and aux_sid false-positive Rebuild x-arp-session-id from forterToken/arkose/ff_session_guid instead of ranking long Cookie pairs (e.g. aux_sid=…) as opaque ARP, which caused colligo HTTP 408. Cache IMS JWT + cookie sessions, rotate ARP on 408 retries, and keep Playwright warm-up opt-in only (headless Forter is rejected). Also expand synthetic ARP shape with bfp/fpjs to match live successful captures. * fix(adobe-firefly): durable session, off-screen Chrome recovery, browser sign-in Rebuild x-arp-session-id from Cookie pieces (sid/ark/forter) so aux_sid is never sent as ARP. Sticky ARP + submit spacing reduce mid-batch colligo 408 thrash. Add optional managed Chrome warm (off-screen headed by default; Forter rejects headless) and POST /api/providers/{id}/login browser sign-in that returns JWT+Cookie after a fresh SSO. Visible sign-in resets off-screen window placement and clears prior Adobe session when adding another account. * fix(adobe-firefly): renew sessions through durable CDP * fix(adobe-firefly): isolate browser sessions per account * fix(adobe-firefly): make account login fresh and deterministic * chore(adobe-firefly): remove obsolete browser fallback * docs(adobe-firefly): document renewal controls * fix(adobe-firefly): harden CDP warm, risk session, and browser sign-in Stop colligo 408 thrash from stale Forter and frozen Google login during Sign in with browser: - CDP warm: clear Firefly origin storage + risk cookies (keep SSO); require forter age under 10 minutes on loop and timeout paths; dual CDP queues; await Runtime.runIfWaitingForDebugger; profile-lock launch retries - Session: connectionId fingerprint; write-back JWT+Cookie; warm-fail cooldown; fail closed risk_session_stale when forter is known-stale - Client: submit gate around generate-async; max 2 attempts when forter known-stale; poll 401 one refresh; pass sessionBrowserKey through handlers - Login route: pure system Chrome/Edge CDP only; camelCase credential persist - Unit: browser-login + firefly suites green (60) --------- Co-authored-by: artickc <artur1992123@mail.ru> * fix(db): resolve ccr migration version collision (#9884) Renumber the CCR block-store migration from 134 to 139, reconcile databases that already applied the legacy slot, and add regression coverage for both upgrade paths. Co-authored-by: fenix007 <fenix007@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9629 (conflict-resolved fallback) (#9885) * fix(compression): add Lite tool truncation toggle * fix(antigravity): add missing antigravityProjectPersistence.ts module The quota-strategy engine (quotaStrategies.ts) imports from antigravityProjectPersistence.ts, but only antigravityProjectPersist.ts existed in the tree. Add the missing module with the expected preferAntigravityConnectionsWithStoredProject() helper and re-export the existing persistDiscoveredAntigravityProjectId(). Co-authored-by: diegosouzapw <diegosouza.pw@outlook.com> * fix(file-size): rebaseline strategySelector.ts for Lite truncation toggle The PR adds one line to threading options?.config?.lite into applyLiteCompression. Update the frozen size from 1060 to 1061. Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Refs #9629 --------- Co-authored-by: Xiangzhe <xiangzhedev@gmail.com> Co-authored-by: xz-dev <xz-dev@users.noreply.github.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9704 (conflict-resolved fallback) (#9889) * fix(sse): persist per-tool-call JSON escape state across SSE delta chunks escapeJsonStringValues() reset its inString/pendingEscape state on every call instead of carrying it forward per tool-call index, so a raw newline byte (or an already-escaped \n) split across two delta chunks got corrupted in transit — the model's own output was correctly escaped, OmniRoute broke it. Root-caused via a dispatched investigation into real OpenClaw traffic that looked like model-generation quality but wasn't. Fix: escapeJsonStringValues now takes and mutates a persistent per-call state object (JsonStringEscapeState), keyed per tool-call index in the translator's init state and cleared when a tool call is superseded. * chore(quality): rebaseline openai-responses.ts for the escape-state fix Own growth from the extracted per-tool-call JSON escape-state fix (previous commit): open-sse/translator/response/openai-responses.ts 1204->1249 (+45). --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * maint: follow-up cherry-pick fix-in-place #9711 (conflict-resolved fallback) (#9891) * fix(sse): grace period before finalizing a client disconnect as 499 (#9653) A client that closes its connection right after reading a fully-completed SSE stream can race OmniRoute's own completion bookkeeping: the bytes already reached the client, but the transform stream's own completion callback (onStreamComplete, which flips streamCompletionRecorded) hasn't finished bubbling up when the disconnect handler fires, so the request gets persisted as a false 499 with zero token usage even though it delivered its full response. Confirmed live on real traffic before this fix: a request whose server log showed "disconnect: request_signal_aborted" at 18236ms was persisted with status 200 and full token usage (82814/1292) once the grace period let the real completion win the race, matching what the client actually received. createClientDisconnectGraceHandler (new leaf in streamFailureFinalization.ts) polls isStreamCompletionRecorded() for up to STREAM_DISCONNECT_GRACE_PERIOD_MS (default 10s, env-configurable, 0 disables) before finalizing as a failure. If a real completion lands within the window, handleStreamFailure's own guard is a no-op and the genuine 200 stands. Covered by tests/unit/stream-disconnect-grace-period-9653.test.ts (fake-timer driven: already-recorded completion short-circuits, disabled-grace-period finalizes immediately, a completion landing mid-window skips finalize entirely, and no completion ever landing finalizes once the deadline passes). (cherry picked from commit 5d0fe28c4246518f7c7b588795a6da1573b5df16) * chore(quality): rebaseline chatCore.ts for the disconnect grace-period fix Own growth from the disconnect grace-period fix: 5030->5039 (+9, the createClientDisconnectGraceHandler wiring at the existing onClientDisconnectFinalize call site). --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * chore: ignore playwright cli artifact dir * maint: final follow-up cherry-pick #9619 (#9901) * fix(quality): clears two release/v3.8.50 base-red gates Unblocks Merge integrity and Docs Gates for every PR against release/v3.8.50, not just this branch: - changelog.d/features/9415-newapi-sub2api-aggregator-balance.md had a non-standard YAML frontmatter header that no other fragment in the tree uses. check-changelog-integrity.mjs reads a fragment's first non-blank line to validate it starts with a markdown bullet; the frontmatter's leading `---` made that check fail regardless of the actual bullet content further down. Removed the frontmatter and reformatted the body to match the documented changelog.d/README.md bullet convention. - docs/ops/VM_DEPLOYMENT_GUIDE.md documented OMNIROUTE_MAX_POOL_SIZE and OMNIROUTE_DB_POOL_SIZE as tunable env vars, but neither is read anywhere in the codebase (confirmed via full-repo grep) — this repo uses SQLite, which has no connection-pool concept these vars could plausibly control. check:fabricated-docs --strict correctly flags fabricated env-var claims; removed the bullet rather than implementing a feature to match invented documentation. * fix(i18n): completes Vietnamese parity, fixes empty migration query Two more release/v3.8.50 base-red items, both surfaced while chasing CI failures on unrelated PRs: - vi.json was missing 8 keys that #9539 (NewAPI/Sub2API aggregator balance) added to en.json without a matching i18n:sync-ui run — pt-BR.json already had all 8, only Vietnamese drifted. Added translations for the 6 provider-settings strings, the feature-flag description, and the quota tooltip; verified against tests/unit/i18n-vi-completeness.test.ts (parity, placeholder preservation, ICU parse — all 5 assertions pass). - src/lib/db/migrations/120_interception_rules.sql was pure comments documenting a no-schema-change key_value namespace, with no executable SQL statement — the migration runner logged "FAILED: 120_interception_rules — Query contained no valid SQL statement" on every fresh DB init. 118_provider_param_filters.sql (same pattern, two migrations earlier) already ends with a bare `SELECT 1;` no-op for exactly this reason; 120 was just missing it. Verified directly against better-sqlite3 that the file now executes without error. * fix(types): clears 6 pre-existing release/v3.8.50 typecheck errors typecheck:core is its own blocking CI job (quality.yml), separate from Docs Gates/Merge integrity. Confirmed pre-existing and unrelated to any current work by branching this worktree directly from upstream/release/v3.8.50 with no other merges applied. - accountSemaphore.ts: isBypassed() already excludes null/<=0 maxConcurrency before ensureGate() is called, but a boolean- …
diegosouzapw
pushed a commit
that referenced
this pull request
Aug 25, 2026
…e2e assertion and 10 integration reds Electron Package Smoke — a packaging defect that had been hidden behind another packaging defect for nine days. Once the loginHeaderCapture fix let the main process start, the server underneath died on 'Cannot find module next': resources/app/server.js shipped without resources/app/node_modules. electron-builder discards the ROOT node_modules in code, not by configuration — app-builder-lib/out/util/filter.js:42 has a hard-coded `if (relative === "node_modules") return false` that runs before any filter pattern. The second extraResources entry pointing INTO ../.build/electron-standalone/node_modules is what sidesteps it, because those relative paths are never equal to "node_modules". #10325 removed that entry as an apparent duplicate and flipped the test to assert "exactly once", freezing the regression as if it were the contract. Restored, and the unit guard now pins both entries — proven by mutation: reverting package.json to the post-#10325 shape fails the guard 3/4, restoring it passes 4/4. group-b-quota-plans-config — the assertion was impossible to satisfy on ANY route, and the page was never broken. layout.tsx hands the whole message catalogue to NextIntlClientProvider, React serialises that prop into the RSC payload, and en.json carries "Internal Server Error" twice, so page.content() always contains it: probing /dashboard, /dashboard/costs, /dashboard/settings and /login showed the string present with every page rendering fine, and a pageerror probe on the failing run captured zero client exceptions. This is the same trap that killed the sibling not.toContain("500") in fc77100 ("raw HTML is unreliable") — that one was removed, this one was kept. Now asserts on rendered text, which still catches a real error boundary. The pageerror capture stays: the CI failure carried no stack trace, which is why it was misread twice. Integration — 10 of the 14 shard-2 reds, all sibling-test gaps behind security fixes: monitoring health now takes a Request and requires management auth (GHSA-mvf8-qc78-5mxm); the OAuth import routes moved to requireManagementAuth (GHSA-mg76) — the test accepts both guard shapes and gained a stronger anchor that every exported handler awaits a guard on its own request, mutation-verified; skill tool names are derived from encodeSkillToolName() and the fake upstream now returns the encoded name so decodeSkillToolName() is exercised too; previous_response_id now fails closed (#10262); proxy_logs persist as an async batch (#11182) so the test flushes first; providerQuotaOverrides joined GET /api/resilience (#9871); the reasoning fixture used a model that stopped being thinking-incompatible, replaced and pinned with a premise assert so it cannot rot silently again. A vacuous assert.ok(true, "all 10 streams completed without hanging") was replaced with real anchors — content must arrive on every stream and the active Timeout count must not grow. Four are deliberately left red rather than aligned, each now tracked: #11551 (the /v1/models after() wiring is dead — the route passes a third argument to a two-parameter function and catalogCache never imports after, so the #8728 contract is unimplemented), #11552 (~27% of requests emit an extra discarded upstream call; the delivered distribution is exactly 0.70, so weighted routing is correct and the waste is the real finding), the fixed-account combo pin (aligning it would destroy the per-step attribution the test exists for), and the web_search fallback already tracked as #11524. Package Artifact — the provenance stamp I added last round used git rev-parse HEAD, which under pull_request is the ephemeral merge commit and therefore never an ancestor of the release branch. Now takes the PR head sha. Refs #10692
TheDemonTuan
added a commit
to TheDemonTuan/OmniRoute
that referenced
this pull request
Aug 26, 2026
* fix(release): restore diegosouzapw#10534 quota recovery and validate the volcengine connect bodies Two base-reds on the v3.8.50 tip, found by the release pre-flight. 1. diegosouzapw#11355 regressed diegosouzapw#10534. It replaced the per-window recovery check with an unconditional `hasActiveCooldown()` stop, which is right for an upstream-derived cooldown but also blocks the case diegosouzapw#10534 exists for: a Claude-subscription 429 persists a SYNTHETIC 1h rateLimitedUntil because the upstream sends no parseable reset. When the later poll shows every governing window has really reset with quota left, holding that synthetic cooldown just deadlocks the connection for an hour. The orphaned `windowStillExhaustedAfterRealReset()` helper and the three unused claudeExtraUsage imports that ESLint flagged were the fingerprint of this regression, not dead code: they are the two halves of the original gate. Re-wired as `isQuotaExhaustedCooldownReleasable()`, deliberately narrow — only lastErrorType "quota_exhausted" is eligible, one still-exhausted or unknown-reset window keeps the lock, and an extra-usage POLICY block stays locked even though its quota windows do look recovered in the same fetch. diegosouzapw#11277/diegosouzapw#11355 semantics are untouched (both guards still pass). Regression guard: tests/unit/provider-limits-recovery.test.ts already pinned this contract and was red on the tip. 15/15 now. 2. The three volcengine-plan connect routes read `request.json()` and handed the raw fields to a headless-browser login service after ad-hoc typeof checks (`check:route-validation:t06`, Hard Rule #7). `String(body.code ?? "")` turned 123 into "123" and an absent code into "", both reaching the service as a plausible SMS code. Now parsed with Zod schemas, before the session lookup, so a malformed body answers 400 instead of a misleading 404. New: tests/unit/volcengine-plan-connect-validation.test.ts (8 cases, red before the fix). Gate: 687 route files scanned, PASS. Also drops a genuinely dead import (formatVideoTimestamp in videoBridge.ts — only used inside the helpers module that defines it). * fix(release): drain the v3.8.50 docs/golden/GLM base-reds and restore a masked assert Second base-red batch from the release pre-flight, measured on the .113 with a clean npm ci (the devbox tree resolves eslint-plugin-react-hooks 7.1.1 from a stray pnpm store instead of the lockfile 7.0.1 and reports 925 phantom errors). Provider count 350 -> 352, one root cause behind three reds. Two providers landed this cycle (volcengine-agent-plan, volcengine-coding-plan) without regenerating the artifacts that quote the count: - docs/reference/PROVIDER_REFERENCE.md regenerated (gen:provider-reference). - README / AGENTS / llm.txt (+42 mirrors) / package.json description / 4 SVG diagrams updated, including the section heading AND the anchor that links to it, so the link does not break. - tests/snapshots/provider/translate-path.json regenerated. The diff is purely additive: 46 insertions, 0 deletions, exactly the two new providers. GLM effort tiers. diegosouzapw#11415 added the explicit glm-5.3-max tier and left two sibling vitest specs pinning the old 16-model inventory and an empty tier list for it. Aligned to the shipped contract (inventory order matches glmProvider.ts; glm-5.3-max declares ["max"]). Test-masking. Four assert reductions surfaced once the deleted-file signal was resolved. Three are legitimate and are allowlisted with their reasoning: diegosouzapw#11355 inverted the startup-cooldown contract (preserve future quota cooldowns), diegosouzapw#11280 replaced two unrolled hops with a 3-hop loop that asserts MORE, and the Gemini 3.5 Flash retirement removed the models those capability asserts described. The fourth was real masking: diegosouzapw#10960 rewrote the oneproxy status test to install a stream mock, immediately overwrite it with a passthrough to the real fetch, and assert `calls.length >= 0` — always true. Restored to assert what the test name claims (the JSON-RPC tools/call carries omniroute_oneproxy_stats and its result reaches the caller), with a scope note that it pins the MCP client contract rather than the commander wiring. Also allowlists the Gemini 3.5 Flash test deletion as _deletedWithReplacement (the model was retired by 2764812; gemini-models-parser.test.ts pins the new "excluded from the parsed list" contract), and rebaselines bundleSize 8045 -> 8461 with per-entry measurements — every entrypoint stays far below its absolute budget. * test(a2a): call the agent-card route handlers with a NextRequest PR diegosouzapw#11418 (S2 topology sanitisation) removed the hardcoded localhost:20128 from both well-known agent-card routes and made them derive the base URL from `request.nextUrl.origin` via `getBaseUrl(request)` (src/lib/wellKnown.ts). That changed the handler contract: `GET` now requires the request Next.js always passes it. Three sibling test files were never aligned and still invoked the handler as a bare `GET()`, so every case blew up with `TypeError: Cannot read properties of undefined (reading nextUrl)` before reaching a single assertion — 8 base-reds from one moved contract, not from a skill-count drift. Align the callers to the shipped contract with a local `makeCardRequest()` helper mirroring tests/unit/security-s1-s2-s4.test.ts (a Request with a defined `nextUrl`). No assertion was removed, loosened or skipped; the assert counts are unchanged and the cases now actually execute. Refs diegosouzapw#11418 * test(router-eval): give spawned CLI children an explicit DATA_DIR The router-eval CLI test spawns the CLI with spawnSync and asserts stderr stays empty. NODE_TEST_CONTEXT is inherited by those children, so since diegosouzapw#10432 (guard diegosouzapw#10428) resolveWritableDataDir() detects a test context with no DATA_DIR and warns on stderr before falling back to a throwaway dir - 194 chars that broke three cases. Pass an isolated DATA_DIR in the child env (the resolution the guard message itself prescribes) instead of loosening the assertions. * test(quota): freeze the clock in the GLM absolute-ISO-reset regression test The diegosouzapw#11353 regression test pins an ABSOLUTE upstream reset instant (2026-08-29 21:01:21) in its production fixture body but measured the resulting cooldown against the real wall clock. The remaining window therefore shrank every day: from 2026-08-25 it dropped under the 5-day floor the two assertions use, and past 2026-08-29 it would parse to null and collapse onto the 24h WEEKLY_QUOTA_COOLDOWN_MS default - a guaranteed future red. The shipped parser (parseIsoDateTimeResetMs / parseDayGranularityResetMs / buildWeeklyQuotaFallback / checkFallbackError) is correct: it returned the real multi-day reset, just measured from today instead of the fixtures NOW. Freeze Date at NOW via node:test mock timers in the two time-dependent cases so they assert the parser rather than the calendar. No assertion weakened, no production code touched. * test(providers): use a non-reserved prefix in the CC-compatible node create cases 93da24c ("fix(providers): reject reserved provider prefixes on compatible-node create/update") made createProviderNodeSchema reject any prefix that is a built-in REGISTRY id or alias. "cc" is the alias of the built-in `claude` provider, so the two provider-nodes create cases in cc-compatible-provider.test.ts started getting a 400 schema rejection before the route ever reached its feature-flag gate (403) or the create path (201) — the guard PR updated its own tests but missed this sibling file, leaving a base-red on release/v3.8.50. The operator-chosen prefix is incidental to what these cases assert (the ENABLE_CC_COMPATIBLE_PROVIDER gate, the dedicated anthropic-compatible-cc- id prefix, baseUrl sanitization and the nulled modelsPath), so switch it to a non-reserved "cc-proxy". No assertion was removed or loosened. * fix(quality): drain the three v3.8.50 inventory/coverage base-reds All three guards were drifting behind legitimate cycle growth, not catching a defect. Nothing was weakened: no assertion removed, no floor lowered, no blanket-allow added. providers-constants-split: APIKEY_PROVIDERS 231 -> 233. The delta is exactly the two Volcano Ark plan providers (volcengine-agent-plan, volcengine-coding-plan) added to the regional family in d732cf6. The invariant the guard exists for still holds, measured on the tip: 233 merged keys, 233 unique, family sum 233 (gateways 92 + frontier-labs 25 + inference-hosts 29 + enterprise-cloud 17 + regional 43 + specialty-media 27) with an empty cross-family duplicate set and an empty symmetric difference between the merged object and the family union - so the six files are still a strict partition, no loss and no dup. openapi-coverage: the operation floor (34.6%) is untouched. The cycle grew the denominator 985 -> 1002 while covered only moved 343 -> 345 (34.4%). Fixed by DOCUMENTING five real public operations rather than moving the floor, taking it to 350/1002 = 34.9%: GET /api/health, GET /api/v1/voices, POST /api/v1/speech-to-text, POST /api/v1/text-to-speech/{voiceId} and GET /api/v1/explain/routing. Each entry was written from the route source (auth mode, path-param pattern, limit clamp, upstream relay behaviour and the 400 / 401 / 429 branches), not from memory. hard-session-lease-bypass-inventory: three new connection-query sites classified, none silenced. open-sse/services/combo.ts (readConnectionForCooldownGate) reads the row backing the pre-dispatch persisted-cooldown gate, so it sits on the routing path and joins the class-B list next to combo/providerWildcard.ts and autoComboCandidates.ts. src/lib/providers/volcenginePlanBinding.ts and src/lib/providers/volcPlanAutoSyncBackfill.ts are connection persistence, not dispatch - the first resolves update-vs-create during connect, the second is a one-shot boot backfill of a providerSpecificData flag with no upstream call - so both stay class C alongside oauth/connectionPersistence.ts. * fix(release): drain two v3.8.50 base-reds (volcengine vision metadata, antigravity BYOP contract) Two unrelated real reds on the release tip: * fix(providers): flag MiniMax M3 as multimodal on the Volcengine Ark plans. d732cf6 ("feat(volcengine): add Ark plan providers") added volcengine-agent-plan/minimax-m3 and volcengine-coding-plan/minimax-m3 without supportsVision, breaking the LEDGER-4 invariant that every minimax-m3 registry entry except PromptQL (text-only upstream) is flagged multimodal. Every other provider carrying the model (opencode-zen, opencode-go, bazaarlink, ollama-cloud, codebuddy-cn, trae) sets it. Registry metadata defect, not a stale test. * test(antigravity): align the empty-projectId onboarding test to the contract shipped by diegosouzapw#11284/diegosouzapw#11358 (6de542b). That change made an onboardUser 200 whose body carries NO cloudaicompanionProject mean Google BYOP — no project was created and none ever will be — so it short-circuits before the retry loadCodeAssist. The older test still mocked onboardUser with the bare { done: true } BYOP shape while asserting the retry path, so it pinned a contract that was deliberately moved. The mock now returns a real onboarding-success body; every assertion is kept, and the id in the onboard body deliberately differs from the expected one so the test still proves the projectId came from the retry discovery. Refs diegosouzapw#11284 * fix(cli): do not treat a tmpfs mount as proof a config path reaches the host hasBindMountAt() accepted ANY mount as evidence that a would-be CLI config write reaches the operator's host: it matched on the mount point alone and never looked at the filesystem type. An in-memory mount therefore cleared the ephemeral flag, so guardCliConfigWrite() let the write through and both POST /api/cli-tools/apply and the dashboard's guide-settings writer answered 200 instead of the safe 422 that diegosouzapw#10057 added. That is the exact case the guard exists to refuse, and the worst one: a container running with `--tmpfs /tmp` (or a home on tmpfs) loses the file even before the container is recreated, while the UI reports success. Parse the filesystem type from mountinfo (the field after the lone "-" separator) and skip mounts backed by RAM or kernel state. Real bind mounts (ext4/xfs/nfs/virtiofs/fuse.*) still count, including one nested under a tmpfs path, so the compose `host` profile is unaffected. A line carrying no separator proves nothing and is skipped too. Regression cover added to tests/unit/container-env-detect.test.ts; this also un-reds tests/unit/cli-tools-apply-container-422.test.ts and tests/unit/api/cli-tools/apply-container-guard.test.ts, which were failing on any box whose /tmp is a tmpfs. * chore(docs): commit the next-dev agent-rules block into AGENTS.md `next dev` writes and re-adds this block (see node_modules/next/dist/server/lib/generate-agent-files.js), so leaving it out of a diff only recreates the uncommitted change on the next dev run. Committing it keeps the working tree clean, which is what the block's own note prescribes. * docs(changelog): aggregate the 363 changelog.d fragments into [3.8.50] Release reconciliation (Phase 0a.1). `scripts/release/aggregate-changelog.mjs` folds each changelog.d/<section>/*.md fragment into its heading in the living [3.8.50] section and deletes the fragment, which is the whole point of the fragment convention: two PRs never touch the same file, so the CHANGELOG never conflicts mid-cycle and no bullet is eaten by a merge auto-resolve. Section bullets 731 -> 1041. The remaining uncovered commits (mostly merges from diegosouzapw#11397 onward, which landed without a fragment) are reconciled separately. * docs(changelog): reconcile the v3.8.50 section with the cycle's uncovered commits Adds 131 consolidated bullets (45 features, 71 fixes, 15 maintenance) covering the ~490 user-facing commits and the ~100 chore/ci/test/refactor/docs commits that landed in the cycle without a CHANGELOG entry, grouped by subsystem and citing their PR references. Uncovered report: 594 -> 175 (the remainder are commits carrying no #N in their subject, which the matcher can never resolve; they are covered in prose). * docs(changelog): date the 3.8.50 header, inject contributors and sync the 42 i18n mirrors * fix(ci): raise the unit-shard heap ceiling to 8192 MB to stop the SIGABRT OOM The 8 unit shards run under V8 coverage instrumentation, which retains far more memory than the bare suite. With the 4096 MB ceiling they began aborting with exit 134 ("Ineffective mark-compacts near heap limit") at ~4086 MB as the provider catalogue grew during the v3.8.50 cycle: every test in the shard passed and the process died at the end, which reads as a test failure without being one. Aligns test:unit:ci:shard and test:unit:serial with the 8192 MB the non-sharded variants already use. GitHub-hosted runners have 16 GB, so the headroom is real. Validated by the CI run on this commit — the shards are the gate. Refs diegosouzapw#10692 * fix(ci): set NODE_OPTIONS on the unit-shard step and prune stale eslint suppressions The previous heap bump only touched test:unit:ci:shard, i.e. the node the shard script spawns. The process that actually runs out of memory is the `c8` wrapper around it — it aggregates ~577 MB of raw V8 coverage JSON — so the ceiling stayed at the V8 default (~4 GB) and the shards kept aborting at ~4083 MB, byte for byte the same failure. Setting NODE_OPTIONS on the step covers c8 and every child, which is the pattern the coverage-merge job already uses. Also prunes three eslint suppression entries whose violations no longer exist: videoBridgeContactSheet.ts and videoBridgeRuntime.ts (no-unused-vars, fixed during this cycle) and cli-oneproxy-commands.test.ts (no-explicit-any 14 -> 13, a consequence of restoring the real mock in that test). Stale entries make `npm run lint` exit 2 with 'There are suppressions left that do not occur anymore'. Pruned and verified on an uncontaminated checkout, not the devbox. Refs diegosouzapw#10692 * fix(ci): drain four inherited base-reds in packaging, electron and integration tests All four predate this session — each reproduces identically on f95b03d (2026-08-24), so none is a cycle regression. Draining them here because the release pre-flight is where inherited reds get resolved. 1. Package Artifact: the job runs `build:cli`, which assembles dist/ but never writes dist/BUILD_SHA — only `build:release` does, via write-build-sha.mjs. The diegosouzapw#10427 provenance guard inside check:pack-artifact then rejects the artifact as untraceable, and rejects it even under OMNIROUTE_ALLOW_CANARY_BUILD. The job's build+validate pair was structurally incompatible and failed 100% of the time. Stamps the SHA between the two steps. 2. Electron Package Smoke: electron/package.json's build.files allowlist enumerates each lib/*.js by hand and never got lib/loginHeaderCapture.js, added alongside its require() in diegosouzapw#9984. The file therefore stayed out of app.asar and the packaged app died at startup on 'Cannot find module ./lib/loginHeaderCapture'. 3. proxy-pipeline: the breaker assertion grepped chat.ts for executeChatWithBreaker(, but that call moved behind the chatDispatch.ts seam. Rather than drop the check, it now pins both hops — chat.ts dispatches through the seam and the seam calls the breaker — so the extraction cannot silently take the breaker off the path. 4. skills-pipeline: diegosouzapw#9058 began encoding skill tool names as omr_skill_<base64url> because providers require ^[a-zA-Z0-9_-]+$, and these assertions still expected the raw name@version. They now derive the expected name from encodeSkillToolName(), the same helper production uses, so the test tracks the contract instead of duplicating it. Only the assertions about names on the wire were converted; the identifiers passed straight to skillExecutor.execute() stay raw, because those are not encoded. Integration suite for these two files: 54/55. The one still red — 'web_search fallback preserves Responses API output' — is a separate pre-existing defect, deliberately left failing rather than papered over: on the /v1/responses path resolveSearchCredentials() returns null for the seeded serper-search connection, so executeWebSearch.ts:185-200 falls through to the cheapest fallbackOnly provider (duckduckgo-free) and the results come back empty. The sibling chat-path test seeds identically and does resolve serper-search. Needs its own investigation. Refs diegosouzapw#10692 * test(vitest): realign five sibling-test contracts unmasked by the green mcp shard None of these are cycle regressions. The Vitest job runs test:vitest (mcp shard) then test:vitest:ui; the mcp shard was failing on a missing glm-5.3-max and aborted the job before the ui shard ever ran. Fixing that shard this cycle unmasked 34 ui failures that had been broken since 18-23 Aug — four separate PRs that moved a contract and updated their own tests but not their siblings. - ProviderCard gained useRouter() in diegosouzapw#10448; four test files render it without mocking next/navigation and died on 'invariant expected app router to be mounted'. The sibling created alongside diegosouzapw#10448 already had the mock — it just was not applied to the other four consumers. - SkillCoverage gained a required config category. The four fixtures in agent-skills-page still described only api/cli, so the component read config.have off undefined. Values were chosen per scenario rather than pasted: full coverage gets 2/2 so its bar stays emerald, the amber fixture gets 3/4 so it stays amber. CoverageBar renders api -> config -> cli, so the new bar lands in the MIDDLE and the cli assertions moved from index [1] to [2]; without that the cli checks would have passed while measuring the config bar. The aria test now pins all three bars. - CliAgentsPage hardcoded AGENT_IDS, which had already drifted once (6 -> 8 with omp/letta, per its own comment) and drifted again with prime-agent (diegosouzapw#11166). It is now derived from CLI_TOOLS. This is why an agent missing from that list is not cosmetic: it never enters the status map, defaults to not_installed, and adds a phantom card to the filter and count tests. Deriving keeps the fixture in sync by construction instead of waiting for the next agent. - claudeTlsClient asserted proxyUrl was undefined inside a test literally named 'falls back to env var when per-call proxyUrl not provided' — it pinned the old behaviour where testOverride bypassed proxy resolution. diegosouzapw#10910 moved resolution ahead of the override on purpose ('so test overrides and the real path both see it'), so the assertion now checks the fallback the test name promises. test:vitest:ui goes from 34 failures to 14. The remaining 14 sit in six files none of this commit touches (AutoComboCatalog, CoolingConnectionsPanel, ProxyRegistryManager x2, connectionsSearchFilter) plus one claudeTlsClient case that passes in isolation and only fails in the full run — i.e. cross-file pollution. They need a clean environment to judge: this devbox resolves part of its tree through a stray pnpm store and has already produced one phantom failure count this cycle. Refs diegosouzapw#10692 * fix(i18n): unstale trafficInspectorSubtitle across 32 locales and regenerate the omni-inference skill Two gates in the Lint job, both inherited — each was hidden behind the one before it. i18n value drift: diegosouzapw#11283 rewrote sidebar.trafficInspectorSubtitle in en.json without touching the 42 translations, so 32 locales kept serving a sentence the English no longer says. Most take the documented __MISSING__: placeholder, which makes the runtime serve the corrected English until the translation pipeline catches up. Three do not: - vi cannot take a placeholder at all — tests/unit/i18n-vi-completeness.test.ts bans any __MISSING__/__TODO__ value outright, so it needs a real translation. - pt and pt-BR are translated for real rather than placeheld, because a placeholder there means this project's own maintainer reads the sidebar in English. Each file changed by exactly one line; the JSON was not reserialised wholesale. agent-skills-sync: skills/omni-inference/SKILL.md was missing the ElevenLabs voices and speech-to-text routes added by diegosouzapw#11312, so the generator reported one file out of date and the gate exited 2. Regenerated — purely additive, 48 lines, no deletions. Verified: check-ui-value-drift PASS, i18n:check-ui-coverage PASS (42 locales), i18n-vi-completeness 5/5, check:agent-skills-sync 46 unchanged. Refs diegosouzapw#10692 * test(vitest): pay heavy module imports at collection time, not out of the per-test budget The remaining ui-shard reds were one class, not six bugs: every one of them did `await import(<heavy component>)` INSIDE an `it()`, so Vite's transform of the dependency graph was billed to that test's timeout. Measured costs against the budgets they had to fit in: ProxyRegistryManager 86s import vs 30s / 60s / 5s budgets (render itself: 567ms) claudeTlsClient ~12s import vs 5s default useProviderConnections 1050-line hook, whole dashboard graph, vs 5s default That is why they looked like cross-file pollution: on an idle box the import squeaked under the limit, and under the ui suite's 20 parallel workers it did not. Running claudeTlsClient ALONE on a loaded box reproduces it — the trigger is CPU contention, not a neighbouring file. The sibling chatgptTlsClient/grokTlsClient tests import the same graph and never fail, because they import statically at module scope, where the cost falls on the collection phase which has no per-test budget. Every fix here does the same: static import or a beforeAll with its own budget. AutoComboCatalog also explains its own blast radius: the timeout aborted inside an open act(), leaking an unbalanced act scope that then failed the file's three remaining tests in ~20ms with 'overlapping act() calls'. One slow import, four reds. CoolingConnectionsPanel is the one production change. It imported providerText from the ../providerPageHelpers barrel, but that symbol is DEFINED in the ../providerCredentialText leaf and only re-exported by the barrel — which drags providerRegistry (352 providers) and the rest of the provider-page graph into a "use client" component for one string helper. Verified before accepting: the component used nothing else from the barrel, the barrel has no top-level side-effect to lose (the empty-registry hazard this repo has hit before does not apply), typecheck:core is clean, and the panel's first test drops from ~4s to 95ms. The import was suboptimal, never broken — the screen was not failing for users. No assertion was weakened anywhere. expect() counts are unchanged (25/25, 4/4) or up by one (AutoComboCatalog 11 -> 12); the diegosouzapw#8855 autofill sentinels, the data-1p-ignore / data-lpignore guards and the dead-status round-trip are intact. The diegosouzapw#5918 TDZ guard was proven still live by mutation, not by absence of red: moving useProxyBatchOperations(load) above its const reproduced 'ReferenceError: Cannot access load before initialization' in 207ms, then the production file was restored (diff empty). tests/unit/ui under load: 17 failed files / 45 failed tests -> 4 failed files / 4 failed tests, none of them these. The four left are compression-guidance-7530, compressionPanel, compressionUltraTier and lobe-provider-icons-stepfun, untouched and uninvestigated. Refs diegosouzapw#10692 * test(integration): realign three suites to security and version contracts that moved All three are the sibling-test gap again: a PR moved a contract, updated its own tests, and left these behind. None is a production defect — in two of the three the production side is a deliberate security fix. v1-contracts-behavior (4 failures, one cause): the job env sets INITIAL_PASSWORD, which makes isAuthRequired() true, and diegosouzapw#9320 (b07182c) made the /v1 catalogue gate on-by-default instead of opt-in via settings.requireAuthForModels. The four contract reads were calling the catalogue routes with no credential and correctly getting 401. Bisected the job's four env vars to confirm INITIAL_PASSWORD alone reproduces it (5 pass / 4 fail with it, 9 / 0 without). The tests now send a Bearer token; the shape assertions are untouched, and the auth contract itself stays owned by tests/unit/v1-models-auth-leak-9320.test.ts rather than being duplicated here. opencode-config-startup: two independent drifts. OPENCODE_VERSION was pinned to 1.18.8 while the installed opencode-ai is 1.18.18 (Dependabot 7f69589, diegosouzapw#10626) — now read from require("opencode-ai/package.json").version, which is exactly as strict but cannot drift on the next bump. And the no-limit-metadata case asserted limit === undefined, but diegosouzapw#11054 made the generator always emit a limit; it now pins the actual fallback {context: 128_000, output: 8_192} instead of an absence. memory-pipeline: diegosouzapw#11040 (GHSA-cpv3-xr7r-xf8q) made the resolved caller principal always win over a caller-supplied apiKeyId, so a spoofed id can no longer write into another principal's store. That PR updated the unit sibling but not this one. The test now asserts the stronger property — and deliberately not just the absence: the spoofed principal's store is empty AND the caller can still read the entry, which proves the write was redirected rather than dropped and keeps the emptiness check from passing vacuously with a disabled store. (The old assertion was count === 0, which a switched-off memory store would satisfy.) Assertion counts: 43 -> 43, 13 -> 14, 76 -> 81. Nothing weakened or removed. Verified: 24/24 pass, with and without the CI env vars. Refs diegosouzapw#10692 * fix(ci): drain the electron packaging regression, the models-catalog e2e assertion and 10 integration reds Electron Package Smoke — a packaging defect that had been hidden behind another packaging defect for nine days. Once the loginHeaderCapture fix let the main process start, the server underneath died on 'Cannot find module next': resources/app/server.js shipped without resources/app/node_modules. electron-builder discards the ROOT node_modules in code, not by configuration — app-builder-lib/out/util/filter.js:42 has a hard-coded `if (relative === "node_modules") return false` that runs before any filter pattern. The second extraResources entry pointing INTO ../.build/electron-standalone/node_modules is what sidesteps it, because those relative paths are never equal to "node_modules". diegosouzapw#10325 removed that entry as an apparent duplicate and flipped the test to assert "exactly once", freezing the regression as if it were the contract. Restored, and the unit guard now pins both entries — proven by mutation: reverting package.json to the post-diegosouzapw#10325 shape fails the guard 3/4, restoring it passes 4/4. group-b-quota-plans-config — the assertion was impossible to satisfy on ANY route, and the page was never broken. layout.tsx hands the whole message catalogue to NextIntlClientProvider, React serialises that prop into the RSC payload, and en.json carries "Internal Server Error" twice, so page.content() always contains it: probing /dashboard, /dashboard/costs, /dashboard/settings and /login showed the string present with every page rendering fine, and a pageerror probe on the failing run captured zero client exceptions. This is the same trap that killed the sibling not.toContain("500") in fc77100 ("raw HTML is unreliable") — that one was removed, this one was kept. Now asserts on rendered text, which still catches a real error boundary. The pageerror capture stays: the CI failure carried no stack trace, which is why it was misread twice. Integration — 10 of the 14 shard-2 reds, all sibling-test gaps behind security fixes: monitoring health now takes a Request and requires management auth (GHSA-mvf8-qc78-5mxm); the OAuth import routes moved to requireManagementAuth (GHSA-mg76) — the test accepts both guard shapes and gained a stronger anchor that every exported handler awaits a guard on its own request, mutation-verified; skill tool names are derived from encodeSkillToolName() and the fake upstream now returns the encoded name so decodeSkillToolName() is exercised too; previous_response_id now fails closed (diegosouzapw#10262); proxy_logs persist as an async batch (diegosouzapw#11182) so the test flushes first; providerQuotaOverrides joined GET /api/resilience (diegosouzapw#9871); the reasoning fixture used a model that stopped being thinking-incompatible, replaced and pinned with a premise assert so it cannot rot silently again. A vacuous assert.ok(true, "all 10 streams completed without hanging") was replaced with real anchors — content must arrive on every stream and the active Timeout count must not grow. Four are deliberately left red rather than aligned, each now tracked: diegosouzapw#11551 (the /v1/models after() wiring is dead — the route passes a third argument to a two-parameter function and catalogCache never imports after, so the diegosouzapw#8728 contract is unimplemented), diegosouzapw#11552 (~27% of requests emit an extra discarded upstream call; the delivered distribution is exactly 0.70, so weighted routing is correct and the waste is the real finding), the fixed-account combo pin (aligning it would destroy the per-step attribution the test exists for), and the web_search fallback already tracked as diegosouzapw#11524. Package Artifact — the provenance stamp I added last round used git rev-parse HEAD, which under pull_request is the ephemeral merge commit and therefore never an ancestor of the release branch. Now takes the PR head sha. Refs diegosouzapw#10692 * test(ui): unmount before asserting so the auto-sync timer cannot outlive the test Vitest went red on 'synchronizes upstream models only when autoFetchModels is explicitly true' with new URL throwing inside a fetch dispatched from Timeout._onTimeout (useProviderModels.ts:69). It is intermittent: green in the two previous CI runs, green every time in isolation, red only under the ui suite's 20 parallel workers. The hook schedules its auto-sync in a setTimeout whose callback only checks the flag on entry — and that flag stays false while the component is mounted. Both tests asserted first and unmounted last, so under contention the timer escaped the test window, fired after afterEach had already run vi.unstubAllGlobals(), and reached the REAL fetch with a relative URL. Unmounting before the assertions closes the window: cleanup flips , the callback returns early, and the calls already recorded on fetchMock are still there to assert against. No assertion changed. Not a regression from this cycle — the file's last change is fd76271 (diegosouzapw#10603). Fixed rather than tracked because an intermittent red in a blocking job is worse than a permanent one: it teaches people to re-run instead of to look. Refs diegosouzapw#10692 * fix(search): prefer the configured search connection over duckduckgo-free When no explicit provider is requested and the auto-selected cheapest provider has no credentials, executeWebSearch ran the fallbackOnly loop first. duckduckgo-free (costPerQuery 0, authType none) always won there with an empty credentials object, so a configured paid connection such as serper-search was silently ignored and the caller got success:true with zero results. Move the sweep for other credentialed regular providers ahead of the fallbackOnly loop (and exclude fallbackOnly ids from it, so a free last-resort provider never outranks a configured one on cost). The fallbackOnly loop stays as the true last resort. The chat path only appeared correct because duckduckgo-free happened to fail there and handleSearch retried the alternate provider; on /v1/responses it "succeeded" with no results. Closes diegosouzapw#11524 * fix(api): restore the after() injection point for the /v1/models SWR refresh `/v1/models` has passed a third argument to `getUnifiedModelsResponse()` (`{ scheduleBackgroundRefresh: (task) => after(task) }`) ever since diegosouzapw#10198, but diegosouzapw#9199 had already removed the parameter: the function takes two, so the object was silently dropped and `catalogCache` kept scheduling the stale-while- revalidate rebuild with `setTimeout(..., 0)`. The builder is overwhelmingly synchronous under the single-threaded App Router, so it pinned the event loop before the stale response was flushed — the diegosouzapw#8728 guarantee did not exist. - `catalogCache` now imports `after` from `next/server` and exposes `defaultBackgroundRefreshScheduler`, which defers to `after()` and falls back to a macrotask outside a Next request scope (instrumentation warm-up, tests). - `resolveCachedCatalogResponse` takes `scheduleBackgroundRefresh` and `getStaleWhileRevalidateMs` on its existing options object. - `getUnifiedModelsResponse` accepts the options object the route already passes and propagates the scheduler. The excess argument was invisible to CI: `tsconfig.typecheck-core.json` is a curated 27-file allowlist, `check:dashboard-typecheck` only covers `src/app/(dashboard)`, and `next.config.mjs` sets `ignoreBuildErrors: true`. Closes diegosouzapw#11551 * fix(sse): stop re-summarizing a universal handoff that never parses A universal handoff whose summary comes back unparseable persists nothing, so shouldGenerateUniversalHandoff keeps answering "generate" and the very next model switch in the same session re-issues the same full-history summarization call and discards the answer again — forever. With a switch-heavy combo strategy (weighted, random, round-robin, p2c) the models alternate on almost every turn, so that background call lands on a large fraction of requests: an upstream call whose response nobody reads is real money on a paid provider and real quota on a metered one. Measured on the weighted 70/30 matrix, that inflated the observed openai share to 0.895 where routing actually delivered 0.70, and at the unit level 199 of 200 model switches issued a fresh discarded summarization call. Back off per (session, combo) after an answer that is not a usable handoff: exponential 5min -> 1h, cleared on the first successful generation, capped at 500 tracked keys. Deliberately narrow — a transient upstream failure (!response.ok) is NOT tracked, so it still retries on the next switch, which is the behavior the context-relay path already depends on. After the fix, same harness at n=200: 201 upstream calls for 200 requests (1 extra, 0.5%), and the measured openai share equals the delivered share (0.725). The unit guard drops 199 discarded calls to 1. Refs diegosouzapw#11552 * fix(sse): honor per-step connection pins instead of rotating accounts on fallback A combo step pinned with an explicit `connectionId` (or a request pinned via `x-omniroute-connection`) is an operator instruction, not a hint. The generic account-fallback branch in handleSingleModelChat excluded the pinned connection after an upstream failure and re-selected a sibling account of the same provider, so a priority combo repeating one provider/model with two different fixed accounts ran both attempts under the FIRST step: the second step, with its own pin, never executed and per-step attribution (comboStepId / comboExecutionKey) was wrong. Gate the rotation on `!hasForcedConnection`, matching the antigravity stream-readiness, pre-response-timeout and account-semaphore branches that already let pinned steps fall through to combo orchestration. Cooldown recording via markAccountUnavailable is unchanged, and unpinned selection still excludes burned connections. Refs tests/integration/combo-routing-e2e.test.ts * fix(ci): point the pack-artifact provenance gate at the branch under test The guard added for diegosouzapw#10427 checks ancestry against origin/main by default. That is the right ref at publish time (npm-publish.yml runs on main), but in a pull_request context it can never hold: while the PR is open its head is by construction not an ancestor of main, and the shallow checkout does not even bring origin/main into the local graph, so the probe answers false regardless. The job therefore failed 100% of the time and only became visible now that it stopped being cancelled behind Build. Pre-merge the one checkable invariant is that the stamp matches the branch under test, so resolve the ref from refs/pull/<N>/head — which exists on origin even for fork PRs, unlike head.ref, which only exists on the author's repository. * test(dashboard): open the proxy toolbar overflow menu before bulk assign PR diegosouzapw#9870 moved the proxy-registry bulk-assign action into the toolbar's "More actions" (…) overflow menu, which renders its items only while open. The e2e smoke flow still clicked the testid directly, so the locator never resolved and the test burned its full 180s budget. Open the menu first and assert it is visible before clicking; every existing assertion is unchanged. * fix(dashboard): restore expert-mode manual model entry in the combo builder PR diegosouzapw#8285 (global model search) replaced the expert-mode "Manual model" block positionally with the new GlobalModelSearchPanel, dropping the only way to type a provider/model pair by hand in expert mode. The supporting state and handlers (manualModelInput, manualModelError, manualModelHasDuplicate, handleAddManualModel) survived as dead code, so neither typecheck nor lint flagged the loss. Re-render the block above the search panel, unchanged from its pre-diegosouzapw#8285 form. Regression guard: tests/e2e/combos-flow.spec.ts "expert mode shows a single-page combo form with manual model entry", which had been failing with a 180s timeout on locator.fill for combo-manual-model-input. * test(api): accept the post-diegosouzapw#9320 auth gate on the /v1/models e2e check b07182c (diegosouzapw#9320) inverted the catalog auth rule: /v1/models now requires auth whenever management auth is configured, unless requireAuthForModels is explicitly false. The e2e harness boots with INITIAL_PASSWORD set, so the endpoint has been answering 401 since 2026-08-04 and this check has been red ever since — invisible only because the job kept being cancelled behind Build. Mirror the sibling /api/providers check in this same file: assert the catalog shape whenever the catalog is actually served, and otherwise pin the auth gate by status AND error type, so a 401 from an unrelated misroute cannot pass for the deliberate one. --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: Xiangzhe <diegosouza.pw@gmail.com> Co-authored-by: Xiangzhe <bakryun0718@proton.me> Co-authored-by: TheDemonTuan <nguyenviettuanbp@gmail.com>
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…iegosouzapw#9871) Co-authored-by: herjarsa <herjarsa@users.noreply.github.com>
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
… A2A TaskManager (PRD RF1) (diegosouzapw#8080) * fix(api): enforce model permissions on gateway mirrors (#9854) Co-authored-by: Xiangzhe <xiangzhedev@gmail.com> * cherry-pick(pr-9787): fix(sse): apply Azure param rules on azure-ai and clamp gpt-4o-mini output tokens (#9855) * fix(sse): apply Azure request-param rules on the azure-ai wire path Azure rejects several stock Chat Completions params on its newer deployments and returns HTTP 400 rather than ignoring them: max_tokens -> 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead. reasoning_effort -> Function tools with reasoning_effort are not supported. Those rules lived inline in AzureOpenAIExecutor, so they only covered the azure-openai provider. azure-ai (Azure AI Foundry) had no executor entry and fell through to the bare DefaultExecutor, so the SAME Azure deployment succeeded on one connection and 400'd on the other. Every agentic client sends tools on every turn, so azure-ai failed on the first request. Extract the rules to open-sse/executors/azureParamRules.ts, add an AzureAiExecutor that inherits DefaultExecutor's azure-ai URL/header/apiType handling unchanged and applies the shared rules, and register it for azure-ai. Also widen the deployment pattern to cover gpt-chat-latest: it is a moving alias that resolves to a GPT-5-era model and rejects max_tokens, but carries no version number for the token-boundary pattern to key on. Verified against the base regex - gpt-chat-latest did not match, which is exactly the observed 400. Regression guard: tests/unit/azure-param-rules.test.ts, including an assertion that getExecutor("azure-ai") no longer resolves to a bare DefaultExecutor. * fix(sse): clamp Azure gpt-4o-mini completion tokens to its 16384 ceiling Azure gpt-4o-mini deployments accept at most 16384 completion tokens and 400 on anything larger: max_tokens is too large: 32000. This model supports at most 16384 completion tokens, whereas you provided 32000. The 32000 is OmniRoute's own doing: adjustMaxTokens raises any smaller max_tokens to DEFAULT_MIN_TOKENS (32000) whenever tools are present, to avoid truncated tool arguments. That floor has no upper bound, so an agentic client asking for far less still trips the model ceiling on its first turn. Add scoped maxOutputCap rules in paramSupport.ts for both Azure wire paths. PROVIDER_MAX_TOKENS is the wrong lever here - it is provider-wide, and the same Azure resource also serves GPT-5 deployments with a much higher ceiling. Regression guard: tests/unit/azure-max-output-clamp.test.ts, which also pins that the clamp does not leak to gpt-5.1 or to gpt-4o-mini on other providers. --------- Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * maint: final follow-up cherry-pick #9783 (#9904) * fix(deps): bump transitive deps for 6 Dependabot + remaining audit vulns on main Same overrides as #9464 (ip-address, hono, fast-uri, socket.io-parser, undici) applied directly to main. Also covers brace-expansion (scoped), js-yaml v4 copies, and mermaid. npm audit: 6→0 vulnerabilities. Closes Dependabot #161-#166. * fix(deps): bump nanoid, dompurify for 2 new Dependabot alerts (#189, #190) Bumps: nanoid ^3.3.17 (was transitive, now overridden), dompurify ^3.4.13 (with monaco-editor scoped override). Closes Dependabot #189, #190. Remaining #182-#188 (js-yaml + mermaid) already closed by #9651 merge — awaiting Dependabot re-scan. npm audit → 0 vulnerabilities. * fix(repo): harden .gitignore to also ignore a _tasks symlink (/_tasks) _tasks is a SEPARATE nested git repo (gitignored). The pattern _tasks/ (trailing slash) ignores only a directory, not a SYMLINK named _tasks. A self-referential _tasks symlink can slip in via git add -A and, once pulled, checkout materializes it over the real _tasks repo (destroying plans/specs/hands-off). Anchored /_tasks ignores the symlink too, preventing re-capture. * fix(translator): keep Responses namespace identity across the hub-and-spoke pivot Step 1 of the pivot (openai-responses -> openai) flattens namespace sub-tools to a qualified wire name (#8295) and records the `{namespace, name}` pair on a non-enumerable `_toolNameMap`. Step 2 (openai -> target) returns a brand-new object, so the property was dropped for every non-OpenAI target. chatCore then handed `null` to the #7936 response seam and namespace sub-tool calls reached the client under their flattened name, which Codex rejects with `unsupported call: <name>` — the symptom #7936 was opened to fix. Copying `_toolNameMap` through is not viable: openai-to-claude and openai-to-gemini publish their own `Map<string, string>` alias map on that same property during step 2, so it carries two incompatible types. This adds a dedicated `_namespaceToolIdentityMap`, propagated by translateRequest across the pivot; chatCore prefers it and falls back to `_toolNameMap` for the non-pivot producers. Both keys are stripped from the cliproxyapi wire body. Fixes #9780 * fix(chat): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reduce combined file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reduce combined file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Co-authored-by: VXNCXNX <vincent@preuve.ai> * fix(sse): route claude/<provider>/<model> aliases for catalog-only providers (#9856) The /v1/models catalog mirrors `claude/<provider>/<model>` ids purely from the alias gate -- ccAliasPredicate.ts consults no provider registry. The request path additionally required the prefix to be an open-sse REGISTRY entry or an operator-defined custom node. Enterprise-cloud providers such as azure-ai / azure-openai live only in the provider catalog (src/shared/constants/providers/apikey/enterprise-cloud.ts). They route fine directly -- `azure-ai/Phi-4` returns 200 -- but have no open-sse registry entry, so the two sides disagreed: the catalog advertised `claude/azure-ai/<model>` while stripCcDiscoveryAlias refused to strip it. The unstripped id then fell through to normal resolution, which splits on the first / and parsed `claude` as the provider. Every Claude Code request for an Azure model was routed to the Claude provider instead: ROUTING: Provider: claude, Model: azure-ai/DeepSeek-V4-Flash Extract the predicate as `isRoutableProviderPrefix()` and widen it to the provider catalog (id + alias) alongside the open-sse registry, so the request path recognises exactly what the catalog can advertise. Regression guard: tests/unit/cc-discovery-alias-routable-prefix.test.ts pins azure-ai/azure-openai/azure as routable, keeps openai/anthropic routable, and keeps an unknown prefix non-routable. Verified failing before the widening. Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * fix(i18n): translate validation model keys in 34 locales (#9857) The provider-connection dialog (AddApiKeyModal / EditConnectionModal) rendered humanized key names instead of real copy for providers.validationModelId{Label,Placeholder,Hint} in 34 of 43 locales — the values read "Validation Model Id Label", "Validation Model Id Placeholder" and "Validation Model Id Hint" verbatim. Each translation follows the terminology and register already used by the neighbouring provider keys in its own file — e.g. de Anbieter/API-Schlüssel with formal Sie, fr fournisseur/clé API, ru провайдер/ключ API — and each locale's own "e.g." convention (z. B., 例:, напр., ör., cth., hal.). Source of truth is en.json, which labels the field "Validation Model" (no "ID"); a few older locales say "validation model ID" and were left untouched rather than propagating that divergence. Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * cherry-pick(pr-9770): chore(repo): ignore Electron build output unpacked into repo root (#9858) * chore(repo): ignore Electron build output unpacked into repo root electron-builder (squirrel-windows target) unpacks the packaged app -- the entire Chromium runtime, ~24k files -- directly into the repository root: OmniRoute.exe, chrome_*.pak, *.dll, locales/, resources/, icudtl.dat, snapshot blobs and the Chromium license files. None of it was covered by .gitignore, so `git add -A` would commit the whole runtime. Every rule is root-anchored (leading `/`) because a bare `locales/` or `resources/` would also swallow tracked sources -- notably the CLI translations in bin/cli/locales/*.json. Verified with `git check-ignore`: all artifact paths ignored, and bin/cli/locales/{en,de}.json remain tracked. * chore(electron): sync package-lock for windows installer deps Adds the lockfile entries for the Windows installer/signing toolchain that the electron build now pulls in: electron-builder-squirrel-windows, electron-winstaller and @electron/windows-sign (plus their transitive fs-extra/jsonfile/universalify/mkdirp pins), and bumps app-builder-lib and builder-util-runtime. Lockfile-only change; no source or runtime behaviour is affected. --------- Co-authored-by: Mihaly Bodo <michael@proton-quantum.com> * fix(skills): normalize web fetch credentials (#9859) Co-authored-by: backryun <bakryun0718@proton.me> * fix(types): narrow DeepSeek tool calls (#9860) Co-authored-by: backryun <bakryun0718@proton.me> * fix(perf): memoize synced pricing reads (#9861) Co-authored-by: chloeassistant <279834366+chloeassistant@users.noreply.github.com> * cherry-pick(pr-9744): test(integration): add general live-test tool for the real "default" combo + rootless wire capture (#9862) * test(integration): add general live-test tool for the real "default" combo Temporary WIP commit on this deferred branch — lands in its own separate PR once the bug-fix extraction batch is done (never bundled into a bug-fix PR). Unlike liveGeminiShared.ts (provisions its own narrow 2-model Gemini-only combo), this reads the REAL "default" combo currently configured on the target instance directly from the DB and exercises every provider/model step in it directly, bypassing combo routing, so live-test coverage always matches whatever is actually configured instead of a hardcoded snapshot. Live-verified against omniroute-beta (seeded with the real 18-model, 5-provider default combo): 14/18 models pass consistently across non-streaming + streaming Chat Completions and streaming Responses API. The 4 consistent failures are real external state (cerebras credits_exhausted, one deprecated openrouter free-tier model), not code regressions. (cherry picked from commit c40b13a48fd897259c56f5122e9e57a3dc7654ba) * test(integration): add rootless wire-capture correlation to the live-test tool Temporary WIP commit on this deferred branch — lands in the same final live-test-tool PR as the general default-combo suite, never bundled into a bug-fix PR. liveContainerHarness.ts spins up a dedicated, throwaway podman container (same runner-base image target as the operator's local dev/beta containers) so wire-capture tests are fully self-contained: builds the image if missing, starts the container with a persistent data dir, waits for health, seeds the real "default" combo + provider connections from the operator's local omniroute-dev instance (idempotent — only runs once per data dir), and provisions API keys via the running instance's own auth flow. wireCapture.ts captures the container's actual network traffic via `podman unshare nsenter --net=<container netns> -- tcpdump` — no root needed, verified working live (this generalizes the root-requiring `sudo nsenter -t $PID` command scripts/sre/tcp-close-analyzer.py already documented for the same rootless-Podman netns problem; that script's docstring now documents both). Capture and analysis needed two real fixes found only by running the pipeline live: `-U` (unbuffered tcpdump writes) plus a `pkill -f <pcap path>` fallback, since `podman unshare -> nsenter -> tcpdump` is a 3-level subprocess chain and SIGTERM to the top-level process doesn't reach the tcpdump grandchild, leaving an orphaned process and a truncated/unreadable pcap; and filtering on the container's internal listening port (20128) rather than the dynamically-assigned host port, since capture happens inside the container's own network namespace where only the internal port is meaningful. live-default-combo-wire-capture.test.ts (gated on RUN_LIVE_WIRE_CAPTURE=1) ties it together: sends a small representative sample of requests through the real default combo, then cross-checks each one's app-level JSON status against the actual HTTP status line observed on the wire via scripts/sre/tcp-close-analyzer.py's stream reassembly — catching bugs where the app layer claims success but the wire shows a truncated/reset stream, not just what liveDefaultComboShared.ts's existing breadth suite already covers. Live-verified end-to-end: 4/4 sampled requests correlated correctly across 8 captured TCP streams, container + capture process fully torn down afterward (verified no orphaned podman container or tcpdump process left running). sendModelRequest/filterActiveModelTargets (liveDefaultComboShared.ts) gain optional baseUrl/apiKey overrides, defaulting to the existing module-level omniroute-beta target, so the wire-capture suite can point the same request-sending logic at its own dedicated container instead. (cherry picked from commit 914a7e42cbe914f257db9f72eedc902ee1532083) --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * maint: follow-up cherry-pick fix-in-place #9741 (conflict-resolved fallback) (#9895) * fix(responses-api): sync reasoning-cache write index with the fixed read side The turn-index-hardcoding fix updated the reasoning-cache read side (translator/index.ts's main replay loop) to key lookups by the assistant message's real position in the messages array, but two other spots still used the old hardcoded convention: - chatCore.ts's write side (both the streaming and non-streaming completion paths) still cached every response under a hardcoded messageIndex: 0. - translator/index.ts's own plain-turn (non-tool-call) cache-key lookup ALSO still hardcoded messageIndex 0 at its call site — a second, previously undiscovered instance of the same class of bug, found while re-verifying this fix against the current upstream tip (the original fix only addressed the write side). Past the first assistant turn these conventions no longer matched, so DeepSeek/Xiaomi-mimo plain-turn reasoning replay silently missed the cache and fell back to the placeholder (or, once #9573 removed the placeholder fallback, to an absent field) in ordinary multi-turn conversations. Compute the write-side index from the incoming request's message count instead, and use the real loop-provided messageIndex on the read-side lookup, both matching the position the response occupies once the client appends it to history for the next turn. Note: this was originally part of a larger squashed fix (output_index collision prevention across reasoning/message/tool_call items, reasoning-content-alias generalization) that has since been superseded by upstream's own independent fix — translator/response/openai-responses.ts now has its own dense-output-index-sort + getReadableReasoningValue implementation (own comment: "mirrors upstream PR #721"). Only this narrower, still-genuinely-broken write/read index sync survives as a distinct bug. Test plan: - TDD: tests/unit/reasoning-cache.test.ts's new end-to-end "write side (chatCore's messageIndex) and read side (translateRequest) agree on the same key end-to-end" test, plus the pre-existing "should inject placeholder for a plain (non-tool-call) DeepSeek turn" and "should replay cached reasoning for a plain (non-tool-call) DeepSeek turn when available" tests — confirmed failing against the pre-fix code on a clean release/v3.8.50 checkout (both the hardcoded-0 write side AND the hardcoded-0 read-side lookup independently reproduce the mismatch), passing after both fixes - npm run typecheck:core — clean - npm run lint — clean - npm run check:file-size — clean (chatCore.ts rebaselined 5034->5042 for the messageIndex computation at both call sites; reasoning-cache.test.ts frozen at 1035, matching the original fix's own rebaseline) - 2 pre-existing, unrelated test failures in the same file ("should replace empty-string reasoning_content with NON_ANTHROPIC_THINKING_PLACEHOLDER on cache miss", "should inject placeholder for a plain (non-tool-call) DeepSeek turn missing reasoning_content") confirmed present on a completely clean, untouched release/v3.8.50 checkout — these test obsolete placeholder-injection behavior the code deliberately removed per #9573 (see the code's own comment); not touched by this PR * fix(chat): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reconcile file-size baseline Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * cherry-pick(pr-9738): feat(logging): make the chat-log truncation limit configurable, bumped default 128x (#9863) * feat(logging): make the chat-log truncation limit configurable, bumped default 128x The 8KB cap on logged request/response bodies (open-sse/handlers/chatCore/logTruncation.ts::truncateForLog()) was hardcoded — trivially exceeded by any real multi-turn agentic conversation, meaning the dashboard's "Full Conversation" panel could only ever show a placeholder instead of the actual messages for nearly every logged row of any conversation with real substance. - Added CHAT_LOG_MAX_BODY_KB env var (src/lib/logEnv.ts:: getChatLogMaxBodyBytes()), default 1024 KB (1MB) — a 128x bump from the old hardcoded 8KB — following the same configurable-limit pattern as the sibling CHAT_LOG_TEXT_LIMIT/CHAT_LOG_ARRAY_TAIL_ITEMS/etc. vars. - Documented in .env.example and docs/reference/ENVIRONMENT.md. estimateSizeFast() (open-sse/utils/estimateSize.ts) has been substantially rewritten upstream since this bug was first found (now an iterative Frame-based walker with a separate node-visit budget, not the simple stack loop originally patched) — re-implemented the fix against the current algorithm rather than porting the old diff: the byte early-exit was unconditionally the module-level ESTIMATE_SIZE_BYTE_LIMIT (256 KiB) with no way for a caller to raise it, so any caller comparing against a bigger configured threshold could never see a size above ~256 KiB — every payload between 256 KiB and the caller's real limit looked "under threshold" and truncation never fired, the opposite of intended. Added an optional byteLimit parameter (default unchanged at ESTIMATE_SIZE_BYTE_LIMIT, so isSmallEnoughForSemanticCache's existing behavior is untouched) threaded through both the byte-check early-exit and the node-budget-exhaustion fail-closed fallback, with truncateForLog() now passing its own configured getChatLogMaxBodyBytes() value through. * feat(dashboard): show conversation session tag in request detail metadata Adds a "Conversation" field to the request detail panel's metadata grid (after "Combo"), showing the request's conversation id (sessionTag) for quick reference/copy. --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * cherry-pick(pr-9735): feat(logging): bump CHAT_LOG_ARRAY_TAIL_ITEMS default 24 -> 128 (#9864) * feat(logging): bump CHAT_LOG_ARRAY_TAIL_ITEMS default 24 -> 128 Real agentic CLIs with many MCP servers routinely declare 40-50+ tools in a single request — a live OpenClaw session logged 47. The tail-24 default silently dropped the array's earlier entries behind an _omniroute_truncated_array marker, so investigating why a specific tool call (apply_patch) behaved oddly turned up nothing: its declared shape (function vs custom type) was unrecoverable from the call log across 40 recent requests, even though the calls themselves succeeded. Bumped the configurable default to comfortably cover real large tool lists with headroom. Updated .env.example and docs/reference/ ENVIRONMENT.md to match (env-doc-sync check passes). * test(logging): pin CHAT_LOG_ARRAY_TAIL_ITEMS default at 128 The bump commit had no dedicated test asserting the literal default value; the existing chatcore-log-truncation.test.ts derives its expectations from getChatLogArrayTailItems() itself, so it can't discriminate a regression back toward the old, too-small 24 default. --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * fix(logging): use configurable max-depth when bounding logged tool_calls (#9865) requestLogger.ts's cloneBoundedForLog had its own hardcoded depth cap of 6, independent of the existing configurable getChatLogMaxDepth(). A typical Chat Completions response body's responseBody.choices[0].message.tool_calls[0].function sits at exactly depth 6, so every logged tool call's function field (name+arguments) was silently replaced with the literal string "[MaxDepth]" before ever being stored — corrupting the data, not just how it renders. Bumped the shared default 6->20 and switched requestLogger.ts to read it instead of using its own literal. (cherry picked from commit a2df6cf289cbab7cd618b8e55272434812f7a4a7) Co-authored-by: Markus Hartung <mail@hartmark.se> * fix(executors): repair DuckDuckGo AI Chat challenge solver (418 ERR_CHALLENGE) (#9866) Every duckduckgo-web chat request failed with HTTP 418 ERR_CHALLENGE while duck.ai worked normally in a browser from the same IP. Ground truth was established by driving a real headful Chromium at duck.ai from that IP (it returned 200), so the environment was never the problem — the anti-abuse challenge solver was. Six independent defects were found; the first alone disabled the solver completely. 1. Module syntax inside the vm sandbox source. CHALLENGE_STUBS is executed with vm.runInContext, which compiles in SCRIPT mode. A refactor mass-added `export` to the five `function` declarations inside that template literal (they read as ordinary top-level TS functions), so every solve threw SyntaxError. The executor swallows solve failures and posts the raw unsolved challenge, which upstream answers with 418. 2. Double-escaped regex in a String.raw template. `\\s` in __parseCssDisplay reached the sandbox as a literal backslash, so the display regex never matched and a getComputedStyle probe silently read empty. 3. buildHtmlLookup undercounted descendants by one. `count` backs el.querySelectorAll('*').length; that returns DESCENDANTS and countHtmlElements already skips the #document-fragment root, so the `- 1` was wrong. Chromium reports 3 for '<li><div></li><li></div'; we reported 2, and a variant multiplies innerHTML.length by that count. 4. Browser-fidelity probes. Newer challenge variants assert JS/DOM invariants a flat stub cannot satisfy: real prototype chains (HTMLDivElement -> HTMLElement -> Element), NodeList identity, a live body.children HTMLCollection, native-code toString, and sloppy-mode `this === window`. Nine of thirteen failed. Notably Math must NOT be sealed — Chromium reports Object.isSealed(Math) === false, and sealing it made our vector differ by one. 5. The solved payload dropped meta.origin / meta.stack / meta.duration. The duck.ai bundle always sends all three; captured browser requests confirm it. Without them upstream returns 418 even when every client_hash is correct. 6. reasoningEffort is now mandatory on duckchat/v1/chat. An otherwise byte-identical payload returns 200 with the field and 400 ERR_BAD_REQUEST without it (A/B verified live, repeated). Also removes the throwaway "seed" chat POST that ran before every real request. It existed to coax a usable challenge out of the upstream while the solver was broken; it only doubled chat calls against an IP-rate-limited endpoint, showing up as spurious 429 ERR_RATE_LIMIT. Verification: the solver now reproduces real Chromium's probe vectors exactly for all 8 captured challenge variants, and the executor returns 200 end-to-end live (non-streaming, streaming, claude-haiku-4-5, and a math prompt returning "42"). Tests: tests/unit/duckduckgo-challenge-solver-regression.test.ts (32 tests) and tests/unit/duckduckgo-reasoning-effort-required.test.ts (5 tests), backed by tests/fixtures/duckduckgo/challenge-variants.json — real captured challenge programs plus the probe vectors a real browser produced for them, so the suite asserts against recorded browser behaviour rather than our own output. Each fix was confirmed to fail its test when individually reverted. Co-authored-by: Mynacol <git@mynacol.xyz> * cherry-pick(pr-9730): fix(compression): persist RTK renderer configuration (#9867) * fix(compression): persist RTK renderer configuration * docs(changelog): add fragment for #9730 Adds the changelog.d/fixes/9730-persist-rtk-renderers.md fragment required by check:changelog-integrity for the RTK enableRenderers persistence fix in PR #9730. --------- Co-authored-by: Isaac <isaaclyons98@gmail.com> * fix(dashboard): unregister leftover service workers in dev mode (#9868) A phone that previously loaded a production build on this origin (or an old dev build from before the registration was gated) kept an active service worker across dev restarts. It intercepted every navigation/asset fetch, occasionally serving a JS chunk that didn't match the running dev server, which tripped Next's dev-client chunk-mismatch auto-reload — visible as an unexplained, unstoppable refresh loop on that device only (confirmed via a clean private tab on the same phone/URL not looping). PwaRegister now actively unregisters any existing service worker registrations and clears their caches outside production, instead of just skipping a new registration. (cherry picked from commit 66a2515cbce7a6132639614d88d48349a83bdcde) Co-authored-by: Markus Hartung <mail@hartmark.se> * fix(combo): remove stray brace from #9630 error handling (#9894) Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com> * feat(oauth): add Openference OAuth and API key provider integration (#9869) Wire Openference as a first-party OAuth gateway (PKCE, rotating refresh) and an API-key catalog entry on api.openference.com, with live model discovery, connection testing, free-tier badges, and regression tests. Co-authored-by: Anh Tran <anhlead@outlook.com> * maint: follow-up cherry-pick fix-in-place #9719 (conflict-resolved fallback) (#9893) * fix(db): clear combo pins when connections are deleted * docs: add changelog entry for #9719 --------- Co-authored-by: Zartharas <1402357+Zartharas@users.noreply.github.com> * cherry-pick(pr-9718): feat(src): proxy-pool-toolbar-minor-improvements (#9870) * feat(proxy-pool): streamline pool actions * test(proxy-pool): cover toolbar layout * refactor(settings): extract proxy registry helpers Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(settings): reduce proxy registry component size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Agnes <linkscrazy2@gmail.com> * feat(resilience): expose providerQuotaOverrides via /api/resilience (#9871) Co-authored-by: herjarsa <herjarsa@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9712 (conflict-resolved fallback) (#9892) * fix(build): colocateLlmlinguaOptionals skip-check treated a Next-traced stub as fully copied Debugging the omniroute-beta Docker rebuild: `npm run build` (and the Dockerfile's own post-build verification) failed with `Cannot find module '.../node_modules/@atjsh/llmlingua-2/dist/index.js'`. Root cause, reproduced directly (both against a live Docker builder image and in a unit test): Next.js's own standalone trace creates a stub directory for `@atjsh/llmlingua-2` containing only `package.json` — it references the package (a dynamically-imported optional dependency) but can't fully bundle it. colocateLlmlinguaOptionals's skip checks (both the closure-level early return and the per-package loop) only tested `existsSync(dest)`, so that stub was indistinguishable from "already fully co-located" — the function skipped copying the real `dist/` output entirely, silently shipping a package with a manifest but no code. Fix: check for the package's declared `main` entry file when it has one (the real-world case for every actual SLM optional). Packages with no `main` field fall back to comparing the destination's top-level entries against the source's — correct both for genuinely multi-file packages and for a metadata-only source (package.json is then its complete, faithfully- copied contents), which the existing idempotency test exercises. Covered by tests/unit/colocate-optionals.test.ts's new stub-reproduction case (fails against the pre-fix code, passes after — confirmed directly) plus the 6 pre-existing cases, all still green. (cherry picked from commit 359aba59c7b362a5efaa4f0cd4d482d48ee2df66) * fix(build): register onnxruntime-node's native bin/ as a standalone asset (#9687) Docker/standalone builds of the LLMLingua SLM compression tier failed at runtime with "Error: libonnxruntime.so.1: cannot open shared object file: No such file or directory" (open-sse/services/compression/engines/llmlingua's worker, via @huggingface/transformers -> onnxruntime-node). onnxruntime-node's dist/binding.js is a normal JS file Next.js's standalone trace bundles correctly, but binding.js dlopen()s a platform-specific native library shipped under bin/napi-v3/<platform>/<arch>/libonnxruntime.so.1 — a dynamic native load static file tracing can't see (same blind-spot class as the separate colocateLlmlinguaOptionals stub bug, just for a .so instead of a JS import, via NATIVE_ASSET_ENTRIES instead). That directory was simply never registered, unlike better-sqlite3's native binary, which already goes through the exact same mechanism correctly. Fix: add an entry for onnxruntime-node/bin, mirroring the existing better-sqlite3 entry. Confirmed against a real Docker build of the Dockerfile's own post-build verification step: this was the very next failure once the separate llmlingua-2 stub bug was fixed and the build progressed far enough to reach it. Covered by tests/unit/assemble-standalone-onnxruntime-native-asset.test.ts (fails against the pre-fix code on both assertions, passes after). (cherry picked from commit 8c98a59f26a27e844678e673b31c1c21aaf72b0e) --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * maint: follow-up cherry-pick fix-in-place #9707 (conflict-resolved fallback) (#9890) * fix(db): renumber ccr_blocks migration 134 -> 139 134 was taken by 134_proxy_logs_egress_ip, so two migrations shared the same numeric prefix and check-migration-numbering failed. Move ccr_blocks to the next free slot and add the retroactive isSchemaAlreadyApplied guard so a DB that already applied it under 134 skips the re-run. * fix(combo): restore missing preferAntigravityConnectionsWithStoredProject quotaStrategies imported the reset-aware pool filter from ../antigravityProjectPersistence.ts, a module that does not exist — the helper belongs in antigravityProjectPersist.ts and was never added there, breaking typecheck. Add the helper alongside the persist path, point the import at the real module, and cover the filter with unit tests. * chore: add Makefile wrapping the canonical npm scripts * fix(compression): remove duplicate Antigravity project helper The release branch already includes the generic project-aware connection selection helper. Keep that implementation and remove the duplicate introduced while cherry-picking #9707. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Matias Baglieri <168452313+matiasbaglieri@users.noreply.github.com> * cherry-pick(pr-9695): fix(docker): make the webpack build-arg escape hatch actually work (#9872) * build(docker): make the bundler build-arg actually take effect A bare ENV shadows a same-named ARG for the rest of the stage, so --build-arg OMNIROUTE_USE_TURBOPACK=0 was silently ignored and the webpack escape hatch the surrounding comment advertises only ever worked through -e at runtime, never at build time. That mattered because Turbopack compiles in native Rust memory living outside the V8 heap, so OMNIROUTE_BUILD_MEMORY_MB cannot bound it. A build host with a memory ceiling gets SIGKILLed by the cgroup OOM killer with no error text at all, which reads like a hung build rather than an out-of-memory one. * docs(docker): correct the builder stage facts and document its cost The stage table described a builder that no longer exists: it named node:24.15.0-trixie-slim where every stage now derives from node:26-trixie-slim, and said the stage runs `npm run build -- --webpack` where it runs plain `npm run build`, which is Turbopack by default. That second one is worse than stale. A reader who needs the webpack fallback would conclude the Docker build already uses it and never look for the switch. Adds a Build-time resources section covering the two build args, why the V8 heap arg cannot bound Turbopack, and measured ceilings for both bundlers. The runtime paragraphs that followed get their own heading so they no longer read as part of the build-time story. * docs(docker): correct the runtime heap defaults Same drift as the builder stage, in the paragraphs just below it. The image exports OMNIROUTE_MEMORY_MB=1024 and derives NODE_OPTIONS from it, but the guide reported 512 in three places, including the environment variable table. The "if unset, the launcher uses 512" line was misleading in both readings: the image always sets the variable so that branch cannot fire under Docker, and outside Docker the launcher calibrates from host RAM rather than using a flat 512. * docs(changelog): add fragment for #9695 --------- Co-authored-by: Minxi Hou <houminxi@gmail.com> * maint: follow-up cherry-pick fix-in-place #9693 (conflict-resolved fallback) (#9887) * fix(web-tools): anchor tool contract at prompt tail + user-turn reminder The <tool> contract from prepareToolMessages was prepended as the first system message. Web executors fold all system messages into one block, so with agentic clients whose system prompts exceed ~28K chars the contract sat at the head of a huge block and web models ignored it, refusing tool calls with "tool X is not in my tool set" (chatgpt-web, 0/3 at 30K chars). Two changes, both required in testing: - Dual placement: the full contract now rides as a trailing system message (folds to the tail of the system block) and a one-line reminder naming the tools is appended to the latest user message. - Rewording: the contract now frames injected tools as client tools invoked via a plain-text protocol, distinct from the model's native tool registry (web.run, python.exec, ...), and instructs the model to never claim they are unavailable. Without this the model resolved tool names against its native registry and refused even when it had seen the contract. Measured on cgpt-web gpt-5.5-thinking/gpt-5.6-thinking/o3: prepend 0/3 tool calls at 30K chars; dual placement 16/17 across 30K-250K system prompts, 30-tool sets, multi-turn tool history, streaming, and 3-way concurrency, with no spurious calls on no-tool prompts. Known limit: ~40K-char single user messages still flake (2/3) due to the upstream model's own injection heuristics. All prepareToolMessages consumers parse system messages position-independently and select the current user turn by role scan, so the trailing system message is shape-safe for every web executor. * test(web-tools): cover contract placement edge cases --------- Co-authored-by: Ryan Brosas <ryanjoserbrosas@gmail.com> * maint: follow-up cherry-pick fix-in-place #9631 (conflict-resolved fallback) (#9886) * feat(db): add a job registry for scheduled background work Background jobs each ship their own timer today, so there is no list of what is scheduled, no history of what ran, and no way to pause one without an environment variable and a restart. The registry gives them one home: a jobs table holding the schedule, a job_runs table holding the outcomes, and a loopback-only API to inspect and control both. Cron jobs read their expression through an optional cronGetter rather than the stored column, so an operator changing OMNIROUTE_WARMUP_CRON does not need the row rewritten. register() is an idempotent upsert that refreshes the schedule but never overwrites `enabled` or `created_at`, which is what lets a job be re-registered on every boot without discarding the operator's toggle. Run history is pruned per job rather than globally, and safeRun records a failure for a handler that throws as well as one that returns success:false, so a crashing job leaves a trail instead of a gap. The API is under /api/jobs and gated to loopback in the route guard. It can trigger a run and flip a job off, which is runtime administration and does not belong on a remotely reachable surface. Signed-off-by: Minxi Hou <houminxi@gmail.com> * feat(jobs): move the budget reset and token health check onto the registry Both jobs owned their own timer and started themselves as an import side effect, so nothing could report whether they were running, when they last ran, or why a run failed. They now register with the job registry and are started from it, which also means their schedule and run history are visible through /api/jobs. startAll() runs each interval job's first tick synchronously, so both entry points start the registry only after initializeCloudSync() has been awaited. The old wiring reached that ordering two different ways: the budget reset was started after the init call, and the health check's first sweep sat behind a 10s timer. Replacing both with one startAll() would otherwise have moved the two handlers in front of the initialisation they run against. Both entry points also register the same pair of jobs. Registering one and not the other is how a background job goes missing without anything failing. sweep() now returns how many connections it swept, so the health check can record a real records_affected the way the budget reset does. The migration documents that column as a per-job count, and hardcoding zero would have left one of the two jobs reporting a number the schema promises but the code never produces. A skipped or empty sweep reports zero. Every existing caller ignores the return value. The token health check keeps its own disable semantics: the handler still calls isHealthCheckDisabled() before sweeping, so OMNIROUTE_DISABLE_TOKEN_HEALTHCHECK, the production-build phase and the automated-test guard behave as before. Its registry adapter lives in src/lib/jobs/ next to the budget reset rather than in tokenHealthCheck.ts, which is already above its frozen size ceiling on the base branch and should not grow further. The adapter lets a failing sweep throw rather than reporting it itself, matching the budget reset: safeRun records a thrown error as a failure run with its message. The warmup job is seeded disabled. Its handler arrives with the warmup scheduler, and startAll() filters on enabled before it looks for a handler, so seeding it enabled here would warn about the missing handler on every boot. * fix: allowlist cron-parser dep and document OMNIROUTE_RUNNOW_TIMEOUT_MS env var Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> * fix: pass max reasoning effort through by default, add global model registry fallback (#8057) (#9883) Co-authored-by: Mo'men Qatr <momen.qatr04@eng-st.cu.edu.eg> * cherry-pick(pr-9605): ci(test): route orphaned Vitest tests through blocking CI (#9875) * ci(test): route orphaned Vitest tests through blocking CI * docs: fix advisory status in AGENTS.md and refresh baseline note * fix(changelog): fix fragment format for #9415 * fix(changelog): preserve upstream fragment format --------- Co-authored-by: MohitRawat017 <rawatmohit17906@gmail.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> * cherry-pick(pr-9601): feat(responses): add encrypted reasoning replay opt-in (#9876) * feat(codex): add encrypted reasoning replay opt-in * feat(responses): generalize encrypted reasoning replay * docs: clarify encrypted reasoning provider scope * fix(ui): group reasoning replay with connection controls * fix(logs): omit encrypted reasoning payloads * fix(chat): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(chat): reduce combined file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: jackjinke <jack.kejin@gmail.com> * cherry-pick(pr-9572): fix(providers): reject the dashboard password as a connection API key (#9877) * fix(providers): refuse to store the dashboard password as a connection API key A browser autofilled the management password into a connection's API-key field. The resulting credential authenticates against nothing, so every request routed through that connection came back 401, and because the field looks like any other password input the same autofill fired again while the connection was being repaired by hand. The refusal belongs on the write path rather than in the form. Twenty routes create or update connections and all of them funnel through createProviderConnection and updateProviderConnection, so one check there covers every entry point including a future one. The two other places that write api_key are left alone on purpose: one re-encrypts rows that already exist and the other is the one-time db.json import, and neither takes a value an operator just typed. Update checks the incoming value, never the merged one. A connection that already holds the password has to stay editable or the operator cannot repair the exact state this prevents, and re-checking the merged value would spend a bcrypt round on every unrelated field edit. Only a real match blocks the write. An unreadable settings row or a throwing bcrypt call logs and allows, because a guard against one specific mistake must not turn into a way to lock out every connection write. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(providers): compare the untrimmed credential, and cover the guard's branches The guard trimmed the incoming value before comparing it, which catches a paste carrying whitespace the password does not have. It missed the mirror case: neither the login route nor the set-password route trims, so a dashboard password may itself begin or end with a space, and an autofill reproducing it exactly was trimmed into a value that no longer matched the stored hash. The write then went through, which is the state this guard exists to prevent. Both forms are compared now, the second only when the first fails on a string that differs, so an ordinary key still costs a single bcrypt round. Two branches carried no coverage and both are load-bearing. The catch that logs and allows is the only path that lets a write through; a stored hash bcrypt cannot parse reaches it without needing a mock, since the shape check accepts an impossible cost factor that the comparison then rejects. The early return is what keeps a token renewal -- a write carrying tokens but no apiKey -- from paying for a settings read and a bcrypt round every time it fires, and the same unparseable hash makes that path observable, so an absent warning is proof the return happened. The narrower scope is deliberate and now says so in the code: the OAuth tokens arrive from a provider's token endpoint rather than from a form, so extending the comparison to them would charge every renewal for a field no autofill can reach. Signed-off-by: Minxi Hou <houminxi@gmail.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> * cherry-pick(pr-9569): fix(settings): use provider prefixes in model overrides (#9878) * fix(settings): use provider prefixes in model overrides * refactor(settings): extract pricing tab helpers Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Xiangzhe <xiangzhedev@gmail.com> * fix: address self-review findings (#9900) Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com> * cherry-pick(pr-9675): fix(providers): per-provider opt-out for anonymous no-auth fallback (opencode-go/zen 401s) (#9873) * fix(providers): add per-provider opt-out for anonymous no-auth fallback API-key providers with anonymousFallback: true (opencode-go, opencode-zen, pollinations, kilocode) receive a synthetic "noauth" connection whenever all real connections are terminal (credits_exhausted/banned/expired) or unavailable. The opencode upstream now rejects anonymous requests with 401 Missing API key, so the fallback adds a guaranteed-failing round trip and health/reconnect noise before the combo moves on. Add a noAuthFallbackDisabledProviders settings array (zod-validated, persisted via /api/settings, following the blockedProviders pattern). When a provider is listed, maybeSyntheticNoAuthFallback returns null for anonymousFallback-only providers, so exhausted providers are skipped immediately as allExpired/allRateLimited while real keyed connections keep working and recover automatically once quota state clears. True no-auth providers are unaffected; blockedProviders remains their disable mechanism. Default (absent/empty list) preserves current behavior. Provider detail pages for anonymousFallback providers gain an "Anonymous fallback" toggle (default ON) backed by the new setting. Refs #9674 * fix(auth): reduce file size Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: Hermes Agent <hermes@hermes-chloe.hyades.io> * cherry-pick(pr-9634): fix(test): reconcile base-drifted test expectations on release/v3.8.50 (#9874) * fix(combo): restore routing module load * fix(db): resolve ccr migration version collision Renumber the CCR block-store migration from 134 to 139, reconcile databases that already applied the legacy slot, and add regression coverage for both upgrade paths. Co-Authored-By: GPT-5 <noreply@openai.com> * fix(changelog): format the aggregator balance fragment as a bullet The fragment landed with YAML frontmatter rather than the bullet the aggregator reads, so check:changelog-integrity exits 1 on every branch and takes the merge-integrity job down with it regardless of what the branch changed. Only the format changes. The entry text is the author's, unedited, and now carries the link to the pull request that shipped it. * fix(test): update expected auth/vision/provider schema for base-drifted expectations * fix(test): narrow this branch to the drifted test expectations Three other PRs already cover what this one was carrying. #9618 renumbers the colliding ccr_blocks migration, #9632 repairs the malformed aggregator changelog fragment, and #9676 restores the combo module load by implementing the selection helper the import was reaching for, rather than deleting the caller the way this branch did. Keeping any of it here would put two files back on the same migration slot and overwrite a better fix with a worse one. What survives is the part none of them touch. Once the combo barrel loads again, three assertions in the context-window filter suite start failing: they demand that catalog-too-small targets be dropped, while the file's own header and its four neighbouring tests say those targets stay available as runtime fallback. The unresolved import was masking them. A new case pins the output-token limit as a genuine hard requirement so the relaxation cannot drift further. The provider count assertion kept one literal at the old value after the rest of the file moved to 198, so the partition check failed on a sum that was correct. * chore(quality): re-time migrationRunner for the 139 guard on the new tip --------- Co-authored-by: alexey.nazarov@softmg.ru <alexey.nazarov@softmg.ru> Co-authored-by: GPT-5 <noreply@openai.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> * cherry-pick(pr-9556): fix(translator): preserve Kimi K3 Responses reasoning (#9879) * fix(translator): preserve Kimi K3 Responses reasoning * fix(translator): make K3 reasoning preservation model-driven * fix(translator): replay cached Kimi reasoning before fallback * fix(translator): keep authentic K3 reasoning through cleanup * refactor(reasoning): use replay policy for K3 --------- Co-authored-by: jackjinke <jack.kejin@gmail.com> * maint: follow-up cherry-pick fix-in-place #9510 (fallback resolution) (#9880) * feat(api): add GET /api/resilience/connections for per-account state The three temporary-failure mechanisms each have their own scope -- the provider circuit breaker covers a whole provider, connection cooldown covers one account, model lockout covers a provider/connection/model triple -- and until now nothing showed them side by side. Diagnosing "why is this key being skipped" meant reading three separate surfaces and correlating by hand, which is exactly what the docs' own debugging guidance asks an operator to do. The route returns all three keyed by connection, plus the breaker's transition history so a flapping provider is visible as a sequence rather than a single current state. getStatus() already assembled everything except that history; it now returns a copy of it and carries an explicit CircuitBreakerStatus type instead of an inferred one. Reading raw connection rows for this meant widening getRawProviderConnections' column projection, so the existing allowlist is exported and the route selects through it. A test asserts every column the route names is in that allowlist, which turns a future typo into a failure here rather than a silent empty field. Each of the three data sources is wrapped independently: one of them throwing degrades that section and sets meta.degraded rather than failing the whole response, since a partial view still answers most of the questions the page exists for. Loopback-gated. It spawns nothing, unlike every other entry on that list, but it exposes per-account operational state and the comment says so to keep it from being read as precedent for gating read-only routes generally. Tests are real isolated-DB integration tests rather than mocks -- ESM mocking is unavailable here (no mock.module, non-configurable exports) and the codebase already has the isolated-DB pattern, which exercises more than a mock would anyway. Signed-off-by: Minxi Hou <houminxi@gmail.com> * feat(dashboard): add the per-account resilience connections page Renders what the API added: every connection with its cooldown, its provider breaker, and its model lockouts in one table, with a detail view per connection and the breaker's transitions drawn as a timeline. The timeline is the part that is hard to get from the existing surfaces -- a breaker sitting at CLOSED right now looks healthy, and only the sequence shows it has opened four times in the last hour. Polls rather than streams. The state it displays changes on the order of seconds to minutes and the page is loopback-gated, so an SSE channel would buy nothing over an interval. ModelCooldownsCard had its own formatRemaining. The new table needs the same countdown format and two copies would drift, so it moves to shared/utils/formatRemaining.ts and both import it -- behaviour unchanged, the extracted version differs from the deleted one only in local variable names. DataTable's column and row interfaces are exported for the same reason: the new table types against them rather than restating their shape. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(i18n): translate new resilience-connections screen strings PR #9510 added the "Connection Resilience" dashboard screen but the sync-added i18n keys (sidebar.resilienceConnections/Subtitle and the full resilienceConnections namespace) were left as __MISSING__: in every non-English locale, dropping i18nUiCoverage.pct below the 99 ratchet baseline. Translate all ~78 new leaf strings into all 41 non-English locales. Pre-existing unrelated __MISSING__ debt (hermesRole*, apiProtocol*, grokAutoTopUp*, featureFlagExposeFunctionalGatewayMirrorsDescription) is left untouched — out of scope for this fix. Co-authored-by: HouMinXi <HouMinXi@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Co-authored-by: HouMinXi <HouMinXi@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9549 (conflict-resolved fallback) (#9881) * fix(adobe-firefly): open browser sign-in and resolve provider slug in /login POST /api/providers/[id]/login passed the connection DB id to inAppLoginService.startLogin, but that service looks up the provider by slug in TOKEN_EXTRACTION_CONFIGS. The lookup always missed and returned "No extraction config" without launching a browser — so the VibeProxy "Sign in" button for Adobe Firefly (and every other web-cookie provider) never opened a browser. Adobe Firefly additionally had no extraction config because its IMS JWT is never in cookies/localStorage — it only rides on the Authorization: Bearer header of firefly-3p.ff.adobe.io XHRs. - Resolve the provider slug from the connection row and pass the slug (not the DB id) to inAppLoginService.startLogin. - Add open-sse/services/adobeFireflyBrowserLogin.ts: a Playwright service that launches a visible browser at firefly.adobe.com and intercepts firefly-3p requests to capture the IMS JWT + sherlockToken cookie. Wire it into the /login route for the adobe-firefly slug. - Fix latent bug: updateProviderConnection reads camelCase keys (apiKey, providerSpecificData), so the previous snake_case call never persisted extracted credentials. * fix(adobe-firefly): open browser sign-in and resolve provider slug in /login POST /api/providers/[id]/login passed the connection DB id to inAppLoginService.startLogin, but TOKEN_EXTRACTION_CONFIGS is keyed by provider slug — so browser login never launched for web-cookie providers. Adobe Firefly also cannot use cookie extraction: the IMS JWT only appears on Authorization headers to firefly-3p.ff.adobe.io. Add a dedicated Playwright interceptor and persist credentials with camelCase keys that updateProviderConnection actually reads. * fix(adobe-firefly): use system Chrome/Edge CDP for browser sign-in Playwright is not available inside the pkg-packaged VibeProxyServices.exe, so import('playwright') always failed with 'Playwright not installed' and never opened a window. Launch Chrome/Edge with --remote-debugging-port and capture the firefly-3p Authorization Bearer via pure CDP WebSocket instead. * fix(adobe-firefly): live x-arp-session-id / Arkose wire (stop 408 under load) Browser generate-async requires x-arp-session-id as base64({sid,ark,ftr}) with a real Arkose blob (sherlockToken). JWT alone frequently returns colligo HTTP 408 system under load while credits still work. - Match live ftr magic __UDF43-m4_31ck + Arkose pk in synthetic ARP fallback - Ranked extract of sherlockToken / x-arp from Cookie, HAR, fetch() paste, and space-joined JWT+ARP (PasswordBox newline collapse) - Reuse one ARP for storage upload + generate-async - Clearer 408 errors when browser ARP is missing vs stale - Unit suite 42/42 * fix(adobe-firefly): durable session ARP rebuild and aux_sid false-positive Rebuild x-arp-session-id from forterToken/arkose/ff_session_guid instead of ranking long Cookie pairs (e.g. aux_sid=…) as opaque ARP, which caused colligo HTTP 408. Cache IMS JWT + cookie sessions, rotate ARP on 408 retries, and keep Playwright warm-up opt-in only (headless Forter is rejected). Also expand synthetic ARP shape with bfp/fpjs to match live successful captures. * fix(adobe-firefly): durable session, off-screen Chrome recovery, browser sign-in Rebuild x-arp-session-id from Cookie pieces (sid/ark/forter) so aux_sid is never sent as ARP. Sticky ARP + submit spacing reduce mid-batch colligo 408 thrash. Add optional managed Chrome warm (off-screen headed by default; Forter rejects headless) and POST /api/providers/{id}/login browser sign-in that returns JWT+Cookie after a fresh SSO. Visible sign-in resets off-screen window placement and clears prior Adobe session when adding another account. * fix(adobe-firefly): renew sessions through durable CDP * fix(adobe-firefly): isolate browser sessions per account * fix(adobe-firefly): make account login fresh and deterministic * chore(adobe-firefly): remove obsolete browser fallback * docs(adobe-firefly): document renewal controls * fix(adobe-firefly): harden CDP warm, risk session, and browser sign-in Stop colligo 408 thrash from stale Forter and frozen Google login during Sign in with browser: - CDP warm: clear Firefly origin storage + risk cookies (keep SSO); require forter age under 10 minutes on loop and timeout paths; dual CDP queues; await Runtime.runIfWaitingForDebugger; profile-lock launch retries - Session: connectionId fingerprint; write-back JWT+Cookie; warm-fail cooldown; fail closed risk_session_stale when forter is known-stale - Client: submit gate around generate-async; max 2 attempts when forter known-stale; poll 401 one refresh; pass sessionBrowserKey through handlers - Login route: pure system Chrome/Edge CDP only; camelCase credential persist - Unit: browser-login + firefly suites green (60) --------- Co-authored-by: artickc <artur1992123@mail.ru> * fix(db): resolve ccr migration version collision (#9884) Renumber the CCR block-store migration from 134 to 139, reconcile databases that already applied the legacy slot, and add regression coverage for both upgrade paths. Co-authored-by: fenix007 <fenix007@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9629 (conflict-resolved fallback) (#9885) * fix(compression): add Lite tool truncation toggle * fix(antigravity): add missing antigravityProjectPersistence.ts module The quota-strategy engine (quotaStrategies.ts) imports from antigravityProjectPersistence.ts, but only antigravityProjectPersist.ts existed in the tree. Add the missing module with the expected preferAntigravityConnectionsWithStoredProject() helper and re-export the existing persistDiscoveredAntigravityProjectId(). Co-authored-by: diegosouzapw <diegosouza.pw@outlook.com> * fix(file-size): rebaseline strategySelector.ts for Lite truncation toggle The PR adds one line to threading options?.config?.lite into applyLiteCompression. Update the frozen size from 1060 to 1061. Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Refs #9629 --------- Co-authored-by: Xiangzhe <xiangzhedev@gmail.com> Co-authored-by: xz-dev <xz-dev@users.noreply.github.com> Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> * maint: follow-up cherry-pick fix-in-place #9704 (conflict-resolved fallback) (#9889) * fix(sse): persist per-tool-call JSON escape state across SSE delta chunks escapeJsonStringValues() reset its inString/pendingEscape state on every call instead of carrying it forward per tool-call index, so a raw newline byte (or an already-escaped \n) split across two delta chunks got corrupted in transit — the model's own output was correctly escaped, OmniRoute broke it. Root-caused via a dispatched investigation into real OpenClaw traffic that looked like model-generation quality but wasn't. Fix: escapeJsonStringValues now takes and mutates a persistent per-call state object (JsonStringEscapeState), keyed per tool-call index in the translator's init state and cleared when a tool call is superseded. * chore(quality): rebaseline openai-responses.ts for the escape-state fix Own growth from the extracted per-tool-call JSON escape-state fix (previous commit): open-sse/translator/response/openai-responses.ts 1204->1249 (+45). --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * maint: follow-up cherry-pick fix-in-place #9711 (conflict-resolved fallback) (#9891) * fix(sse): grace period before finalizing a client disconnect as 499 (#9653) A client that closes its connection right after reading a fully-completed SSE stream can race OmniRoute's own completion bookkeeping: the bytes already reached the client, but the transform stream's own completion callback (onStreamComplete, which flips streamCompletionRecorded) hasn't finished bubbling up when the disconnect handler fires, so the request gets persisted as a false 499 with zero token usage even though it delivered its full response. Confirmed live on real traffic before this fix: a request whose server log showed "disconnect: request_signal_aborted" at 18236ms was persisted with status 200 and full token usage (82814/1292) once the grace period let the real completion win the race, matching what the client actually received. createClientDisconnectGraceHandler (new leaf in streamFailureFinalization.ts) polls isStreamCompletionRecorded() for up to STREAM_DISCONNECT_GRACE_PERIOD_MS (default 10s, env-configurable, 0 disables) before finalizing as a failure. If a real completion lands within the window, handleStreamFailure's own guard is a no-op and the genuine 200 stands. Covered by tests/unit/stream-disconnect-grace-period-9653.test.ts (fake-timer driven: already-recorded completion short-circuits, disabled-grace-period finalizes immediately, a completion landing mid-window skips finalize entirely, and no completion ever landing finalizes once the deadline passes). (cherry picked from commit 5d0fe28c4246518f7c7b588795a6da1573b5df16) * chore(quality): rebaseline chatCore.ts for the disconnect grace-period fix Own growth from the disconnect grace-period fix: 5030->5039 (+9, the createClientDisconnectGraceHandler wiring at the existing onClientDisconnectFinalize call site). --------- Co-authored-by: Markus Hartung <mail@hartmark.se> * chore: ignore playwright cli artifact dir * maint: final follow-up cherry-pick #9619 (#9901) * fix(quality): clears two release/v3.8.50 base-red gates Unblocks Merge integrity and Docs Gates for every PR against release/v3.8.50, not just this branch: - changelog.d/features/9415-newapi-sub2api-aggregator-balance.md had a non-standard YAML frontmatter header that no other fragment in the tree uses. check-changelog-integrity.mjs reads a fragment's first non-blank line to validate it starts with a markdown bullet; the frontmatter's leading `---` made that check fail regardless of the actual bullet content further down. Removed the frontmatter and reformatted the body to match the documented changelog.d/README.md bullet convention. - docs/ops/VM_DEPLOYMENT_GUIDE.md documented OMNIROUTE_MAX_POOL_SIZE and OMNIROUTE_DB_POOL_SIZE as tunable env vars, but neither is read anywhere in the codebase (confirmed via full-repo grep) — this repo uses SQLite, which has no connection-pool concept these vars could plausibly control. check:fabricated-docs --strict correctly flags fabricated env-var claims; removed the bullet rather than implementing a feature to match invented documentation. * fix(i18n): completes Vietnamese parity, fixes empty migration query Two more release/v3.8.50 base-red items, both surfaced while chasing CI failures on unrelated PRs: - vi.json was missing 8 keys that #9539 (NewAPI/Sub2API aggregator balance) added to en.json without a matching i18n:sync-ui run — pt-BR.json already had all 8, only Vietnamese drifted. Added translations for the 6 provider-settings strings, the feature-flag description, and the quota tooltip; verified against tests/unit/i18n-vi-completeness.test.ts (parity, placeholder preservation, ICU parse — all 5 assertions pass). - src/lib/db/migrations/120_interception_rules.sql was pure comments documenting a no-schema-change key_value namespace, with no executable SQL statement — the migration runner logged "FAILED: 120_interception_rules — Query contained no valid SQL statement" on every fresh DB init. 118_provider_param_filters.sql (same pattern, two migrations earlier) already ends with a bare `SELECT 1;` no-op for exactly this reason; 120 was just missing it. Verified directly against better-sqlite3 that the file now executes without error. * fix(types): clears 6 pre-existing release/v3.8.50 typecheck errors typecheck:core is its own blocking CI job (quality.yml), separate from Docs Gates/Merge integrity. Confirmed pre-existing and unrelated to any current work by branching this worktree directly from upstream/release/v3.8.50 with no other merges applied. - accountSemaphore.ts: isBypassed() already excludes null/<=0 maxConcurrency before ensureGate() is called, but a boolean- …
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…e2e assertion and 10 integration reds Electron Package Smoke — a packaging defect that had been hidden behind another packaging defect for nine days. Once the loginHeaderCapture fix let the main process start, the server underneath died on 'Cannot find module next': resources/app/server.js shipped without resources/app/node_modules. electron-builder discards the ROOT node_modules in code, not by configuration — app-builder-lib/out/util/filter.js:42 has a hard-coded `if (relative === "node_modules") return false` that runs before any filter pattern. The second extraResources entry pointing INTO ../.build/electron-standalone/node_modules is what sidesteps it, because those relative paths are never equal to "node_modules". diegosouzapw#10325 removed that entry as an apparent duplicate and flipped the test to assert "exactly once", freezing the regression as if it were the contract. Restored, and the unit guard now pins both entries — proven by mutation: reverting package.json to the post-diegosouzapw#10325 shape fails the guard 3/4, restoring it passes 4/4. group-b-quota-plans-config — the assertion was impossible to satisfy on ANY route, and the page was never broken. layout.tsx hands the whole message catalogue to NextIntlClientProvider, React serialises that prop into the RSC payload, and en.json carries "Internal Server Error" twice, so page.content() always contains it: probing /dashboard, /dashboard/costs, /dashboard/settings and /login showed the string present with every page rendering fine, and a pageerror probe on the failing run captured zero client exceptions. This is the same trap that killed the sibling not.toContain("500") in feb0709 ("raw HTML is unreliable") — that one was removed, this one was kept. Now asserts on rendered text, which still catches a real error boundary. The pageerror capture stays: the CI failure carried no stack trace, which is why it was misread twice. Integration — 10 of the 14 shard-2 reds, all sibling-test gaps behind security fixes: monitoring health now takes a Request and requires management auth (GHSA-mvf8-qc78-5mxm); the OAuth import routes moved to requireManagementAuth (GHSA-mg76) — the test accepts both guard shapes and gained a stronger anchor that every exported handler awaits a guard on its own request, mutation-verified; skill tool names are derived from encodeSkillToolName() and the fake upstream now returns the encoded name so decodeSkillToolName() is exercised too; previous_response_id now fails closed (diegosouzapw#10262); proxy_logs persist as an async batch (diegosouzapw#11182) so the test flushes first; providerQuotaOverrides joined GET /api/resilience (diegosouzapw#9871); the reasoning fixture used a model that stopped being thinking-incompatible, replaced and pinned with a premise assert so it cannot rot silently again. A vacuous assert.ok(true, "all 10 streams completed without hanging") was replaced with real anchors — content must arrive on every stream and the active Timeout count must not grow. Four are deliberately left red rather than aligned, each now tracked: diegosouzapw#11551 (the /v1/models after() wiring is dead — the route passes a third argument to a two-parameter function and catalogCache never imports after, so the diegosouzapw#8728 contract is unimplemented), diegosouzapw#11552 (~27% of requests emit an extra discarded upstream call; the delivered distribution is exactly 0.70, so weighted routing is correct and the waste is the real finding), the fixed-account combo pin (aligning it would destroy the per-step attribution the test exists for), and the web_search fallback already tracked as diegosouzapw#11524. Package Artifact — the provenance stamp I added last round used git rev-parse HEAD, which under pull_request is the ephemeral merge commit and therefore never an ancestor of the release branch. Now takes the PR head sha. Refs diegosouzapw#10692
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Mantainer-executed cherry-pick of the community PR #9714 from @herjarsa.
This branch preserves the original commits/authorship and keeps the work
available in this repository for merge without requiring author push access.