Skip to content

Release v3.8.40 - #5225

Merged
diegosouzapw merged 62 commits into
mainfrom
release/v3.8.40
Jun 29, 2026
Merged

diegosouzapw merged 62 commits into
mainfrom
release/v3.8.40

Conversation

@diegosouzapw

@diegosouzapw diegosouzapw commented Jun 28, 2026 •

Copy link
Copy Markdown
Owner

Release v3.8.40

Integration PR for the v3.8.40 cycle → main. Living release notes — mirrors the finalized CHANGELOG.md [3.8.40] section. origin/main merged in to resolve cycle drift (#5278/#5234/#5228 landed directly on main).

[3.8.40] — TBD

In development — bullets added per PR; finalized at release.

✨ New Features

  • feat(compression): relevance extractive engine — a new opt-in compression engine that scores each sentence by term-overlap (Jaccard) with the user's last query minus a length/boilerplate penalty, greedily keeps the most relevant within a budget, and reconstructs the original order. Pure-string, deterministic, ReDoS-safe (char-code tokenization, no RegExp over user input), fail-open, default off. Ideal for trimming long pasted RAG context / tool output to what's relevant. Sentences carrying real signal (digits/URLs/errors/code/paths) are never dropped; overlapThreshold/budgetPercent/boilerplateWeight are configurable. Tier-2 item of the compression feature-extraction roadmap (chore(env): sync .env.example with current .env structure #7). (#5289)
  • feat(compression): hard-budget mode — compress to ≤ N tokens — a deterministic post-pass (targetTokens / targetRatio, default unset → no-op) that trims a body to a token budget. It ranks sentences/lines by average scoreToken ascending and drops the lowest-saliency ones until the body fits (measured by the exact cl100k countTextTokens), preserving original order. Lines carrying real signal (digits, URLs, Error:-family, code fences, stack at-frames, multi-segment paths, key=value) are never dropped; the budget is distributed proportionally across messages so the total stays ≤ target; an unreachable target (all-preserved) surfaces a validationWarnings note instead of failing silently. Does NOT touch the estimateCompressionTokens budget-gate estimator. Tier-3 item of the compression feature-extraction roadmap (fix(ci): explicit .npmrc auth for npm publish #17). (#5288, follow-up #5291)
  • feat(compression): result memoization for deterministic engines (opt-in) — caches (input, config) → result for provably pure, stateless modes (lite/standard/rtk and stacked pipelines of {lite,caveman,rtk}) to skip recompute on the hot path. Opt-in via memoizeCompressionResults (default off → zero behavior change). Conservative opt-in whitelist (stateful ccr/session-dedup — which write the cross-request CCR store — and model-backed ultra/aggressive/llmlingua are never cached), principal-scoped (skipped without a principal, so no cross-principal body leak), and clone-on-store + clone-on-read. Tier-3 item of the compression feature-extraction roadmap (feat(security): FASE-01 to FASE-09 — Security Hardening & Advanced Features #21). (#5286)
  • feat(compression): inline transparency annotation — surfaces tokens=847→312; rules: filler×8, dedup×2 derived from existing compression stats. The X-OmniRoute-Compression response header is extended append-only (the mode; source=X prefix stays byte-identical, so existing header parsers don't break) and the compression studio cockpit shows a matching badge. Zero new computation — it aggregates the rulesApplied/techniquesUsed already on the stats. Tier-3 item of the compression feature-extraction roadmap (fix(ci): add environment for npm token access #18). (#5284)
  • feat(compression): saliency heatmap in the compression studio — the preview studio can now color each token by saliency: ultra per-token scoreToken (0–1, green→red gradient) or universal kept/removed from the existing diff. A dry-run visualization behind a toggle (no cost on a normal preview; backward-compatible when off). Completes the visualization half of roadmap item deps: bump qs from 6.14.1 to 6.14.2 #13 (the A/B comparison shipped in #5080). (#5285)
  • feat(compression): composite-command splitter for RTK detection — cd /x && git status now detects as git-status (previously the whole string was treated as one command and matched no filter). A quote-aware top-level tokenizer splits on &&/||/; (never inside quotes or $(…)/backtick subshells) and feeds the last segment to RTK command detection, so every RTK filter/renderer fires on commands wrapped in cd … &&/||/; chains. O(n), no RegExp over the command (ReDoS-safe). Tier-3 item of the compression feature-extraction roadmap (fix(ci): fix npm publish auth — support vars.NPM_TOKEN #16). (#5283)
  • feat(mcp): omniroute_tool_search tool + one-line TS signatures — new MCP tool that does lexical keyword search over every MCP tool's name/description and returns the top matches as compact one-line TypeScript signatures (~half the JSON-schema token cost), so agents discover tools on demand instead of carrying all ~88 schemas every turn. Search is ReDoS-safe (substring scoring, never new RegExp on the query) and deterministic; tools/list stays complete (no hidden tools). Adds the read:tools scope. Tier-1 item of the compression feature-extraction roadmap. (#5269)
  • feat(compression): RTK semantic command-output renderers (opt-in) — adds a second, opt-in compaction layer to the RTK engine that rewrites structured command output into a far more compact semantic form: git diff → file headers + @@ hunks + changed lines only; an all-green pytest/jest/vitest/eslint run → its one-line summary; terraform/tofu plan → Plan: +N ~M -K plus the resource list; kubectl/aws JSON arrays → a minimal table. Each renderer is conservative (no-op when the shape doesn't match) and the integration is fail-open; the test-green renderer never collapses output that carries any failure signal. Gated by RtkConfig.enableRenderers (default off → zero behavioral change). Eighth item of the compression feature-extraction roadmap. (#5268)
  • feat(compression): QuantumLock cache-prefix stabilization (opt-in, default off) — recovers upstream prompt-cache hits that a volatile fragment in the system prompt would otherwise bust. When a caller injects a session UUID, unix timestamp, request-id, JWT, API-key shape, or long hex digest into the role:system message every turn, the longest common prefix across turns ends at that changing byte → the whole system prompt after it is re-billed and re-processed each turn. QuantumLock replaces each non-semantic volatile fragment with a positional, value-independent placeholder ⟦Q{i}⟧ and appends the real values in a delimited ⟦QUANTUMLOCK⟧ tail. The rewrite is sent to the model (lossless — not restored), so the system-prompt body becomes byte-identical across turns and the provider caches the long stable prefix while only the small tail differs. Opt-in, default off, applied only for caching providers (isCachingProvider && config.quantumLock.enabled); bounded ReDoS-safe patterns; idempotent; no date/time patterns (semantically meaningful — explicit non-goal). Studio gets a toggle + a "🔒 N volatile fragment(s) stabilized" dry-run badge. Seventh item of the compression feature-extraction roadmap (bench: #5080, gate: #5127, fuzzy: #5143, ionizer: #5148, TOON: #5163, CCR ranged: #5187, risk-gate: #5243). (#5260)
  • kilocode: anonymous (no-auth) access to Kilo Code's free models, mirroring the opencode/mimocode pattern. With no Kilo account connected, requests now fall back to the gateway's anonymous tier (Authorization: Bearer anonymous on api.kilo.ai/api/openrouter) so the free models work without signup; a connected OAuth account is still used unchanged for the paid tier (#5259, [feature] Anonymous (no-auth) usage for Kilo Code free models, like opencode #4019 — thanks @Theadd for the reference implementation)
  • feat(logging): call-log correlation ID (end-to-end) — every request now gets a unique correlation id, returned in the X-Correlation-Id response header, persisted in call_logs (migration 109), filterable via /api/usage/call-logs, and surfaced in the dashboard request logger (per-chunk stream timestamps + active-requests-first sort). This is the safe, cohesive core subset of the larger feat: add CorrelationId and fix lazy loading [wating autor] #5275 — landed on its own so the low-risk value isn't blocked by the parts of that PR still under review. (#5279 — thanks @hartmark)
  • feat(providers): Microsoft 365 Copilot individual provider — adds the copilot-m365-web provider (the 237th), wiring the M365 BizChat framing/connection helpers into a selectable web-session provider backed by m365.cloud.microsoft/chat for individual Microsoft 365 plans. Builds on the M365 pure-framing groundwork from feat(executors): land M365 Copilot pure framing + connection helpers (#4042) #4696. Regression guard: tests/unit/copilot-m365-web-executor.test.ts. (#5302 — thanks @skyzea1)

🔧 Bug Fixes

  • ci(docker): re-point the Docker Hub / GHCR :latest (and :latest-web) tags to the just-published release. On a release: released event the freshly-created git tag is often not yet visible to git fetch --tags when docker-publish runs, so the :latest-promotion gate built its candidate set purely from git tag -l and resolved the highest semver to the previous version — leaving latest one release behind (3.8.39 published, latest still 3.8.38). The decision now lives in scripts/ci/should-promote-latest.sh, which folds the current VERSION into the candidate set before picking the highest stable semver, making promotion independent of tag-sync timing (a patch published after a higher minor still won't grab latest). Regression guard: tests/unit/build/should-promote-latest-5301.test.ts (#5301)
  • command-code: treat a non-positive max_tokens/max_completion_tokens (e.g. Zoo Code's -1 "let the server choose") as "no limit" — omit the field instead of forcing it to 1. clampMaxTokens previously did Math.max(1, …), so a client -1 was sent upstream as max_tokens: 1, truncating the response to a single token (the observed completion_tokens: 1, content: null, reasoning_content: "The" with finish_reason: stop). Now any value ≤ 0 is dropped so Command Code applies the model's native default; positive values are still floored and clamped to the 200k ceiling. Regression guard: tests/unit/command-code-maxtokens-negative-5166.test.ts (#5166 — thanks @Stazyu)
  • fix(auth): compare-and-swap guard on the OAuth refresh persist — under multi-agent load, the per-connection refresh mutex makes [network refresh + DB write] atomic for one connection, but it does not protect against a third writer (a sibling request, a concurrent HealthCheck, or a replica) landing a fresher refresh_token rotation on the same connection_id between the staleness read and the persist. Overwriting that fresher row reverts the sibling's rotation; the next caller then loads the now-consumed token, Auth0/Anthropic flag it as refresh_token_reused, and the whole token family gets revoked (the 1352× claude/aa5dd5cf invalidation storm). getAccessToken now re-reads the row's current refresh_token immediately before persisting (inside the mutex) and skips the write when it has rotated past the token the caller presented — the caller still receives the freshly-issued access token, only the DB overwrite is skipped. Opt-in via runWithCasGuard (no active guard ⇒ byte-identical behavior); skip/persist counters exposed via getCasGuardStats(). Regression guard: tests/unit/token-refresh-cas-guard-4038.test.ts. (#4038 — thanks @KooshaPari for the root-cause diagnosis)
  • mcp: break the schemas/tools.ts ↔ schemas/toolSearch.ts import cycle introduced when the tool_search defs (feat(mcp): omniroute_tool_search + one-line TS signatures — roadmap #4 #5269) were extracted into their own module — toolSearch.ts imported McpToolDefinition from tools.ts while tools.ts imported toolSearchTool from toolSearch.ts, failing check:cycles on release/v3.8.40. The shared AuditLevel + McpToolDefinition types now live in a leaf schemas/toolDefinition.ts that both import; tools.ts re-exports them for backward compatibility.
  • compression (analytics): record attempted-but-no-op compression runs so Stacked is no longer invisible when it saves nothing. Previously a compression_analytics row was written only on a net-positive saving, so a Stacked (RTK→Caveman) pipeline that ran on already-compact context produced no row — indistinguishable from "never dispatched" (byMode.stacked.count stayed flat while Ultra climbed). Such runs are now recorded with skip_reason and surfaced as a per-mode skipped count plus totalSkipped/bySkipReason in the analytics summary and the Mode Breakdown; the existing net-saving totals/averages are unchanged (skip rows are excluded from them) ([BUG] Stacked RTK + Caveman compression is unclear/unreliable; Ultra works but Stacked often records no savings #4268 — thanks @abdulkadirozyurt, @androw)
  • cli (tray): fix omniroute server --tray showing no tray on macOS/Linux with no error printed. The wired Unix tray path loaded systray2 through an inline loader that called require("module") inside an ESM .mjs file ("type":"module") → ReferenceError: require is not defined, silently swallowed (regressed in v3.8.34); even if it had loaded, systray2 isn't in node_modules (it's lazily installed into ~/.omniroute/runtime). The loader now delegates to the runtime loader, the icon path (icon.png) is corrected, isTemplateIcon is false (the full-color icon rendered as a white square under macOS template mode), and tray start failures are surfaced to stderr instead of being swallowed ([BUG] Omniroute #4605 — thanks @ProgMEM-CC)
  • agent-bridge (antigravity): unwrap the cloudcode-pa .request envelope when converting Antigravity IDE requests. The real IDE sends cloudcode-pa.googleapis.com/v1internal:generateContent with the Gemini request nested under .request ({ project, model, request: { contents, systemInstruction, generationConfig } }), but the bridge read those fields at the top level — yielding an empty conversation, so prompts hung mid-execution. The legacy /v1beta/models/<model>:generateContent top-level shape still works ([BUG] Antigravity-IDE login does not work with agent-bridge #4294 — thanks @shabeer)
  • dashboard: add a GitHub releases fallback to the "Update Available" lookup. After the v3.8.28 fix added an npm-registry HTTP fallback, the banner could still stay hidden on networks that reach GitHub (where the news feed already loads) but not registry.npmjs.org. resolveLatestVersion() now tries npm CLI → npm registry → GitHub releases (/repos/diegosouzapw/OmniRoute/releases/latest) before giving up, and logs a warning only when all three fail ([bug] Home page "Update Available" banner no longer appears when a newer version exists #4100)
  • command-code: omit max_tokens when the client omits it so the upstream applies the model's native default, fixing 400 "expected <=200000" on /alpha/generate for high-cap models; an explicit oversized client value is clamped to the 200k endpoint ceiling (fix(command-code): omit max_tokens when client omits it; correct registry caps #5221 — thanks @adivekar-utexas)
  • combo: wire session stickiness into the round-robin dispatch path. Multi-turn conversations from clients that send no session id (Codex CLI, Claude Code, most OpenAI-compatible tools) were rotated to a different connection on every turn by round-robin combos, busting the upstream prompt-cache → cold high-reasoning starts, intermittent 504s and throughput collapse under concurrency. The weighted/priority paths already honored per-conversation stickiness; the round-robin handler returned before reaching it. Round-robin now starts the rotation at the conversation's sticky connection (failover to the other targets is preserved), and different conversations still spread across connections — only intra-conversation rotation is removed (#5248, [BUG] Frequent 504 Errors and Significant TPS Degradation After Upgrading Beyond v3.8.14 #3825 — thanks @bypanghu, @jpsn123, @xz-dev)
  • kiro: replace the synthesized trailing "Continue" turn with a neutral filler ("...") — when an OpenAI→Kiro request ends on an assistant/tool turn, the translator synthesizes the protocol-required trailing user turn, and the literal word "Continue" could be read by Kiro/CodeWhisperer as a real user instruction and trigger unintended agent action. A trailing tool-result turn is still promoted as-is (it already collapses to a real user turn); only the assistant-text-ending case is affected. Regression guards: tests/unit/kiro-continue-filler-5231.test.ts. (#5231)
  • combo: advance to the next combo target on a 400 "requested model is not supported" instead of hard-failing. The 400 guard in the priority strategy treated MODEL_CAPACITY as a block-fallback reason, so a combo that hit a provider lacking a specific model returned a hard 400 even when other targets (different providers) supported it. Such 400s now fall through to the next target. (#5249 — thanks @Chewji9875)
  • dashboard: disabled no-auth providers no longer vanish from the All Providers page. Disabling a no-auth provider (the "No authentication required" toggle, which adds it to blockedProviders) silently removed its card because the page dropped blocked no-auth entries from its render list — the only way back was buried under Settings → Security → Blocked Providers. The page now partitions no-auth entries: visible providers render as before, blocked ones appear in a "Disabled" sub-group with an Enable button that un-blocks them in place. Aggregates, counts and /v1/models still consume the visible-only list (blocked providers stay out of routing). Regression guard: tests/unit/noauth-blocked-partition-5183.test.ts. (#5183, follow-up from #5166 — thanks @WslzGmzs)
  • dashboard: add a parent /dashboard/context page so RSC prefetches of the compression-context hub no longer 404. The route only had sub-routes (settings, combos, ultra, …) and no parent page, so the App Router returned 404 for the bare segment. The parent now redirects to its canonical sub-route (/dashboard/context/settings), honoring a legacy ?tab= query for deep links. Regression guard: tests/unit/dashboard/context-parent-redirect-5298.test.ts (#5298 — thanks @KooshaPari)
  • i18n: add the missing sidebar.gamificationGroup message across all 42 locales — the Gamification sidebar group referenced a titleKey that existed in no locale, logging MISSING_MESSAGE: sidebar.gamificationGroup (en) at runtime (the group still rendered via its titleFallback). The key is now present everywhere so the warning is gone and locale coverage is unaffected (#5298 — thanks @KooshaPari)
  • api(stream): /v1/chat/completions no longer returns SSE for a non-stream OpenAI-compatible request when stream is omitted and the client sends Accept: application/json, text/event-stream — the Vercel AI SDK / OpenAI SDK non-stream signature (doGenerate()/generateText()), which then failed with Invalid JSON response (Unexpected token 'd', "data: {"id"...). The route-level Accept override (fix(docker): use /api/monitoring/health for Docker healthcheck (#296) #302) and resolveStreamFlag now treat an Accept header that explicitly lists application/json as a JSON opt-in even when it also lists text/event-stream; only a pure Accept: text/event-stream (no application/json) still opts an omitted-stream request into SSE, and an explicit body stream value always wins. The shared decision now lives in acceptHeaderForcesStream. Regression guard: tests/unit/sse-nonstream-accept-5305.test.ts. (#5305 — thanks @md-riaz)
  • providers: drop the retired GPT‑5.2 / GPT‑4.5 models from the direct ChatGPT‑web and Codex surfaces (OpenAI removed them there), so OmniRoute stops advertising/routing models that no longer exist. Scoped on purpose to those two providers — third‑party proxies that still expose the ids are untouched. (#5280 — thanks @backryun)
  • codex: drop the deprecated local_shell hosted tool type before forwarding to OpenAI's Responses API, resolving the omni-combo 400 "The local_shell tool is no longer supported." spike. Inbound Responses local_shell is still accepted and mapped to a caller-side Chat shell function for compatibility. (#5250, #5256 — thanks @KooshaPari)
  • antigravity: retry excluded accounts via the fallback LRU. The combo same-model retry loop accumulates excluded Antigravity connection ids after account-level failures, but auth selection only treated a single excludeConnectionId as a fallback scenario — once exclusions accumulated through excludedConnectionIds, selection could fall back to normal sticky/priority behavior instead of LRU-selecting the next eligible account for the same model/family. Any non-empty accumulated exclude set is now treated as fallback mode. Builds on the family-scoped lockout work in fix(antigravity): retry accounts by quota family #5180 (v3.8.39). (#5222 — thanks @Ardem2025)
  • grok-cli: strip unsupported sampling params (presencePenalty, frequencyPenalty, logprobs, topLogprobs) before sending to the Grok Build API, fixing 400 'Model does not support parameter presencePenalty' when clients (MiMoCode, Cursor, etc.) send OpenAI-style params. (#5273 — thanks @fulorgnas)
  • grok-cli: accept the full ~/.grok/auth.json object in the dashboard import-token endpoint. The oauthImportTokenSchema only accepted a bare string token while the UI sends the whole auth.json object → 400 Bad Request; the schema now accepts the object and stores the original under providerSpecificData.rawAuthJson for diagnostics and token refresh. (#5258 — thanks @fulorgnas)
  • qoder: coalesce concurrent PAT→job-token exchanges per PAT so high-concurrency / multi-agent bursts no longer stampede openapi.qoder.sh/api/v1/jobToken/exchange before the first exchange populates the completed-token cache; the shared exchange is also decoupled from any single caller's AbortSignal so one aborted waiter can't cancel it for the others. (#5254, #5265 — thanks @KooshaPari)
  • proxy: scope the fallback reachability cache by normalized target URL instead of by hostname, so a failed probe for one endpoint on a shared API host no longer suppresses a later probe for a different endpoint on that same host for the full TTL (host fallback is preserved for malformed URLs). (#5261 — thanks @KooshaPari)
  • proxy: cache failed fast-fail health probes with a short negative TTL instead of the full positive health TTL, so a single transient timeout/load blip no longer marks a working residential SOCKS5 proxy unreachable for the whole window ([BUG] Proxies fail with "Proxy Fast-Fail unreachable" under high concurrency despite working in curl/other tools #5109 regression coverage added). (#5255 — thanks @KooshaPari)
  • mcp: forward HTTP auth to internal tool fetches so MCP tools that call back into the local API surface carry the caller's authorization. (#5218 — thanks @KooshaPari)
  • logging: preserve the outbound provider request headers in the detailed call-log Provider Request payload (previously the upstream response headers were shown there). chatCore now keeps executor-returned request headers when wrapping streaming and non-streaming responses; response headers stay scoped to the Response. (#5257 — thanks @rdself)
  • sse: scope textual <think>/<thinking> tag extraction so generic OpenAI-compatible paths don't rewrite prompt-format content into reasoning_content; an explicit opt-in keeps tag-native families (DeepSeek-R1, QwQ) working while Antigravity/Agy stay excluded by provider/model prefix. (#5224 — thanks @rdself)
  • mcp: break the schemas/tools.ts ↔ schemas/toolSearch.ts import cycle introduced when the tool_search defs (feat(mcp): omniroute_tool_search + one-line TS signatures — roadmap #4 #5269) were extracted into their own module — toolSearch.ts imported McpToolDefinition from tools.ts while tools.ts imported toolSearchTool from toolSearch.ts, failing check:cycles on release/v3.8.40. The shared AuditLevel + McpToolDefinition types now live in a leaf schemas/toolDefinition.ts that both import; tools.ts re-exports them for backward compatibility. (#5282)
  • compression (analytics): record attempted-but-no-op compression runs so Stacked is no longer invisible when it saves nothing. A compression_analytics row was previously written only on a net-positive saving, so a Stacked pipeline that ran on already-compact context produced no row — indistinguishable from "never dispatched". Such runs are now recorded with skip_reason and surfaced as a per-mode skipped count plus totalSkipped/bySkipReason; net-saving totals/averages are unchanged (skip rows excluded). (#5277, [BUG] Stacked RTK + Caveman compression is unclear/unreliable; Ultra works but Stacked often records no savings #4268 — thanks @abdulkadirozyurt, @androw)
  • cli (tray): fix omniroute server --tray showing no tray on macOS/Linux with no error printed. The Unix tray path loaded systray2 through an inline loader that called require("module") inside an ESM .mjs file → ReferenceError: require is not defined, silently swallowed (regressed in v3.8.34); even if loaded, systray2 is lazily installed into ~/.omniroute/runtime, not node_modules. The loader now delegates to the runtime loader, the icon path is corrected, isTemplateIcon is false (the full-color icon rendered as a white square under macOS template mode), and tray start failures surface to stderr. (#5276, [BUG] Omniroute #4605 — thanks @ProgMEM-CC)
  • agent-bridge (antigravity): unwrap the cloudcode-pa .request envelope when converting Antigravity IDE requests. The real IDE sends cloudcode-pa.googleapis.com/v1internal:generateContent with the Gemini request nested under .request, but the bridge read those fields at the top level — yielding an empty conversation, so prompts hung mid-execution. The legacy /v1beta/models/<model>:generateContent top-level shape still works. (#5267, [BUG] Antigravity-IDE login does not work with agent-bridge #4294 — thanks @shabeer)
  • dashboard: add a GitHub releases fallback to the "Update Available" lookup. After the v3.8.28 npm-registry fallback, the banner could still stay hidden on networks that reach GitHub but not registry.npmjs.org. resolveLatestVersion() now tries npm CLI → npm registry → GitHub releases before giving up. (#5266, [bug] Home page "Update Available" banner no longer appears when a newer version exists #4100)
  • command-code: omit max_tokens when the client omits it so the upstream applies the model's native default, fixing 400 "expected <=200000" on /alpha/generate for high-cap models; an explicit oversized client value is clamped to the 200k endpoint ceiling. (#5221 — thanks @adivekar-utexas)

🔒 Security

  • authz: require auth for the /v1beta/* Gemini-compatible client API. next.config.mjs rewrote /v1beta/:path* → /api/v1beta/:path*, but src/proxy.ts didn't match /v1beta before the rewrite and classifyRoute() didn't classify /api/v1beta/* as client API — so unauthenticated /v1beta/models/...:generateContent traffic could reach the model-serving route without the central client-API auth policy. Both alias and rewritten forms are now classified CLIENT_API, enforcing Bearer auth when REQUIRE_API_KEY is enabled. (#5274 — thanks @rdself)
  • sentinel: security hardening pass across request handling. (#5241 — thanks @iamedwardngo)
  • providers: refresh impersonation User-Agents + TLS fingerprint profiles to current real-client versions; several had drifted or were inconsistent across files, a bot-detection/blocking risk. (#5237 — thanks @backryun)
  • authz (public origin): centralize browser-mutation origin validation into src/server/origin/publicOrigin.ts and wire it through the authz pipeline, replacing the per-route same-origin-only check that 403'd dashboard mutations when served behind a reverse proxy on a different public origin. The module resolves the allowed public origin from configured base-URL env vars or trusted forwarded headers (only when OMNIROUTE_TRUST_PROXY is set and the peer is loopback/LAN via peer-stamp), validates Sec-Fetch-Site metadata, and sanitizes Host/Forwarded inputs (rejects control chars, userinfo, path/query in Host). Regression guards: tests/unit/authz/public-origin.test.ts + tests/unit/authz/pipeline.test.ts. (#5278 — thanks @Thinkscape / @abodera)

📝 Maintenance



🙌 Contributors

Thanks to everyone who contributed code, fixes, reports and root-cause diagnoses this cycle:

External contributors & reporters

Maintainer

  • @diegosouzapw — release engineering, reconciliation, and the remaining roadmap/security/maintenance work.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request bumps the version of OmniRoute to 3.8.40 across package manifests, lockfiles, and the OpenAPI specification, while also adding a placeholder for the new version in the internationalized changelogs. The reviewer correctly identified that the new version section was inserted out of chronological order in the changelog files (placed between 3.8.31 and 3.8.39 instead of after 3.8.39), which should be corrected across all language directories.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread docs/i18n/ar/CHANGELOG.md
Comment on lines +9 to +13
## [3.8.40] — TBD

_In development — bullets added per PR; finalized at release._

---

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The new version [3.8.40] is inserted between [3.8.31] (2026-06-20) and [3.8.39] (2026-06-28), which breaks the chronological order of the changelog. Since the changelog is in ascending order, the section for [3.8.40] should be placed after the [3.8.39] section. Please apply this fix to all docs/i18n/*/CHANGELOG.md files as they all share this same ordering issue.

@github-actions

github-actions Bot commented Jun 28, 2026 •

Copy link
Copy Markdown
Contributor

CI Coverage Report

  • Coverage job: success
  • PR test policy: failure

Coverage artifact was not available for this run.

PR Test Policy

This PR changes production code in src/, open-sse/, electron/, or bin/ without accompanying automated tests.

diegosouzapw and others added 26 commits June 28, 2026 10:23
apt-get upgrade -y in the base stage pulls security-patched trixie
packages at build time, and npm install -g npm@latest refreshes the
globally-bundled undici/tar inside the npm CLI. Together these clear the
subset of GitHub container-scan CVE alerts that have an upstream fix
available.

None of the flagged CVEs are in the application dependency tree (app
already resolves undici@8.5.0 / tar@7.5.16, both fixed); they live in
the node:24-trixie-slim base layer and npm's own internals, and none are
reachable from the proxy request surface at runtime. CVEs without a
published fix (local-only TOCTOU, etc.) remain until the distro patches
them and the image is rebuilt.
…#5233)

The ordering assertion matched a builder-stage comment that mentions
'npm run build' (added with the v3.8.40 workspace-deps Docker fix), making it
false-fail even though the real RUN step correctly follows the NODE_OPTIONS
heap line. Match real ENV/RUN instructions only, ignoring '#' comment lines.
…se) (#5235)

The advisory Trivy image scan uploaded every HIGH/CRITICAL into the
Security tab without ignore-unfixed, flooding it with ~150 unfixable
base-image OS CVEs (Debian trixie packages with no upstream patch,
overwhelmingly local-only and not reachable from the proxy request
surface). Operators cannot act on those, so they are pure noise.

Add ignore-unfixed:true to the advisory step so it mirrors the existing
CRITICAL blocking gate and surfaces only actionable, fixable
vulnerabilities. Wire trivyignores to a new repo-root .trivyignore that
documents the accepted-risk policy and is the single auditable home for
the rare fixable CVE we must temporarily accept (none at present).

Takes effect on the next release image build (Trivy only runs on tag
builds, not main pushes); fixed CVEs drop out of the SARIF and GitHub
auto-resolves the corresponding alerts.
Scope textual thinking-tag extraction to tag-native model families; preserve GEMINI_CLI registration. Resubmit of #5216 without the regression. Integrated into release/v3.8.40.
Add lobe provider icons + aliases; rebased onto release tip. Integrated into release/v3.8.40.
Forward MCP HTTP auth to internal tool fetches via AsyncLocalStorage (#5211). Rebased onto release tip. Integrated into release/v3.8.40.
…5230)

serve.mjs imports scripts/build/runtime-env.mjs (added with the #5213 heap
auto-calibration fix) but it was missing from package.json files whitelist,
breaking every global npm install at startup. Add it to files and guard with
a regression test that asserts all bin/ runtime imports of scripts/ are
packaged.
Antigravity: retry excluded accounts via fallback LRU + family-inferred 429 cooldown. Test strengthened into a real LRU regression guard; cooldown constant extracted. Integrated into release/v3.8.40.
Cherry-picked the corrective part of #5221 only: the executor stops fabricating
`max_tokens` (= per-model registry cap) when the client omits it, which caused
`400 "expected <=200000"` on /alpha/generate for high-cap models. An explicit
oversized client value is clamped to the 200k endpoint ceiling. The PR's registry
maxOutputTokens recaps (open-sse/config/providers/registry/command-code/index.ts)
are intentionally NOT included pending reconciliation; #5221 stays open for that.
…stry caps (#5221)

Integrated into release/v3.8.40 — corrective max_tokens part already cherry-picked (e8d13ec); this brings the registry maxOutputTokens caps. Thanks @adivekar-utexas.
…urrent client versions (#5237)

Refresh impersonation UAs + TLS profiles to current client versions. Review fixes (perplexity UA kept on Firefox 148 to match firefox_148 TLS profile; golden snapshot regenerated; release file-size/complexity baselines reconciled) co-authored. Thanks @backryun.
…ossy compression (#5243)

Risk-gate pre-pass — shields sensitive spans (PEM/secret/stack/k8s/migration/legal) from lossy compression via SENTINEL preserveSpans. Default off, fail-open, ReDoS-bounded patterns. strategySelector baseline rebaselined for the wrapper extraction.
Harden docs i18n rendering: path-traversal guard (cookie-controlled locale confined to docs/i18n via pure resolveSafeI18nSectionDir) + markdown XSS sanitization (DOMPurify allowlist). Review fixes: declared dompurify dep + allowlist, cleaned sanitizer config, extracted+tested the real path helper, removed agent scratch. Thanks @iamedwardngo!
…#5248)

sessionStickiness.ts (v3.8.36) keeps a sessionless multi-turn conversation pinned
to the same connection so the upstream prompt-cache stays warm, but the
round-robin handler (handleRoundRobinCombo) returned before reaching the
applySessionStickiness call used by the weighted/priority paths. Clients that send
no session id (Codex CLI, Claude Code, most OpenAI-compatible tools) therefore had
round-robin combos rotate to a different connection every turn -> prompt-cache
miss -> cold high-reasoning starts, intermittent 504s and throughput collapse
under concurrency (#3825).

Reuse the existing mechanism: when a sticky connection is bound to the
conversation and present in the current targets, start the round-robin rotation at
it (failover to the other targets is preserved), and (re)record the binding on
success. Different conversations still spread across connections on their first
turn -- only intra-conversation rotation is removed. A diagnostic load harness
confirmed request-dispatch CPU is not the bottleneck; this is a routing fix.

Adds an integration regression test driving the real handleComboChat: a sessionless
round-robin conversation re-pins across turns (fails before the fix), and distinct
conversations still spread (round-robin distribution preserved).
Integrated into release/v3.8.40 — codex local_shell drop validated (merge-result: eslint clean, 41/41 tests). FQG failure was stale base.
Integrated into release/v3.8.40 — Gemini CLI channel removed; Antigravity-path parseTextualReasoningTags guard restored (regression caught + fixed in-place, all antigravity/gemini tests green).
Integrated into release/v3.8.40 — provider request headers preserved in logs; verified combo reads native Response.headers (no regression).
Integrated into release/v3.8.40 — deprecated stub types removed, Trivy ignore-unfixed, js-yaml override; allowlist + file-size baseline reconciled.
* feat(kilocode): anonymous no-auth access to free models (#4019)

Kilo's gateway serves its free tier without signup: an OpenAI-compatible
request to api.kilo.ai/api/openrouter authenticated with the literal API
key `anonymous` (Authorization: Bearer anonymous) plus an
X-KILOCODE-EDITORNAME header returns free models. Expose it the same way
opencode/mimocode do:

- flag the kilocode dashboard provider `anonymousFallback: true` so a
  request with no connected account synthesizes a noauth credential;
- add `anonymousApiKey` to the registry and have DefaultExecutor send it
  as the bearer token only when no real credential exists, so the OAuth
  paid path is untouched.

Regression test exercises buildHeaders for the anonymous, OAuth-token and
API-key cases plus the anonymousFallback flag (fails before the fix).
Reference implementation pointer courtesy of @Theadd (#4019).

* chore(quality): reconcile inherited file-size base-red for executor-codex.test.ts

The release tip carries tests/unit/executor-codex.test.ts at 1347 lines while
the frozen baseline still reads 1340 — 7 legit lines were added without
ratcheting, so check:file-size fails for every PR that branches off the tip
(proven: the gate already fails on HEAD~1, before this PR's #4019 commit).
Ratchet the testFrozen entry to the current count to unblock; the file itself
is untouched by this PR.

* test: regenerate provider translate-path golden for kilocode editor-name header (#4019)

The #4019 commit adds X-KILOCODE-EDITORNAME to the kilocode registry headers.
The all-providers translate-path golden snapshots per-provider headers, so it
must be regenerated. Diff is exactly the new header on kilocode's apiKey/oauth/
nonStream variants — Authorization (Bearer <TOK>) and every other provider are
unchanged.
Integrated into release/v3.8.40 — proxy fallback cache scoped by target URL (prevents cross-endpoint poisoning).
Integrated into release/v3.8.40 — shell tool kept caller-side in Chat→Responses translation (complements #5250).
…5258)

Integrated into release/v3.8.40 — grok-cli import-token accepts full auth.json object; zod z.record key-type fixed; duplicate Docker hardening dropped (already in release).
fulorgnas and others added 13 commits June 28, 2026 22:20
…e sending (#5273)

Strip unsupported sampling params (presencePenalty/frequencyPenalty/logprobs/topLogprobs) before forwarding to Grok Build. Added regression test (Rule #18, TDD-verified). Integrated into release/v3.8.40.
…elease) (#5282)

Break the tools.ts ↔ toolSearch.ts import cycle (check:cycles red on release) via a leaf toolDefinition.ts. Integrated into release/v3.8.40.
…dmap #16 (#5283)

roadmap #16: composite-command splitter for RTK detection (opt-in, fail-open). Integrated into release/v3.8.40.
Drop retired chatgpt-web/codex models (gpt-5.2, gpt-4.5) — registry/map/UI/docs/tests kept aligned. Integrated into release/v3.8.40.
…5279)

Call-log correlation ID core (migration 109 + storage + X-Correlation-Id header) — safe subset of #5275. Original work by @hartmark. Integrated into release/v3.8.40.
#13 (#5285)

roadmap #13: saliency heatmap in the compression studio (opt-in, dry-run). Locally validated release-green (typecheck + heatmap tests + file-size). Integrated into release/v3.8.40.
roadmap #18: inline transparency annotation. Fixed a live 500 (→/× in the X-OmniRoute-Compression latin-1 header → ByteString throw) → ASCII-only + regression test building real Headers/Response. Locally validated release-green. Integrated into release/v3.8.40.
…in) — roadmap #21 (#5286)

roadmap #21: result memoization for deterministic engines (opt-in, default off). Fixed the cache-key under-specification (folded model+supportsVision so lite image-strip can't serve a wrong cached body across vision/non-vision targets) + regression test. Locally validated release-green. Integrated into release/v3.8.40.
… enable (#5183)

Partition no-auth entries instead of dropping blocked ones; disabled no-auth providers are surfaced in a Disabled group with an in-place Enable button (#5166/#5183). Extracted NoAuthProvidersSection to keep page.tsx under its size freeze.

Admin-merged: the only red checks (Unit/Coverage shard 2/8, Node-compat) are the pre-existing #4076 Dockerfile heap-ordering base-red on main — unrelated to this change (proven: fails locally on a branch that does not touch the Dockerfile or that test) and already de-brittled in the v3.8.40 release line.
…#17 (#5288)

* test(compression): TDD failing tests for hard-budget post-pass (#17)

node:test suite for applyHardBudget — targetTokens, targetRatio, no-op,
force-preserve, determinism, both-wins, techniquesUsed, integration seam.
All tests fail (module not yet created) — TDD red step.

* feat(compression): hard-budget post-pass (#17) — compress to exactly N tokens

- types.ts: add targetTokens? and targetRatio? to CompressionConfig after contextBudget
- hardBudget.ts: applyHardBudget(body, {targetTokens?,targetRatio?}) → CompressionResult;
  splits prose into sentences/lines, ranks by avg scoreToken ascending, drops lowest-
  saliency units until ≤ target; UNIT_PRESERVE_RE guards numbers/URLs/errors/code;
  targetTokens wins when both set; techniquesUsed:["hard-budget"]
- strategySelector.ts: hard-budget post-pass in runStackedCompression +
  runStackedCompressionAsync after engine loop, before finalizeStackedResult;
  gated on config.targetTokens || config.targetRatio; mergeStackStep + compressed=true

* fix(compression): preserve digit-less sensitive lines in hard-budget

UNIT_PRESERVE_RE only matched digits/URLs/error-headers/code-fences, so
stack-trace at-frames, key=value credential lines, and digit-less paths
were droppable and could be cut to hit the budget. Add specific anchors
(^\s*at\s, \/[\w.-]+\/, [A-Za-z_]\w*=\S) that never match a bare
end-of-sentence period (which would re-break the feature into a no-op).

Regression tests drive each unit to target=1 (drops every non-preserved
unit) and assert the sensitive line survives; plus a guard that plain
prose ending in a period stays droppable.

* fix(compression): distribute hard-budget across messages + warn when unreachable

Two related correctness fixes in applyHardBudget:

- Aggregate target was passed verbatim to compressText for EACH message,
  so an N-message body could come back ~N× over budget. Distribute the
  target proportionally per message (floor(target * msgTokens/total)) so
  the SUM stays <= target.

- When every unit is preserve-guarded (or a single oversized preserved
  unit), the result still exceeds target with no signal. Measure the
  result and push a validationWarnings entry when it remains over budget,
  so callers are not silently left over the limit.

Regression tests: a 4-message body with target 200 ends with TOTAL <= 200;
an all-numeric (all-preserved) body emits the 'could not reach target'
warning.

* fix(compression): run hard-budget post-pass when targetTokens/targetRatio is 0

The seam gate used `||`, so targetTokens:0 or targetRatio:0 (both falsy)
silently skipped the post-pass in runStackedCompression and
runStackedCompressionAsync. Switch to `!= null` so an explicit 0 still
engages the pass.

Regression test: applyStackedCompression with config.targetTokens:0 must
report 'hard-budget' in techniquesUsed.

* docs(changelog): restore hard-budget bullet (#17, eaten by rebase)
* test(compression): failing tests for relevance engine scorer + apply

* feat(compression): relevance extractive engine — scores sentences against last user query

* test(compression): regression tests for relevance engine review fixes (threshold/multimodal/force-preserve/whitespace/query)

* fix(compression): relevance engine hardening from core review

- overlapThreshold was dead config (|| kept<budget made it unreachable) → real threshold gate
- multimodal: only compress when exactly one text block (was stamping joined text into every block)
- force-preserve sentences are now 'free' (don't consume budget) so they can't starve top-relevance
- preserve inter-sentence whitespace (\n\n survives) instead of flattening to single spaces
- ROOT-CAUSE: stop gating preserve on ultraHeuristic FORCE_PRESERVE_RE — it matches the period
  ending every sentence, so it force-preserved everything (no-op). New SENTENCE_PRESERVE_RE anchors
  on real signals (digits/URL/Error:/code/at-frame/path/key=value), mirroring #17's UNIT_PRESERVE_RE
- core-review suggestion to skip the query message: rejected (self-overlap already protects it;
  skipping would no-op the single-message RAG case) — documented + test updated

* docs(changelog): restore relevance engine bullet (#7, eaten by rebase)
A 400 classified as MODEL_CAPACITY (e.g. 'requested model is not supported') hit the #2101 anti-loop stop-branch and halted the combo instead of advancing. Drop the MODEL_CAPACITY trigger from that stop condition so model-specific 400s fall through to the next combo target; genuinely body-specific 400s (malformed/invalid/bad-request substrings) still stop. Validated: combo-strategies suite 16/16 (incl. the new regression test) green on current release tip, typecheck clean. Stale pre-merge CI was from an older base.

Co-authored-by: Chewji9875 <Chewji9875@users.noreply.github.com>
…rough the stacked seam (#5291)

Follow-up to #5288: propagate the hard-budget unreachable-target warning through the stacked seam (else branch in both sync/async paths) + TDD propagation test. Integrated into release/v3.8.40.
Comment thread open-sse/services/compression/resultMemo.ts Dismissed
diegosouzapw and others added 14 commits June 29, 2026 02:57
#5300)

#5249 deliberately changed a MODEL_CAPACITY 400 ('model X not supported with
this account') from STOP to ADVANCE-to-next-combo-target, since a different
model in the combo may be supported. It added a regression test in
combo-strategies.test.ts but did not update the separate
combo-body-specific-400-stop-4279.test.ts, which still asserted the old STOP
behavior for that exact model-not-supported text — leaving the release branch
red on the Unit 2/2 shard (every PR inherited the failure).

Re-point the #4279 test at a genuinely body-specific malformed 400 ('invalid
message format'), which still recurs identically on every target and must STOP.
This preserves the #2101 anti-loop {ok,response} contract coverage while the
advance-on-model-400 path stays covered by combo-strategies.test.ts.

Test-only; no production behavior change.
…ve value (#5166) (#5304)

Integrated into release/v3.8.40 (CHANGELOG re-resolved against post-#5294 tip; code identical to the green commit)
#5231) (#5303)

Integrated into release/v3.8.40 (CHANGELOG re-resolved against post-#5294/#5304 tip; kiro code identical to the green commit)
….gamificationGroup i18n (#5298) (#5306)

Integrated into release/v3.8.40
…SSE when Accept lists application/json (#5305) (#5309)

Integrated into release/v3.8.40
On a `release: released` event the freshly-created git tag is often not yet
visible to `git fetch --tags` when docker-publish runs, so the :latest-promotion
gate built its candidate set purely from `git tag -l` and resolved the highest
semver to the previous version — leaving :latest one release behind.

Extract the decision into scripts/ci/should-promote-latest.sh, which folds the
current VERSION into the candidate set before picking the highest stable semver,
making promotion independent of tag-sync timing. A patch published after a higher
minor still won't grab :latest.

Regression guard: tests/unit/build/should-promote-latest-5301.test.ts (spawns the
real helper with fixture tag lists; proves the race case now promotes).
… (union of REQUIRE_API_KEY + publicOrigin env restores); changelog bullet for #5278 public-origin centralization
…with #5278/#5309

- provider-health-autopilot: cross-site mutation rejection moved from the route
  handler into the authz pipeline (#5278); drive the assertion through
  runAuthzPipeline (the real enforcement point) → 403 + connection untouched.
- chat-pipeline: a mixed 'application/json, text/event-stream' Accept now resolves
  to JSON (#5305/#5309 Vercel/OpenAI SDK non-stream signature); the SSE-opt-in test
  now sends a pure 'text/event-stream' Accept, the case #5309 keeps as streaming.
…drift

+13 inherited cycle drift (any warn-allowed in open-sse/ + tests/) measured via
quality:collect on the release tip including reconciliation; my regression test
fixes add 0 new warnings. Same precedent as prior cycle-close rebaselines.
@diegosouzapw
diegosouzapw merged commit 7c23dab into main Jun 29, 2026
13 checks passed
@sonarqubecloud

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
2 Security Hotspots
0.0% Coverage on New Code (required ≥ 80%)
B Maintainability Rating on New Code (required ≥ A)
D Security Rating on New Code (required ≥ A)
D Reliability Rating on New Code (required ≥ A)

See analysis details on SonarQube Cloud

Catch issues before they fail your Quality Gate with our IDE extension SonarQube for IDE

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.