Release v3.8.40 - #5225
Release v3.8.40#5225
Conversation
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
There was a problem hiding this comment.
Code Review
This pull request bumps the version of OmniRoute to 3.8.40 across package manifests, lockfiles, and the OpenAPI specification, while also adding a placeholder for the new version in the internationalized changelogs. The reviewer correctly identified that the new version section was inserted out of chronological order in the changelog files (placed between 3.8.31 and 3.8.39 instead of after 3.8.39), which should be corrected across all language directories.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| ## [3.8.40] — TBD | ||
|
|
||
| _In development — bullets added per PR; finalized at release._ | ||
|
|
||
| --- |
There was a problem hiding this comment.
The new version [3.8.40] is inserted between [3.8.31] (2026-06-20) and [3.8.39] (2026-06-28), which breaks the chronological order of the changelog. Since the changelog is in ascending order, the section for [3.8.40] should be placed after the [3.8.39] section. Please apply this fix to all docs/i18n/*/CHANGELOG.md files as they all share this same ordering issue.
CI Coverage Report
Coverage artifact was not available for this run. PR Test PolicyThis PR changes production code in |
apt-get upgrade -y in the base stage pulls security-patched trixie packages at build time, and npm install -g npm@latest refreshes the globally-bundled undici/tar inside the npm CLI. Together these clear the subset of GitHub container-scan CVE alerts that have an upstream fix available. None of the flagged CVEs are in the application dependency tree (app already resolves undici@8.5.0 / tar@7.5.16, both fixed); they live in the node:24-trixie-slim base layer and npm's own internals, and none are reachable from the proxy request surface at runtime. CVEs without a published fix (local-only TOCTOU, etc.) remain until the distro patches them and the image is rebuilt.
…#5233) The ordering assertion matched a builder-stage comment that mentions 'npm run build' (added with the v3.8.40 workspace-deps Docker fix), making it false-fail even though the real RUN step correctly follows the NODE_OPTIONS heap line. Match real ENV/RUN instructions only, ignoring '#' comment lines.
…se) (#5235) The advisory Trivy image scan uploaded every HIGH/CRITICAL into the Security tab without ignore-unfixed, flooding it with ~150 unfixable base-image OS CVEs (Debian trixie packages with no upstream patch, overwhelmingly local-only and not reachable from the proxy request surface). Operators cannot act on those, so they are pure noise. Add ignore-unfixed:true to the advisory step so it mirrors the existing CRITICAL blocking gate and surfaces only actionable, fixable vulnerabilities. Wire trivyignores to a new repo-root .trivyignore that documents the accepted-risk policy and is the single auditable home for the rare fixable CVE we must temporarily accept (none at present). Takes effect on the next release image build (Trivy only runs on tag builds, not main pushes); fixed CVEs drop out of the SARIF and GitHub auto-resolves the corresponding alerts.
Scope textual thinking-tag extraction to tag-native model families; preserve GEMINI_CLI registration. Resubmit of #5216 without the regression. Integrated into release/v3.8.40.
Add lobe provider icons + aliases; rebased onto release tip. Integrated into release/v3.8.40.
Forward MCP HTTP auth to internal tool fetches via AsyncLocalStorage (#5211). Rebased onto release tip. Integrated into release/v3.8.40.
…5230) serve.mjs imports scripts/build/runtime-env.mjs (added with the #5213 heap auto-calibration fix) but it was missing from package.json files whitelist, breaking every global npm install at startup. Add it to files and guard with a regression test that asserts all bin/ runtime imports of scripts/ are packaged.
Antigravity: retry excluded accounts via fallback LRU + family-inferred 429 cooldown. Test strengthened into a real LRU regression guard; cooldown constant extracted. Integrated into release/v3.8.40.
Cherry-picked the corrective part of #5221 only: the executor stops fabricating `max_tokens` (= per-model registry cap) when the client omits it, which caused `400 "expected <=200000"` on /alpha/generate for high-cap models. An explicit oversized client value is clamped to the 200k endpoint ceiling. The PR's registry maxOutputTokens recaps (open-sse/config/providers/registry/command-code/index.ts) are intentionally NOT included pending reconciliation; #5221 stays open for that.
…stry caps (#5221) Integrated into release/v3.8.40 — corrective max_tokens part already cherry-picked (e8d13ec); this brings the registry maxOutputTokens caps. Thanks @adivekar-utexas.
…ossy compression (#5243) Risk-gate pre-pass — shields sensitive spans (PEM/secret/stack/k8s/migration/legal) from lossy compression via SENTINEL preserveSpans. Default off, fail-open, ReDoS-bounded patterns. strategySelector baseline rebaselined for the wrapper extraction.
Harden docs i18n rendering: path-traversal guard (cookie-controlled locale confined to docs/i18n via pure resolveSafeI18nSectionDir) + markdown XSS sanitization (DOMPurify allowlist). Review fixes: declared dompurify dep + allowlist, cleaned sanitizer config, extracted+tested the real path helper, removed agent scratch. Thanks @iamedwardngo!
…#5248) sessionStickiness.ts (v3.8.36) keeps a sessionless multi-turn conversation pinned to the same connection so the upstream prompt-cache stays warm, but the round-robin handler (handleRoundRobinCombo) returned before reaching the applySessionStickiness call used by the weighted/priority paths. Clients that send no session id (Codex CLI, Claude Code, most OpenAI-compatible tools) therefore had round-robin combos rotate to a different connection every turn -> prompt-cache miss -> cold high-reasoning starts, intermittent 504s and throughput collapse under concurrency (#3825). Reuse the existing mechanism: when a sticky connection is bound to the conversation and present in the current targets, start the round-robin rotation at it (failover to the other targets is preserved), and (re)record the binding on success. Different conversations still spread across connections on their first turn -- only intra-conversation rotation is removed. A diagnostic load harness confirmed request-dispatch CPU is not the bottleneck; this is a routing fix. Adds an integration regression test driving the real handleComboChat: a sessionless round-robin conversation re-pins across turns (fails before the fix), and distinct conversations still spread (round-robin distribution preserved).
Integrated into release/v3.8.40
Integrated into release/v3.8.40 — codex local_shell drop validated (merge-result: eslint clean, 41/41 tests). FQG failure was stale base.
Integrated into release/v3.8.40 — Gemini CLI channel removed; Antigravity-path parseTextualReasoningTags guard restored (regression caught + fixed in-place, all antigravity/gemini tests green).
Integrated into release/v3.8.40
Integrated into release/v3.8.40 — provider request headers preserved in logs; verified combo reads native Response.headers (no regression).
Integrated into release/v3.8.40 — deprecated stub types removed, Trivy ignore-unfixed, js-yaml override; allowlist + file-size baseline reconciled.
* feat(kilocode): anonymous no-auth access to free models (#4019) Kilo's gateway serves its free tier without signup: an OpenAI-compatible request to api.kilo.ai/api/openrouter authenticated with the literal API key `anonymous` (Authorization: Bearer anonymous) plus an X-KILOCODE-EDITORNAME header returns free models. Expose it the same way opencode/mimocode do: - flag the kilocode dashboard provider `anonymousFallback: true` so a request with no connected account synthesizes a noauth credential; - add `anonymousApiKey` to the registry and have DefaultExecutor send it as the bearer token only when no real credential exists, so the OAuth paid path is untouched. Regression test exercises buildHeaders for the anonymous, OAuth-token and API-key cases plus the anonymousFallback flag (fails before the fix). Reference implementation pointer courtesy of @Theadd (#4019). * chore(quality): reconcile inherited file-size base-red for executor-codex.test.ts The release tip carries tests/unit/executor-codex.test.ts at 1347 lines while the frozen baseline still reads 1340 — 7 legit lines were added without ratcheting, so check:file-size fails for every PR that branches off the tip (proven: the gate already fails on HEAD~1, before this PR's #4019 commit). Ratchet the testFrozen entry to the current count to unblock; the file itself is untouched by this PR. * test: regenerate provider translate-path golden for kilocode editor-name header (#4019) The #4019 commit adds X-KILOCODE-EDITORNAME to the kilocode registry headers. The all-providers translate-path golden snapshots per-provider headers, so it must be regenerated. Diff is exactly the new header on kilocode's apiKey/oauth/ nonStream variants — Authorization (Bearer <TOK>) and every other provider are unchanged.
Integrated into release/v3.8.40 — proxy fallback cache scoped by target URL (prevents cross-endpoint poisoning).
Integrated into release/v3.8.40 — shell tool kept caller-side in Chat→Responses translation (complements #5250).
…5258) Integrated into release/v3.8.40 — grok-cli import-token accepts full auth.json object; zod z.record key-type fixed; duplicate Docker hardening dropped (already in release).
…elease) (#5282) Break the tools.ts ↔ toolSearch.ts import cycle (check:cycles red on release) via a leaf toolDefinition.ts. Integrated into release/v3.8.40.
Drop retired chatgpt-web/codex models (gpt-5.2, gpt-4.5) — registry/map/UI/docs/tests kept aligned. Integrated into release/v3.8.40.
roadmap #18: inline transparency annotation. Fixed a live 500 (→/× in the X-OmniRoute-Compression latin-1 header → ByteString throw) → ASCII-only + regression test building real Headers/Response. Locally validated release-green. Integrated into release/v3.8.40.
…in) — roadmap #21 (#5286) roadmap #21: result memoization for deterministic engines (opt-in, default off). Fixed the cache-key under-specification (folded model+supportsVision so lite image-strip can't serve a wrong cached body across vision/non-vision targets) + regression test. Locally validated release-green. Integrated into release/v3.8.40.
… enable (#5183) Partition no-auth entries instead of dropping blocked ones; disabled no-auth providers are surfaced in a Disabled group with an in-place Enable button (#5166/#5183). Extracted NoAuthProvidersSection to keep page.tsx under its size freeze. Admin-merged: the only red checks (Unit/Coverage shard 2/8, Node-compat) are the pre-existing #4076 Dockerfile heap-ordering base-red on main — unrelated to this change (proven: fails locally on a branch that does not touch the Dockerfile or that test) and already de-brittled in the v3.8.40 release line.
…#17 (#5288) * test(compression): TDD failing tests for hard-budget post-pass (#17) node:test suite for applyHardBudget — targetTokens, targetRatio, no-op, force-preserve, determinism, both-wins, techniquesUsed, integration seam. All tests fail (module not yet created) — TDD red step. * feat(compression): hard-budget post-pass (#17) — compress to exactly N tokens - types.ts: add targetTokens? and targetRatio? to CompressionConfig after contextBudget - hardBudget.ts: applyHardBudget(body, {targetTokens?,targetRatio?}) → CompressionResult; splits prose into sentences/lines, ranks by avg scoreToken ascending, drops lowest- saliency units until ≤ target; UNIT_PRESERVE_RE guards numbers/URLs/errors/code; targetTokens wins when both set; techniquesUsed:["hard-budget"] - strategySelector.ts: hard-budget post-pass in runStackedCompression + runStackedCompressionAsync after engine loop, before finalizeStackedResult; gated on config.targetTokens || config.targetRatio; mergeStackStep + compressed=true * fix(compression): preserve digit-less sensitive lines in hard-budget UNIT_PRESERVE_RE only matched digits/URLs/error-headers/code-fences, so stack-trace at-frames, key=value credential lines, and digit-less paths were droppable and could be cut to hit the budget. Add specific anchors (^\s*at\s, \/[\w.-]+\/, [A-Za-z_]\w*=\S) that never match a bare end-of-sentence period (which would re-break the feature into a no-op). Regression tests drive each unit to target=1 (drops every non-preserved unit) and assert the sensitive line survives; plus a guard that plain prose ending in a period stays droppable. * fix(compression): distribute hard-budget across messages + warn when unreachable Two related correctness fixes in applyHardBudget: - Aggregate target was passed verbatim to compressText for EACH message, so an N-message body could come back ~N× over budget. Distribute the target proportionally per message (floor(target * msgTokens/total)) so the SUM stays <= target. - When every unit is preserve-guarded (or a single oversized preserved unit), the result still exceeds target with no signal. Measure the result and push a validationWarnings entry when it remains over budget, so callers are not silently left over the limit. Regression tests: a 4-message body with target 200 ends with TOTAL <= 200; an all-numeric (all-preserved) body emits the 'could not reach target' warning. * fix(compression): run hard-budget post-pass when targetTokens/targetRatio is 0 The seam gate used `||`, so targetTokens:0 or targetRatio:0 (both falsy) silently skipped the post-pass in runStackedCompression and runStackedCompressionAsync. Switch to `!= null` so an explicit 0 still engages the pass. Regression test: applyStackedCompression with config.targetTokens:0 must report 'hard-budget' in techniquesUsed. * docs(changelog): restore hard-budget bullet (#17, eaten by rebase)
* test(compression): failing tests for relevance engine scorer + apply * feat(compression): relevance extractive engine — scores sentences against last user query * test(compression): regression tests for relevance engine review fixes (threshold/multimodal/force-preserve/whitespace/query) * fix(compression): relevance engine hardening from core review - overlapThreshold was dead config (|| kept<budget made it unreachable) → real threshold gate - multimodal: only compress when exactly one text block (was stamping joined text into every block) - force-preserve sentences are now 'free' (don't consume budget) so they can't starve top-relevance - preserve inter-sentence whitespace (\n\n survives) instead of flattening to single spaces - ROOT-CAUSE: stop gating preserve on ultraHeuristic FORCE_PRESERVE_RE — it matches the period ending every sentence, so it force-preserved everything (no-op). New SENTENCE_PRESERVE_RE anchors on real signals (digits/URL/Error:/code/at-frame/path/key=value), mirroring #17's UNIT_PRESERVE_RE - core-review suggestion to skip the query message: rejected (self-overlap already protects it; skipping would no-op the single-message RAG case) — documented + test updated * docs(changelog): restore relevance engine bullet (#7, eaten by rebase)
A 400 classified as MODEL_CAPACITY (e.g. 'requested model is not supported') hit the #2101 anti-loop stop-branch and halted the combo instead of advancing. Drop the MODEL_CAPACITY trigger from that stop condition so model-specific 400s fall through to the next combo target; genuinely body-specific 400s (malformed/invalid/bad-request substrings) still stop. Validated: combo-strategies suite 16/16 (incl. the new regression test) green on current release tip, typecheck clean. Stale pre-merge CI was from an older base. Co-authored-by: Chewji9875 <Chewji9875@users.noreply.github.com>
#5300) #5249 deliberately changed a MODEL_CAPACITY 400 ('model X not supported with this account') from STOP to ADVANCE-to-next-combo-target, since a different model in the combo may be supported. It added a regression test in combo-strategies.test.ts but did not update the separate combo-body-specific-400-stop-4279.test.ts, which still asserted the old STOP behavior for that exact model-not-supported text — leaving the release branch red on the Unit 2/2 shard (every PR inherited the failure). Re-point the #4279 test at a genuinely body-specific malformed 400 ('invalid message format'), which still recurs identically on every target and must STOP. This preserves the #2101 anti-loop {ok,response} contract coverage while the advance-on-model-400 path stays covered by combo-strategies.test.ts. Test-only; no production behavior change.
…credit every contributor
Integrated into release/v3.8.40
Integrated into release/v3.8.40
On a `release: released` event the freshly-created git tag is often not yet visible to `git fetch --tags` when docker-publish runs, so the :latest-promotion gate built its candidate set purely from `git tag -l` and resolved the highest semver to the previous version — leaving :latest one release behind. Extract the decision into scripts/ci/should-promote-latest.sh, which folds the current VERSION into the candidate set before picking the highest stable semver, making promotion independent of tag-sync timing. A patch published after a higher minor still won't grab :latest. Regression guard: tests/unit/build/should-promote-latest-5301.test.ts (spawns the real helper with fixture tag lists; proves the race case now promotes).
… (union of REQUIRE_API_KEY + publicOrigin env restores); changelog bullet for #5278 public-origin centralization
…with #5278/#5309 - provider-health-autopilot: cross-site mutation rejection moved from the route handler into the authz pipeline (#5278); drive the assertion through runAuthzPipeline (the real enforcement point) → 403 + connection untouched. - chat-pipeline: a mixed 'application/json, text/event-stream' Accept now resolves to JSON (#5305/#5309 Vercel/OpenAI SDK non-stream signature); the SSE-opt-in test now sends a pure 'text/event-stream' Accept, the case #5309 keeps as streaming.
…drift +13 inherited cycle drift (any warn-allowed in open-sse/ + tests/) measured via quality:collect on the release tip including reconciliation; my regression test fixes add 0 new warnings. Same precedent as prior cycle-close rebaselines.
|




Release v3.8.40
Integration PR for the v3.8.40 cycle →
main. Living release notes — mirrors the finalizedCHANGELOG.md[3.8.40]section.origin/mainmerged in to resolve cycle drift (#5278/#5234/#5228 landed directly on main).[3.8.40] — TBD
In development — bullets added per PR; finalized at release.
✨ New Features
RegExpover user input), fail-open, default off. Ideal for trimming long pasted RAG context / tool output to what's relevant. Sentences carrying real signal (digits/URLs/errors/code/paths) are never dropped;overlapThreshold/budgetPercent/boilerplateWeightare configurable. Tier-2 item of the compression feature-extraction roadmap (chore(env): sync .env.example with current .env structure #7). (#5289)targetTokens/targetRatio, default unset → no-op) that trims a body to a token budget. It ranks sentences/lines by averagescoreTokenascending and drops the lowest-saliency ones until the body fits (measured by the exact cl100kcountTextTokens), preserving original order. Lines carrying real signal (digits, URLs,Error:-family, code fences, stackat-frames, multi-segment paths,key=value) are never dropped; the budget is distributed proportionally across messages so the total stays ≤ target; an unreachable target (all-preserved) surfaces avalidationWarningsnote instead of failing silently. Does NOT touch theestimateCompressionTokensbudget-gate estimator. Tier-3 item of the compression feature-extraction roadmap (fix(ci): explicit .npmrc auth for npm publish #17). (#5288, follow-up #5291)(input, config) → resultfor provably pure, stateless modes (lite/standard/rtkand stacked pipelines of{lite,caveman,rtk}) to skip recompute on the hot path. Opt-in viamemoizeCompressionResults(default off → zero behavior change). Conservative opt-in whitelist (statefulccr/session-dedup— which write the cross-request CCR store — and model-backedultra/aggressive/llmlinguaare never cached), principal-scoped (skipped without a principal, so no cross-principal body leak), and clone-on-store + clone-on-read. Tier-3 item of the compression feature-extraction roadmap (feat(security): FASE-01 to FASE-09 — Security Hardening & Advanced Features #21). (#5286)tokens=847→312; rules: filler×8, dedup×2derived from existing compression stats. TheX-OmniRoute-Compressionresponse header is extended append-only (themode; source=Xprefix stays byte-identical, so existing header parsers don't break) and the compression studio cockpit shows a matching badge. Zero new computation — it aggregates therulesApplied/techniquesUsedalready on the stats. Tier-3 item of the compression feature-extraction roadmap (fix(ci): add environment for npm token access #18). (#5284)ultraper-tokenscoreToken(0–1, green→red gradient) or universal kept/removed from the existing diff. A dry-run visualization behind a toggle (no cost on a normal preview; backward-compatible when off). Completes the visualization half of roadmap item deps: bump qs from 6.14.1 to 6.14.2 #13 (the A/B comparison shipped in #5080). (#5285)cd /x && git statusnow detects asgit-status(previously the whole string was treated as one command and matched no filter). A quote-aware top-level tokenizer splits on&&/||/;(never inside quotes or$(…)/backtick subshells) and feeds the last segment to RTK command detection, so every RTK filter/renderer fires on commands wrapped incd … &&/||/;chains. O(n), no RegExp over the command (ReDoS-safe). Tier-3 item of the compression feature-extraction roadmap (fix(ci): fix npm publish auth — support vars.NPM_TOKEN #16). (#5283)omniroute_tool_searchtool + one-line TS signatures — new MCP tool that does lexical keyword search over every MCP tool's name/description and returns the top matches as compact one-line TypeScript signatures (~half the JSON-schema token cost), so agents discover tools on demand instead of carrying all ~88 schemas every turn. Search is ReDoS-safe (substring scoring, nevernew RegExpon the query) and deterministic;tools/liststays complete (no hidden tools). Adds theread:toolsscope. Tier-1 item of the compression feature-extraction roadmap. (#5269)git diff→ file headers +@@hunks + changed lines only; an all-greenpytest/jest/vitest/eslintrun → its one-line summary;terraform/tofu plan→Plan: +N ~M -Kplus the resource list;kubectl/awsJSON arrays → a minimal table. Each renderer is conservative (no-op when the shape doesn't match) and the integration is fail-open; the test-green renderer never collapses output that carries any failure signal. Gated byRtkConfig.enableRenderers(default off → zero behavioral change). Eighth item of the compression feature-extraction roadmap. (#5268)role:systemmessage every turn, the longest common prefix across turns ends at that changing byte → the whole system prompt after it is re-billed and re-processed each turn. QuantumLock replaces each non-semantic volatile fragment with a positional, value-independent placeholder⟦Q{i}⟧and appends the real values in a delimited⟦QUANTUMLOCK⟧tail. The rewrite is sent to the model (lossless — not restored), so the system-prompt body becomes byte-identical across turns and the provider caches the long stable prefix while only the small tail differs. Opt-in, default off, applied only for caching providers (isCachingProvider && config.quantumLock.enabled); bounded ReDoS-safe patterns; idempotent; no date/time patterns (semantically meaningful — explicit non-goal). Studio gets a toggle + a "🔒 N volatile fragment(s) stabilized" dry-run badge. Seventh item of the compression feature-extraction roadmap (bench: #5080, gate: #5127, fuzzy: #5143, ionizer: #5148, TOON: #5163, CCR ranged: #5187, risk-gate: #5243). (#5260)opencode/mimocodepattern. With no Kilo account connected, requests now fall back to the gateway's anonymous tier (Authorization: Bearer anonymousonapi.kilo.ai/api/openrouter) so the free models work without signup; a connected OAuth account is still used unchanged for the paid tier (#5259, [feature] Anonymous (no-auth) usage for Kilo Code free models, like opencode #4019 — thanks @Theadd for the reference implementation)X-Correlation-Idresponse header, persisted incall_logs(migration 109), filterable via/api/usage/call-logs, and surfaced in the dashboard request logger (per-chunk stream timestamps + active-requests-first sort). This is the safe, cohesive core subset of the larger feat: add CorrelationId and fix lazy loading [wating autor] #5275 — landed on its own so the low-risk value isn't blocked by the parts of that PR still under review. (#5279 — thanks @hartmark)copilot-m365-webprovider (the 237th), wiring the M365 BizChat framing/connection helpers into a selectable web-session provider backed bym365.cloud.microsoft/chatfor individual Microsoft 365 plans. Builds on the M365 pure-framing groundwork from feat(executors): land M365 Copilot pure framing + connection helpers (#4042) #4696. Regression guard:tests/unit/copilot-m365-web-executor.test.ts. (#5302 — thanks @skyzea1)🔧 Bug Fixes
:latest(and:latest-web) tags to the just-published release. On arelease: releasedevent the freshly-created git tag is often not yet visible togit fetch --tagswhendocker-publishruns, so the:latest-promotion gate built its candidate set purely fromgit tag -land resolved the highest semver to the previous version — leavinglatestone release behind (3.8.39 published,lateststill 3.8.38). The decision now lives inscripts/ci/should-promote-latest.sh, which folds the currentVERSIONinto the candidate set before picking the highest stable semver, making promotion independent of tag-sync timing (a patch published after a higher minor still won't grablatest). Regression guard:tests/unit/build/should-promote-latest-5301.test.ts(#5301)max_tokens/max_completion_tokens(e.g. Zoo Code's-1"let the server choose") as "no limit" — omit the field instead of forcing it to1.clampMaxTokenspreviously didMath.max(1, …), so a client-1was sent upstream asmax_tokens: 1, truncating the response to a single token (the observedcompletion_tokens: 1,content: null,reasoning_content: "The"withfinish_reason: stop). Now any value≤ 0is dropped so Command Code applies the model's native default; positive values are still floored and clamped to the 200k ceiling. Regression guard:tests/unit/command-code-maxtokens-negative-5166.test.ts(#5166 — thanks @Stazyu)[network refresh + DB write]atomic for one connection, but it does not protect against a third writer (a sibling request, a concurrent HealthCheck, or a replica) landing a fresherrefresh_tokenrotation on the sameconnection_idbetween the staleness read and the persist. Overwriting that fresher row reverts the sibling's rotation; the next caller then loads the now-consumed token, Auth0/Anthropic flag it asrefresh_token_reused, and the whole token family gets revoked (the 1352× claude/aa5dd5cfinvalidation storm).getAccessTokennow re-reads the row's currentrefresh_tokenimmediately before persisting (inside the mutex) and skips the write when it has rotated past the token the caller presented — the caller still receives the freshly-issued access token, only the DB overwrite is skipped. Opt-in viarunWithCasGuard(no active guard ⇒ byte-identical behavior); skip/persist counters exposed viagetCasGuardStats(). Regression guard:tests/unit/token-refresh-cas-guard-4038.test.ts. (#4038 — thanks @KooshaPari for the root-cause diagnosis)schemas/tools.ts ↔ schemas/toolSearch.tsimport cycle introduced when thetool_searchdefs (feat(mcp): omniroute_tool_search + one-line TS signatures — roadmap #4 #5269) were extracted into their own module —toolSearch.tsimportedMcpToolDefinitionfromtools.tswhiletools.tsimportedtoolSearchToolfromtoolSearch.ts, failingcheck:cyclesonrelease/v3.8.40. The sharedAuditLevel+McpToolDefinitiontypes now live in a leafschemas/toolDefinition.tsthat both import;tools.tsre-exports them for backward compatibility.compression_analyticsrow was written only on a net-positive saving, so a Stacked (RTK→Caveman) pipeline that ran on already-compact context produced no row — indistinguishable from "never dispatched" (byMode.stacked.countstayed flat while Ultra climbed). Such runs are now recorded withskip_reasonand surfaced as a per-modeskippedcount plustotalSkipped/bySkipReasonin the analytics summary and the Mode Breakdown; the existing net-saving totals/averages are unchanged (skip rows are excluded from them) ([BUG] Stacked RTK + Caveman compression is unclear/unreliable; Ultra works but Stacked often records no savings #4268 — thanks @abdulkadirozyurt, @androw)omniroute server --trayshowing no tray on macOS/Linux with no error printed. The wired Unix tray path loadedsystray2through an inline loader that calledrequire("module")inside an ESM.mjsfile ("type":"module") →ReferenceError: require is not defined, silently swallowed (regressed in v3.8.34); even if it had loaded,systray2isn't innode_modules(it's lazily installed into~/.omniroute/runtime). The loader now delegates to the runtime loader, the icon path (icon.png) is corrected,isTemplateIconisfalse(the full-color icon rendered as a white square under macOS template mode), and tray start failures are surfaced to stderr instead of being swallowed ([BUG] Omniroute #4605 — thanks @ProgMEM-CC).requestenvelope when converting Antigravity IDE requests. The real IDE sendscloudcode-pa.googleapis.com/v1internal:generateContentwith the Gemini request nested under.request({ project, model, request: { contents, systemInstruction, generationConfig } }), but the bridge read those fields at the top level — yielding an empty conversation, so prompts hung mid-execution. The legacy/v1beta/models/<model>:generateContenttop-level shape still works ([BUG] Antigravity-IDE login does not work with agent-bridge #4294 — thanks @shabeer)registry.npmjs.org.resolveLatestVersion()now tries npm CLI → npm registry → GitHub releases (/repos/diegosouzapw/OmniRoute/releases/latest) before giving up, and logs a warning only when all three fail ([bug] Home page "Update Available" banner no longer appears when a newer version exists #4100)max_tokenswhen the client omits it so the upstream applies the model's native default, fixing400 "expected <=200000"on/alpha/generatefor high-cap models; an explicit oversized client value is clamped to the 200k endpoint ceiling (fix(command-code): omit max_tokens when client omits it; correct registry caps #5221 — thanks @adivekar-utexas)504s and throughput collapse under concurrency. The weighted/priority paths already honored per-conversation stickiness; the round-robin handler returned before reaching it. Round-robin now starts the rotation at the conversation's sticky connection (failover to the other targets is preserved), and different conversations still spread across connections — only intra-conversation rotation is removed (#5248, [BUG] Frequent 504 Errors and Significant TPS Degradation After Upgrading Beyond v3.8.14 #3825 — thanks @bypanghu, @jpsn123, @xz-dev)"Continue"turn with a neutral filler ("...") — when an OpenAI→Kiro request ends on an assistant/tool turn, the translator synthesizes the protocol-required trailing user turn, and the literal word"Continue"could be read by Kiro/CodeWhisperer as a real user instruction and trigger unintended agent action. A trailing tool-result turn is still promoted as-is (it already collapses to a real user turn); only the assistant-text-ending case is affected. Regression guards:tests/unit/kiro-continue-filler-5231.test.ts. (#5231)400 "requested model is not supported"instead of hard-failing. The 400 guard in the priority strategy treatedMODEL_CAPACITYas a block-fallback reason, so a combo that hit a provider lacking a specific model returned a hard400even when other targets (different providers) supported it. Such 400s now fall through to the next target. (#5249 — thanks @Chewji9875)blockedProviders) silently removed its card because the page dropped blocked no-auth entries from its render list — the only way back was buried under Settings → Security → Blocked Providers. The page now partitions no-auth entries: visible providers render as before, blocked ones appear in a "Disabled" sub-group with an Enable button that un-blocks them in place. Aggregates, counts and/v1/modelsstill consume the visible-only list (blocked providers stay out of routing). Regression guard:tests/unit/noauth-blocked-partition-5183.test.ts. (#5183, follow-up from #5166 — thanks @WslzGmzs)/dashboard/contextpage so RSC prefetches of the compression-context hub no longer 404. The route only had sub-routes (settings,combos,ultra, …) and no parent page, so the App Router returned 404 for the bare segment. The parent now redirects to its canonical sub-route (/dashboard/context/settings), honoring a legacy?tab=query for deep links. Regression guard:tests/unit/dashboard/context-parent-redirect-5298.test.ts(#5298 — thanks @KooshaPari)sidebar.gamificationGroupmessage across all 42 locales — the Gamification sidebar group referenced atitleKeythat existed in no locale, loggingMISSING_MESSAGE: sidebar.gamificationGroup (en)at runtime (the group still rendered via itstitleFallback). The key is now present everywhere so the warning is gone and locale coverage is unaffected (#5298 — thanks @KooshaPari)/v1/chat/completionsno longer returns SSE for a non-stream OpenAI-compatible request whenstreamis omitted and the client sendsAccept: application/json, text/event-stream— the Vercel AI SDK / OpenAI SDK non-stream signature (doGenerate()/generateText()), which then failed withInvalid JSON response(Unexpected token 'd', "data: {"id"...). The route-level Accept override (fix(docker): use /api/monitoring/health for Docker healthcheck (#296) #302) andresolveStreamFlagnow treat an Accept header that explicitly listsapplication/jsonas a JSON opt-in even when it also liststext/event-stream; only a pureAccept: text/event-stream(noapplication/json) still opts an omitted-streamrequest into SSE, and an explicit bodystreamvalue always wins. The shared decision now lives inacceptHeaderForcesStream. Regression guard:tests/unit/sse-nonstream-accept-5305.test.ts. (#5305 — thanks @md-riaz)local_shellhosted tool type before forwarding to OpenAI's Responses API, resolving the omni-combo400 "The local_shell tool is no longer supported."spike. Inbound Responseslocal_shellis still accepted and mapped to a caller-side Chatshellfunction for compatibility. (#5250, #5256 — thanks @KooshaPari)excludeConnectionIdas a fallback scenario — once exclusions accumulated throughexcludedConnectionIds, selection could fall back to normal sticky/priority behavior instead of LRU-selecting the next eligible account for the same model/family. Any non-empty accumulated exclude set is now treated as fallback mode. Builds on the family-scoped lockout work in fix(antigravity): retry accounts by quota family #5180 (v3.8.39). (#5222 — thanks @Ardem2025)presencePenalty,frequencyPenalty,logprobs,topLogprobs) before sending to the Grok Build API, fixing400 'Model does not support parameter presencePenalty'when clients (MiMoCode, Cursor, etc.) send OpenAI-style params. (#5273 — thanks @fulorgnas)~/.grok/auth.jsonobject in the dashboard import-token endpoint. TheoauthImportTokenSchemaonly accepted a bare string token while the UI sends the whole auth.json object →400 Bad Request; the schema now accepts the object and stores the original underproviderSpecificData.rawAuthJsonfor diagnostics and token refresh. (#5258 — thanks @fulorgnas)openapi.qoder.sh/api/v1/jobToken/exchangebefore the first exchange populates the completed-token cache; the shared exchange is also decoupled from any single caller'sAbortSignalso one aborted waiter can't cancel it for the others. (#5254, #5265 — thanks @KooshaPari)Response. (#5257 — thanks @rdself)<think>/<thinking>tag extraction so generic OpenAI-compatible paths don't rewrite prompt-format content intoreasoning_content; an explicit opt-in keeps tag-native families (DeepSeek-R1, QwQ) working while Antigravity/Agy stay excluded by provider/model prefix. (#5224 — thanks @rdself)schemas/tools.ts ↔ schemas/toolSearch.tsimport cycle introduced when thetool_searchdefs (feat(mcp): omniroute_tool_search + one-line TS signatures — roadmap #4 #5269) were extracted into their own module —toolSearch.tsimportedMcpToolDefinitionfromtools.tswhiletools.tsimportedtoolSearchToolfromtoolSearch.ts, failingcheck:cyclesonrelease/v3.8.40. The sharedAuditLevel+McpToolDefinitiontypes now live in a leafschemas/toolDefinition.tsthat both import;tools.tsre-exports them for backward compatibility. (#5282)compression_analyticsrow was previously written only on a net-positive saving, so a Stacked pipeline that ran on already-compact context produced no row — indistinguishable from "never dispatched". Such runs are now recorded withskip_reasonand surfaced as a per-modeskippedcount plustotalSkipped/bySkipReason; net-saving totals/averages are unchanged (skip rows excluded). (#5277, [BUG] Stacked RTK + Caveman compression is unclear/unreliable; Ultra works but Stacked often records no savings #4268 — thanks @abdulkadirozyurt, @androw)omniroute server --trayshowing no tray on macOS/Linux with no error printed. The Unix tray path loadedsystray2through an inline loader that calledrequire("module")inside an ESM.mjsfile →ReferenceError: require is not defined, silently swallowed (regressed in v3.8.34); even if loaded,systray2is lazily installed into~/.omniroute/runtime, notnode_modules. The loader now delegates to the runtime loader, the icon path is corrected,isTemplateIconisfalse(the full-color icon rendered as a white square under macOS template mode), and tray start failures surface to stderr. (#5276, [BUG] Omniroute #4605 — thanks @ProgMEM-CC).requestenvelope when converting Antigravity IDE requests. The real IDE sendscloudcode-pa.googleapis.com/v1internal:generateContentwith the Gemini request nested under.request, but the bridge read those fields at the top level — yielding an empty conversation, so prompts hung mid-execution. The legacy/v1beta/models/<model>:generateContenttop-level shape still works. (#5267, [BUG] Antigravity-IDE login does not work with agent-bridge #4294 — thanks @shabeer)registry.npmjs.org.resolveLatestVersion()now tries npm CLI → npm registry → GitHub releases before giving up. (#5266, [bug] Home page "Update Available" banner no longer appears when a newer version exists #4100)max_tokenswhen the client omits it so the upstream applies the model's native default, fixing400 "expected <=200000"on/alpha/generatefor high-cap models; an explicit oversized client value is clamped to the 200k endpoint ceiling. (#5221 — thanks @adivekar-utexas)🔒 Security
/v1beta/*Gemini-compatible client API.next.config.mjsrewrote/v1beta/:path*→/api/v1beta/:path*, butsrc/proxy.tsdidn't match/v1betabefore the rewrite andclassifyRoute()didn't classify/api/v1beta/*as client API — so unauthenticated/v1beta/models/...:generateContenttraffic could reach the model-serving route without the central client-API auth policy. Both alias and rewritten forms are now classifiedCLIENT_API, enforcing Bearer auth whenREQUIRE_API_KEYis enabled. (#5274 — thanks @rdself)src/server/origin/publicOrigin.tsand wire it through the authz pipeline, replacing the per-route same-origin-only check that403'd dashboard mutations when served behind a reverse proxy on a different public origin. The module resolves the allowed public origin from configured base-URL env vars or trusted forwarded headers (only whenOMNIROUTE_TRUST_PROXYis set and the peer is loopback/LAN via peer-stamp), validatesSec-Fetch-Sitemetadata, and sanitizesHost/Forwardedinputs (rejects control chars, userinfo, path/query in Host). Regression guards:tests/unit/authz/public-origin.test.ts+tests/unit/authz/pipeline.test.ts. (#5278 — thanks @Thinkscape / @abodera)📝 Maintenance
docs/, run an accuracy audit, and drop Node 20 from the supported matrix (rebased onto the current release tip). (#5262)npm installwarnings, fix a runtimeCannot find modulecrash onomniroute serve(the runtime-env module was missing from the npmfilesallow-list), and clear the 4 moderatenpm auditadvisories. (#5252 — thanks @yunaamelia), (#5230, [BUG] Global npm installation fails: Missing scripts/build/runtime-env.mjs when running omniroute #5227)combo/targetExhaustion.tshandler (provider-exhausted / connection-error / transient rate-limited classification), registered in the vitest discovery list. (#5296 — thanks @KooshaPari)🙌 Contributors
Thanks to everyone who contributed code, fixes, reports and root-cause diagnoses this cycle:
External contributors & reporters
/v1beta/*auth hardening ([urgent]: fix v1beta authz bypass #5274), provider-request header logging (fix: preserve provider request headers in logs #5257),<think>tag scoping (Scope textual thinking tag extraction #5224), Gemini CLI removal (Remove discontinued Gemini CLI channel #5246)max_tokensomit (fix(command-code): omit max_tokens when client omits it; correct registry caps #5221)max_tokens([BUG] Zoo Code + Command Code via OmniRoute returns empty output or 400 validation error, while the same setup works via 9Router #5166).requestenvelope unwrap ([BUG] Antigravity-IDE login does not work with agent-bridge #4294)Maintainer