Skip to content

Release v3.8.16 - #3385

Merged
diegosouzapw merged 34 commits into
mainfrom
release/v3.8.16
Jun 8, 2026
Merged

diegosouzapw merged 34 commits into
mainfrom
release/v3.8.16

Conversation

@diegosouzapw

@diegosouzapw diegosouzapw commented Jun 7, 2026 •

Copy link
Copy Markdown
Owner

[3.8.16] — 2026-06-08

✨ New Features

  • feat(vision-bridge): auto-routing to the fastest available vision model — when a request carries image content and the selected model does not support vision, OmniRoute now transparently delegates to the best-match vision-capable model instead of returning an error. (#3377 — thanks @herjarsa)
  • feat(web-session): web-session pool observability — new MCP tool get_web_session_pool_health and a health-matrix REST response (GET /api/web-session-pool/health) expose per-provider slot counts, lease ages, and error budgets so operators can diagnose pool exhaustion without digging through logs. (#3395 — thanks @oyi77)
  • feat(web-session): adaptive keepalive threshold — the keepalive heartbeat interval now self-adjusts based on observed provider idle-disconnect behaviour instead of using a fixed constant, reducing both unnecessary pings and unexpected session drops. (#3397 — thanks @oyi77)
  • feat(web-session): bulk credential import endpoint (POST /api/web-session/import) — import a JSON array of session credentials in one call; each entry is validated and inserted atomically, with per-entry success/failure reported in the response. (#3403 — thanks @oyi77)
  • feat(api): REST API for session pool health (GET /api/session-pool/health) — a dashboard-facing endpoint that aggregates live slot usage, wait-queue depth, and error rates across all active session pools; wired to a new dashboard widget. (#3404 — thanks @oyi77)

🔧 Bug Fixes

  • fix(sse): eliminate race window in usageTokenBuffer settings update — a concurrent save + stream-start could race to apply stale settings, causing token counts to roll back by up to 2 000 tokens after a restart; the update now uses an atomic read-modify-write on the shared settings ref. (#3405 — thanks @diegosouzapw)
  • fix(context-cache): server-side context-cache pinning now correctly persists across restarts; proxy message content no longer leaks into the upstream prompt; and the context_cache_protection toggle is properly saved to the DB on change. (#3399 — thanks @k0valik)
  • fix(providers): the provider settings page now refreshes its model list after a successful sync-models call — previously the stale list remained until a full page reload. (#3402 — thanks @0xtbug)
  • fix(stream): empty-choices chunks (choices array present but empty, no finish_reason) are now silently dropped rather than emitted as a retry: SSE event — removes spurious retry lines from streaming responses for providers that emit heartbeat keep-alive chunks. (#3400 — thanks @0xtbug)
  • fix(account-fallback): the connection cooldown deduplication state is now preserved across the fallback retry chain — previously a second concurrent failure on the same account could clear the dedupe flag set by the first, allowing the cooldown window to be extended twice. (#3381 — thanks @oyi77)
  • fix(stream): false-positive textual tool-call marker truncation — containsTextualToolCallMarker now tracks how much of the accumulated streamed content has already been emitted, so it only withholds the unemitted tail rather than re-scanning from the start on every new chunk. (#3382 — thanks @Ardem2025)
  • fix(sanitizer): containsTextualToolCallContent() now requires the complete [Tool call: name]\nArguments: header pattern instead of a bare .includes("[Tool call:") check — prevents the non-streaming response sanitizer from nulling out model responses that merely quote [Tool call:] in prose or code examples. (#3355 — thanks @diegosouzapw)
  • fix(stream): the streaming textual tool-call guard now flushes any remaining buffered content as plain text when the stream ends, regardless of whether the buffer contains "Arguments:" — previously, a partial/incomplete tool-call header that arrived at end-of-stream was silently dropped. (#3355 — thanks @diegosouzapw)
  • fix(executor): Mistral (and any provider in PROVIDERS_REQUIRING_USER_LAST_MESSAGE) no longer receives a trailing assistant message with plain text content — stripTrailingAssistantForProvider drops it on the upstream-send path, fixing the 400: Expected last role User or Tool … but got assistant rejection. (#3396 — thanks @diegosouzapw)
  • fix(mitm): getMitmStatus() in the build-time stub (Docker image) now returns a graceful { running: false } status instead of throwing, so the Agent Bridge UI shows a clean "stopped" state rather than an error banner in containerised deployments. (#3390 — thanks @diegosouzapw)
  • fix(env): corrected casing of OMNIROUTE_TRACE in .env.example and all related documentation files — was previously mixed-case in some places, causing the variable to be silently ignored on case-sensitive file systems. (#3393 — thanks @androw)
  • fix(featureFlags): PRICING_SYNC_ENABLED description now clearly states that the feature requires the corresponding environment variable to be set — removes the ambiguity that led operators to enable it via the UI only and wonder why sync never ran. (#3394 — thanks @androw)

📝 Maintenance

  • ci(docker): the CI pipeline now builds and publishes the -web image variant in the same Docker publish workflow, so both the standard and browser-backed images stay in sync on every release. (#3389 — thanks @zhiru)
  • ci(e2e): E2E shard suite hardened — timeout raised to 45 min for the heaviest shard; build artifact now uses an explicit tar bundle to avoid upload-artifact@v4 LCA path ambiguity; node_modules copied into standalone after download; browser cache added to cut cold-shard time; sync-models endpoint mocked in providers-management.spec.ts so the import modal reaches "done" immediately. (#3387 / #3392 — thanks @diegosouzapw)
  • docs: Codex CLI configuration guide added to the dashboard (/dashboard/codex-config) — covers profile naming, model selection, and the CODEX_* environment variables accepted by OmniRoute. (thanks @diegosouzapw)
  • chore(agentSkills): catalog expanded to 43 entries — config-codex-cli added as a new CONFIG_SKILL_IDS category; all skill-count assertions updated across unit and integration test suites; next-fetch opts cast to satisfy the TypeScript overload signature in the skill runner. (thanks @diegosouzapw)

🙌 Contributors

Thanks to everyone whose work landed in v3.8.16:

Contributor Contribution
@herjarsa Vision-bridge auto-routing to fastest vision model (#3377)
@oyi77 Web-session pool observability (#3395), adaptive keepalive (#3397), bulk credential import (#3403), session pool REST API (#3404), cooldown dedupe fix (#3381)
@Ardem2025 Stream false-positive tool-call marker truncation fix (#3382)
@zhiru Docker -web image variant CI (#3389)
@androw OMNIROUTE_TRACE casing fix (#3393), PRICING_SYNC_ENABLED clarification (#3394)
@k0valik Context-cache pinning + proxy message leak fix (#3399)
@0xtbug Empty-choices chunk drop (#3400), model list refresh after sync (#3402)
@diegosouzapw Release engineering + usageTokenBuffer race fix (#3405), sanitizer+stream hardening (#3355/#3410), Mistral trailing-assistant fix (#3396/#3409), mitm Docker stub (#3390/#3408), E2E shard stabilization (#3387/#3392), and direct release-branch commits

✅ Quality Gate

  • lint: ✅ pass
  • typecheck:core: ✅ pass
  • check:cycles: ✅ pass
  • check:docs-sync: ✅ pass
  • test:unit: ✅ pass
  • test:vitest: ✅ pass (146/146)
  • CI (GitHub Actions run 27152372873): ✅ 85/85 jobs green

📊 Coverage of commits since v3.8.15

  • Range: v3.8.15..HEAD
  • Commits: 32 (excl. cycle-open chore) — all accounted for in the changelog above
  • Contributor PRs: 15 external PRs from 7 contributors

⚠️ After merging: run Phase 2 (Local VPS homologation) before tagging.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request bumps the project version from 3.8.15 to 3.8.16 across multiple configuration files, including package.json, package-lock.json, and openapi.yaml. It also introduces an 'Unreleased' section for version 3.8.16 in the main changelog and all internationalized changelog files. As there are no review comments, I have no feedback to provide.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@github-actions

github-actions Bot commented Jun 7, 2026 •

Copy link
Copy Markdown
Contributor

CI Coverage Report

  • Coverage job: success
  • PR test policy: failure

Coverage artifact was not available for this run.

PR Test Policy

This PR changes production code in src/, open-sse/, electron/, or bin/ without accompanying automated tests.

diegosouzapw and others added 27 commits June 7, 2026 14:57
…droom) (#3387)

The heaviest E2E shard (5/6 — responsive viewport matrix + studio/smoke) overran
the job's 20m timeout-minutes because each shard re-runs `npm run build` (~5m)
before Playwright, then runs ~24 serial tests with retries:2. The job was killed
(CANCELLED mid-run, 'Terminate orphan process') instead of any test failing.

- Bump test-e2e timeout-minutes 20 -> 35 (cumulative build+tests headroom).
- Lower the Playwright per-test timeout 600s -> 180s so a genuine hang fails fast
  and visibly (a clear per-test timeout) instead of silently eating the job budget.
)

The 35m bump still wasn't enough — shard 5/6 (responsive viewport matrix +
studio/smoke, ~24 serial tests after a ~5m build) was still cancelled at 35m,
and the `github` Playwright reporter buffers output so the cancelled log showed
no per-test results (couldn't tell which test was slow).

- e2e timeout-minutes 35 -> 50 (the shard observably needs >35m; other shards
  finish in ~7m so they're unaffected).
- Playwright CI reporter github -> line so per-test progress + timing stream
  live to the job log, making any genuinely slow/hung test diagnosable.
…rify environment variable requirement (#3394)

Integrated into release/v3.8.16
… using emitted content state (#3382)

Integrated into release/v3.8.16
…sist context_cache_protection toggle (#3399)

Integrated into release/v3.8.16
Add a comprehensive guide for configuring Codex CLI to use OmniRoute as an OpenAI-compatible backend.

Document ready-to-use config examples, Responses API routing behavior, context window settings, token limits, model profiles, and troubleshooting guidance to help users avoid direct-provider compatibility issues.
…-effect lint

- docs/guides/CODEX-CLI-CONFIGURATION.md was missing the YAML frontmatter
  block required by fumadocs (title/version/lastUpdated), causing the
  production build to fail with "invalid frontmatter" MDX error.
- CodexCliGuideModal.tsx called setLoading/setError synchronously in a
  useEffect body, triggering the react-hooks/set-state-in-effect lint error.
  Refactored to an internal async function with an `cancelled` guard to
  prevent state updates on unmounted components.
- Remove 429 from PROVIDER_BREAKER_FAILURE_STATUSES; 429 belongs to
  connection cooldown, not whole-provider breaker (CLAUDE.md §resilience).
  PR #3366 correctly added 429 to PROVIDER_FAILURE_ERROR_CODES in
  accountFallback.ts (combo infinite-retry fix) but the parallel change
  to chat.ts was wrong — the integration test from v3.8.10 confirms this.

- Align stream-utils tests to PR #3399 (SYNTHETIC_CLAUDE_EMPTY_RESPONSE_TEXT
  → "", message.content → null) and PR #3355 (malformed tool-call buffer
  now emitted as plain text, not suppressed).

- Align services-branch-hardening test to PR #3399 (pinnedModel always
  null from applyComboAgentMiddleware; server-side session pinning replaced
  client-side <omniModel> tag extraction).

- Align combo-routing-engine context-cache tests to PR #3399 (no <omniModel>
  tag in output, no X-OmniRoute-Model header, priority routing unchanged).
Upload the Next.js build from the build job and reuse it across E2E
shards to avoid rebuilding in each shard. Increase Playwright sharding
from 6 to 9, cache Chromium browsers, and lower the E2E timeout to match
the faster expected runtime.

Add a Codex CLI configuration skill for OmniRoute setup and ignore local
credential-bearing setup prompts.
…s, cp after download) and update skill count to 43
- Unit/integration tests: update hardcoded 42→43 in 7 test files
  (agentSkillTools-mcp, agentSkills-catalog, agentSkills-generator,
  agent-skills-content, agent-skills-discovery, listCapabilities-a2a)
  to match the 43rd skill (config-codex-cli) added in the previous commit.
- Include CONFIG_SKILL_IDS in integration content test ALL_IDS so
  skills/config-codex-cli/ is no longer "unexpected".
- listCapabilities.ts: change totalSkills from literal 42 to catalog.length
  so it adapts to catalog growth automatically.
- computeCoverage assertions: include config.have in totalSkills check.
- CI: switch E2E artifact from upload-artifact path (ambiguous stripping)
  to explicit tar archive. Fixes "Could not find a production build in
  ./.build/next" — the previous approach's download path was double-nested
  (.build/next/next/...) due to upload-artifact LCA computation. tar -czf
  stores .build/next/... relative to CWD; tar -xzf restores them verbatim.
- Also exclude .build/next/cache from the tar to keep archive lean.
- feat(translator): strip client_metadata in Responses→Chat translation
  (Mistral 422 extra_forbidden fix); add regression test.
Update Codex CLI docs and configuration skill to use the v0.137+
profile file naming format: ~/.codex/<name>.config.toml instead of
the deprecated profile- prefix.

Clarify that missing profile files silently fall back to defaults, and
rename the setup workflow heading to match the config-codex-cli skill.
@diegosouzapw
diegosouzapw merged commit 60fc41f into main Jun 8, 2026
6 of 7 checks passed
@sonarqubecloud

sonarqubecloud Bot commented Jun 8, 2026

Copy link
Copy Markdown

@kilo-code-bot

kilo-code-bot Bot commented Jun 9, 2026

Copy link
Copy Markdown

Kilo Code Review could not run — your account is out of credits.

Add credits or switch to a free model to enable reviews on this change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants