chore(release): v3.5.2 — Qoder DashScope Native Integration & Stability - #999
Conversation
…I child instances, plus workflow bumps
…mpty content guard & context-optimized combo strategy - Add isContextOverflowError + isContextOverflow detectors (400 + token-limit signals) - Auto-fallback to next family model on context overflow in chatCore - Add isEmptyContentResponse to catch fake-success empty responses, trigger fallback + recursive retry - Add OAUTH_INVALID_TOKEN error type (T11) with isOAuthInvalidToken signal matching; warn instead of deactivating node - Add getModelContextLimit helper in modelsDevSync (reads limit_context from synced capabilities) - Upgrade getTokenLimit in contextManager to check models.dev DB before registry (fixes gemini-2.5-pro: 1000000→1048576) - Add findLargerContextModel in modelFamilyFallback for context-aware model selection - Add sortModelsByContextSize + context-optimized combo strategy in combo.ts - Update context-manager unit test for corrected gemini-2.5-pro limit Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ecursion, dedup signals, findLargerContextModel - Fix isEmptyContentResponse: check message.tool_calls/delta.tool_calls instead of firstChoice.tool_calls (wrong OpenAI API path, caused tool-call responses to be falsely flagged as empty) - Fix empty content fallback: replace recursive handleChatCore call (infinite recursion risk + wrong model due to original body.model) with non-recursive pattern — call executeProviderRequest, parse fallback response body, reassign responseBody and fall through to existing processing - Fix context overflow: use findLargerContextModel over family candidates first, fall back to getNextFamilyFallback — ensures we pick a model with actually larger context window on overflow - Fix signal dedup: export CONTEXT_OVERFLOW_SIGNALS + CONTEXT_OVERFLOW_REGEX from errorClassifier.ts; import shared regex in modelFamilyFallback.ts, removing duplicate signal list and per-call RegExp construction Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When translating Claude-format responses (e.g. GLM) to Responses API
format for Codex CLI, the sanitizer stripped {event, data} structured
items to {"object":"chat.completion.chunk"}, losing all content and
the critical response.completed event.
Only run sanitizeStreamingChunk on OpenAI Chat Completions chunks,
skipping items that have the Responses API {event, data} structure.
Verifies that {event,data} structured items from the Responses API
translator bypass sanitizeStreamingChunk when translating Claude-format
providers (e.g. GLM) to Responses API format for Codex CLI.
…fix check Use explicit `response.` prefix check instead of generic `event && data` presence check, as recommended in PR review.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: df24a48e26
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const credentialsApiBase = | ||
| (credentials as any).customApiBase || (credentials as any).resourceUrl; |
There was a problem hiding this comment.
Read Qoder custom base URL from providerSpecificData
This endpoint override lookup never checks credentials.providerSpecificData, so resourceUrl/customApiBase values stored with provider metadata are ignored and requests always fall back to the hardcoded DashScope URL. In environments where Qoder OAuth returns a tenant-specific resourceUrl, this causes every request to hit the wrong host and fail despite valid credentials.
Useful? React with 👍 / 👎.
| const fallbackRaw = await fallbackResult.response.text(); | ||
| try { | ||
| responseBody = fallbackRaw ? JSON.parse(fallbackRaw) : {}; |
There was a problem hiding this comment.
Parse fallback body with SSE-aware logic
The empty-content fallback path assumes the fallback response is JSON and immediately calls JSON.parse, but earlier in this same handler we explicitly support providers that return SSE even when stream=false. For non-stream requests that still use upstream streaming (for example Claude Code compatible providers), a successful fallback response will be valid SSE and this parse will throw, incorrectly returning 502 instead of using the fallback result.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Code Review
This pull request updates OmniRoute to version 3.5.2, refactoring the Qoder executor for native DashScope integration and introducing a resilience engine overhaul with context overflow fallbacks and empty-content prevention. It also adds a context-optimized routing strategy, fixes Responses API stream corruption, and pins Next.js to version 16.0.10. Additionally, it updates deployment workflows and build scripts to improve package hygiene. Feedback identifies a bug in the empty-content fallback logic where the providerResponse variable is not updated, which could break streaming by failing to pass the new response stream to the client.
| const fallbackResult = await executeProviderRequest(nextModel, false); | ||
| if (fallbackResult.response.ok) { | ||
| const fallbackRaw = await fallbackResult.response.text(); | ||
| try { | ||
| responseBody = fallbackRaw ? JSON.parse(fallbackRaw) : {}; | ||
| providerUrl = fallbackResult.url; | ||
| providerHeaders = fallbackResult.headers; | ||
| finalBody = fallbackResult.transformedBody; | ||
| reqLogger.logTargetRequest(providerUrl, providerHeaders, finalBody); | ||
| log?.info?.( | ||
| "EMPTY_CONTENT_FALLBACK", | ||
| `Serving ${nextModel} as fallback for ${model}` | ||
| ); | ||
| // Fall through — continue processing with the new responseBody |
There was a problem hiding this comment.
The fallback logic for empty content responses (lines 1872-1885) is inconsistent with the context overflow fallback logic (lines 1661-1672) and contains a bug.
Specifically, it fails to update the providerResponse variable. While it updates responseBody, providerUrl, and providerHeaders, the rest of the handler (especially for streaming requests) will continue to use the original providerResponse. This means the client will receive the original empty response instead of the fallback result.
Additionally, calling .text() on the fallback response (line 1874) buffers the entire content. If the request is streaming, this defeats the purpose of streaming and introduces latency. You should update providerResponse and avoid consuming the body as text if streaming is enabled, allowing the subsequent logic to handle the new response stream correctly.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 837e64de1d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (typeof credentials.refreshToken === "string" && credentials.refreshToken.trim()) { | ||
| return credentials.refreshToken.trim(); |
There was a problem hiding this comment.
Stop using refresh token as request bearer token
getAuthToken falls back to credentials.refreshToken, and execute always forwards that value in Authorization. In OAuth flows a refresh token is not a resource-server credential, so when accessToken is missing this path both guarantees upstream auth failures and exposes a long-lived refresh secret to the provider endpoint/logging chain. This branch should reject and trigger refresh logic instead of sending the refresh token upstream.
Useful? React with 👍 / 👎.
| const currentProvider = currentParsed.provider || currentParsed.providerAlias || "unknown"; | ||
| const currentModelId = currentParsed.model || currentModel; | ||
| const currentLimit = getModelContextLimit(currentProvider, currentModelId) ?? 0; |
There was a problem hiding this comment.
Use real provider id for context-limit fallback lookup
findLargerContextModel derives provider via parseModel, but in this flow models are plain IDs (e.g. "gpt-5", "claude-sonnet-4-6") so parseModel returns no provider and this code falls back to "unknown". That makes getModelContextLimit miss synced capabilities, so the function usually returns null and context-overflow fallback never picks a truly larger-context sibling model.
Useful? React with 👍 / 👎.
…ty (diegosouzapw#999) * feat(qoder): native cosy integration * feat(qoder): implement native COSY encryption algorithm and remove CLI child instances, plus workflow bumps * feat(resilience): context overflow fallback, OAuth token detection, empty content guard & context-optimized combo strategy - Add isContextOverflowError + isContextOverflow detectors (400 + token-limit signals) - Auto-fallback to next family model on context overflow in chatCore - Add isEmptyContentResponse to catch fake-success empty responses, trigger fallback + recursive retry - Add OAUTH_INVALID_TOKEN error type (T11) with isOAuthInvalidToken signal matching; warn instead of deactivating node - Add getModelContextLimit helper in modelsDevSync (reads limit_context from synced capabilities) - Upgrade getTokenLimit in contextManager to check models.dev DB before registry (fixes gemini-2.5-pro: 1000000→1048576) - Add findLargerContextModel in modelFamilyFallback for context-aware model selection - Add sortModelsByContextSize + context-optimized combo strategy in combo.ts - Update context-manager unit test for corrected gemini-2.5-pro limit Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(review): address Gemini code review — tool_calls path, infinite recursion, dedup signals, findLargerContextModel - Fix isEmptyContentResponse: check message.tool_calls/delta.tool_calls instead of firstChoice.tool_calls (wrong OpenAI API path, caused tool-call responses to be falsely flagged as empty) - Fix empty content fallback: replace recursive handleChatCore call (infinite recursion risk + wrong model due to original body.model) with non-recursive pattern — call executeProviderRequest, parse fallback response body, reassign responseBody and fall through to existing processing - Fix context overflow: use findLargerContextModel over family candidates first, fall back to getNextFamilyFallback — ensures we pick a model with actually larger context window on overflow - Fix signal dedup: export CONTEXT_OVERFLOW_SIGNALS + CONTEXT_OVERFLOW_REGEX from errorClassifier.ts; import shared regex in modelFamilyFallback.ts, removing duplicate signal list and per-call RegExp construction Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(UI): add context-optimized strategy to frontend schema and options * fix(sse): preserve Responses API events in stream translation When translating Claude-format responses (e.g. GLM) to Responses API format for Codex CLI, the sanitizer stripped {event, data} structured items to {"object":"chat.completion.chunk"}, losing all content and the critical response.completed event. Only run sanitizeStreamingChunk on OpenAI Chat Completions chunks, skipping items that have the Responses API {event, data} structure. * test(sse): add regression test for Claude→Responses stream sanitization Verifies that {event,data} structured items from the Responses API translator bypass sanitizeStreamingChunk when translating Claude-format providers (e.g. GLM) to Responses API format for Codex CLI. * fix(sse): strengthen Responses API event detection with response. prefix check Use explicit `response.` prefix check instead of generic `event && data` presence check, as recommended in PR review. * fix: pin Next.js to 16.0.10 to prevent Turbopack hashed module bug Remove ^ prefix from next and eslint-config-next to prevent automatic upgrades to 16.1.x+ which introduced content-based hashing for external module references in Turbopack. Also remove duplicate Material Symbols @import from globals.css (font already loaded via <link> in layout.tsx). Fixes diegosouzapw#509 * align cc-compatible cache handling with client passthrough * chore: integrate resilience and turbopack fixes (PRs diegosouzapw#992, diegosouzapw#990, diegosouzapw#987) * chore(release): bump to v3.5.2 — changelog, docs, version sync * docs(i18n): sync documentation updates to 33 languages * fix(qoder): replace any with unknown to comply with strict any-budget --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Co-authored-by: oyi77 <oyi77@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Chris Staley <christopher-s@users.noreply.github.com> Co-authored-by: Ivan <shanin-i2011@yandex.ru> Co-authored-by: R.D. <rogerproself@gmail.com>
…ty (diegosouzapw#999) * feat(qoder): native cosy integration * feat(qoder): implement native COSY encryption algorithm and remove CLI child instances, plus workflow bumps * feat(resilience): context overflow fallback, OAuth token detection, empty content guard & context-optimized combo strategy - Add isContextOverflowError + isContextOverflow detectors (400 + token-limit signals) - Auto-fallback to next family model on context overflow in chatCore - Add isEmptyContentResponse to catch fake-success empty responses, trigger fallback + recursive retry - Add OAUTH_INVALID_TOKEN error type (T11) with isOAuthInvalidToken signal matching; warn instead of deactivating node - Add getModelContextLimit helper in modelsDevSync (reads limit_context from synced capabilities) - Upgrade getTokenLimit in contextManager to check models.dev DB before registry (fixes gemini-2.5-pro: 1000000→1048576) - Add findLargerContextModel in modelFamilyFallback for context-aware model selection - Add sortModelsByContextSize + context-optimized combo strategy in combo.ts - Update context-manager unit test for corrected gemini-2.5-pro limit Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(review): address Gemini code review — tool_calls path, infinite recursion, dedup signals, findLargerContextModel - Fix isEmptyContentResponse: check message.tool_calls/delta.tool_calls instead of firstChoice.tool_calls (wrong OpenAI API path, caused tool-call responses to be falsely flagged as empty) - Fix empty content fallback: replace recursive handleChatCore call (infinite recursion risk + wrong model due to original body.model) with non-recursive pattern — call executeProviderRequest, parse fallback response body, reassign responseBody and fall through to existing processing - Fix context overflow: use findLargerContextModel over family candidates first, fall back to getNextFamilyFallback — ensures we pick a model with actually larger context window on overflow - Fix signal dedup: export CONTEXT_OVERFLOW_SIGNALS + CONTEXT_OVERFLOW_REGEX from errorClassifier.ts; import shared regex in modelFamilyFallback.ts, removing duplicate signal list and per-call RegExp construction Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(UI): add context-optimized strategy to frontend schema and options * fix(sse): preserve Responses API events in stream translation When translating Claude-format responses (e.g. GLM) to Responses API format for Codex CLI, the sanitizer stripped {event, data} structured items to {"object":"chat.completion.chunk"}, losing all content and the critical response.completed event. Only run sanitizeStreamingChunk on OpenAI Chat Completions chunks, skipping items that have the Responses API {event, data} structure. * test(sse): add regression test for Claude→Responses stream sanitization Verifies that {event,data} structured items from the Responses API translator bypass sanitizeStreamingChunk when translating Claude-format providers (e.g. GLM) to Responses API format for Codex CLI. * fix(sse): strengthen Responses API event detection with response. prefix check Use explicit `response.` prefix check instead of generic `event && data` presence check, as recommended in PR review. * fix: pin Next.js to 16.0.10 to prevent Turbopack hashed module bug Remove ^ prefix from next and eslint-config-next to prevent automatic upgrades to 16.1.x+ which introduced content-based hashing for external module references in Turbopack. Also remove duplicate Material Symbols @import from globals.css (font already loaded via <link> in layout.tsx). Fixes diegosouzapw#509 * align cc-compatible cache handling with client passthrough * chore: integrate resilience and turbopack fixes (PRs diegosouzapw#992, diegosouzapw#990, diegosouzapw#987) * chore(release): bump to v3.5.2 — changelog, docs, version sync * docs(i18n): sync documentation updates to 33 languages * fix(qoder): replace any with unknown to comply with strict any-budget --------- Co-authored-by: diegosouzapw <diegosouzapw@users.noreply.github.com> Co-authored-by: oyi77 <oyi77@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Chris Staley <christopher-s@users.noreply.github.com> Co-authored-by: Ivan <shanin-i2011@yandex.ru> Co-authored-by: R.D. <rogerproself@gmail.com>
🚀 Release v3.5.2
Changes
cryptomodules while improving stream fidelity.response.specific SSE prefixes from streaming boundaries (fix(sse): preserve Responses API events in stream translation #992).16.0.10preventing memory leaks and build staleness from recent upstream Turbopack hashed module regressions (Fix: Pin Next.js to 16.0.10 to prevent Turbopack hashed module bug #987).Tests