Repository navigation
fix(github): Implement dynamic fallback for Codex models requiring /responses endpoint - #127
Conversation
There was a problem hiding this comment.
Pull request overview
Implements a reactive fallback in the GitHub provider executor to handle newer GitHub Copilot Codex models that reject /chat/completions and require the /responses endpoint (Fixes #102).
Changes:
- Adds dynamic detection of the specific 400 error and retries the request against
/responses. - Caches models known to require
/responsesto avoid repeating the initial failing request. - Translates OpenAI Chat Completions request/response streaming to/from the OpenAI Responses API format.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| if (parsed.done) { | ||
| controller.enqueue(new TextEncoder().encode("data: [DONE]\n\n")); | ||
| continue; | ||
| } | ||
|
|
||
| const converted = openaiResponsesToOpenAIResponse(parsed, state); | ||
| if (converted) { | ||
| const sseString = formatSSE(converted, "openai"); | ||
| controller.enqueue(new TextEncoder().encode(sseString)); | ||
| } |
There was a problem hiding this comment.
TransformStream creates a new TextEncoder() for every parsed SSE line (new TextEncoder().encode(...) inside the loop). This adds avoidable overhead in high-token streams. Instantiate a single TextEncoder once (outside the loop / outside the transform callback) and reuse it for all enqueues.
| async executeWithResponsesEndpoint({ model, body, stream, credentials, signal, log }) { | ||
| const url = "https://api.githubcopilot.com/responses"; | ||
| const headers = this.buildHeaders(credentials, stream); | ||
|
|
||
| const transformedBody = openaiToOpenAIResponsesRequest(model, body, stream, credentials); | ||
|
|
||
| log?.debug("GITHUB", "Sending translated request to /responses"); | ||
|
|
||
| const response = await fetch(url, { | ||
| method: "POST", | ||
| headers, | ||
| body: JSON.stringify(transformedBody), | ||
| signal |
There was a problem hiding this comment.
executeWithResponsesEndpoint assumes the upstream /responses call returns an SSE stream and immediately pipes response.body through an SSE transformer, but headers/body are still built from the caller-provided stream flag. If stream is false, buildHeaders() sends Accept: application/json, which can cause /responses to return JSON and make the streaming parser/pipeThrough path fail. Either force streaming when using /responses (set Accept to text/event-stream and ensure the request body has stream: true), or add a non-streaming code path that parses JSON and converts it to the expected OpenAI chat completion response format.
| if (buffer.trim()) { | ||
| const parsed = parseSSELine(buffer.trim()); | ||
| if (parsed && !parsed.done) { | ||
| const converted = openaiResponsesToOpenAIResponse(parsed, state); | ||
| if (converted) { | ||
| controller.enqueue(new TextEncoder().encode(formatSSE(converted, "openai"))); | ||
| } | ||
| } | ||
| } |
There was a problem hiding this comment.
In the SSE transformer flush(), a trailing data: [DONE] line can be dropped if it arrives without a terminating newline: parseSSELine(buffer.trim()) returns {done:true}, but the code only forwards buffered data when !parsed.done. This can leave clients waiting indefinitely for [DONE]. Handle the parsed.done case in flush() (enqueue data: [DONE]\n\n), or flush any remaining buffered line through the same parsed.done branch used in transform().
| if (buffer.trim()) { | |
| const parsed = parseSSELine(buffer.trim()); | |
| if (parsed && !parsed.done) { | |
| const converted = openaiResponsesToOpenAIResponse(parsed, state); | |
| if (converted) { | |
| controller.enqueue(new TextEncoder().encode(formatSSE(converted, "openai"))); | |
| } | |
| } | |
| } | |
| const trimmed = buffer.trim(); | |
| if (!trimmed) { | |
| return; | |
| } | |
| const parsed = parseSSELine(trimmed); | |
| if (!parsed) { | |
| return; | |
| } | |
| if (parsed.done) { | |
| controller.enqueue(new TextEncoder().encode("data: [DONE]\n\n")); | |
| return; | |
| } | |
| const converted = openaiResponsesToOpenAIResponse(parsed, state); | |
| if (converted) { | |
| controller.enqueue(new TextEncoder().encode(formatSSE(converted, "openai"))); | |
| } |
| if (result.response.status === HTTP_STATUS.BAD_REQUEST) { | ||
| const errorBody = await result.response.clone().text(); | ||
|
|
||
| if (errorBody.includes("not accessible via the /chat/completions endpoint")) { | ||
| log?.warn("GITHUB", `Model ${model} requires /responses. Switching...`); | ||
| this.knownCodexModels.add(model); | ||
| return this.executeWithResponsesEndpoint(options); | ||
| } |
There was a problem hiding this comment.
The fallback trigger checks errorBody.includes("not accessible via the /chat/completions endpoint") on the raw response text. Since GitHub returns structured JSON errors, it would be more robust to JSON.parse and inspect error.message (and guard parse failures). This avoids false positives and makes the fallback resilient to formatting changes (whitespace, wrapping, localization).
|
Hi @decolua, First off, thanks for this awesome project! I've just submitted PR #127 which fixes the While working on this, I explored a more proactive architectural approach that I wanted to share for your consideration, even though I opted for the safer method in the PR itself. Alternative Idea: Proactive Service Discovery I confirmed that the For example, for {
"id": "gpt-5.2-codex",
...
"supported_endpoints": [
"/responses"
],
"vendor": "OpenAI"
}And for a standard model like {
"id": "claude-opus-4.6",
...
"supported_endpoints": [
"/v1/messages",
"/chat/completions"
],
"vendor": "Anthropic"
}We could implement a "Service Discovery" mechanism that fetches this data on startup (or periodically) and builds an in-memory map ( Pros of this approach:
Why I didn't implement it in the PR:
Suggestion for the Future:
Request for ReviewCould you please review the current implementation? I'd also love to hear your thoughts on the architectural direction:
I am happy to adjust the code based on your preference. Thanks for the great project! |
…esponses endpoint (decolua#127) * fix(github): add dynamic fallback to /responses for Codex models * Refactor GithubExecutor: use config for URL detection (cherry picked from commit 6913129)
…esponses endpoint (decolua#127) * fix(github): add dynamic fallback to /responses for Codex models * Refactor GithubExecutor: use config for URL detection
### Bug Fixes * remove duplicated opencode prefix from model IDs and handle it in executor ([180f884](180f884)) * update opencode endpoint and model IDs ([6be8a03](6be8a03)) ### Features * add opencode provider ([03c846e](03c846e)) ## [0.2.96](5645d0a...v0.2.96) (2026-02-24) ### Bug Fixes * GitHub Copilot model ([95fd950](95fd950)) * **auth:** allow HTTP for local network ([0a394d0](0a394d0)) * **auth:** prevent auto-login after logout ([49df3dc](49df3dc)) * **codex:** use user-agent detection for Droid CLI compatibility ([8c6e3b8](8c6e3b8)) * Correct indentation for clarity in chatCore and claude-to-openai response handlers ([fa06226](fa06226)) * correct token extraction for Claude non-streaming responses ([decolua#131](https://github.com/involvex/involvex-claude-router/issues/131)) ([9fbd6e6](9fbd6e6)) * **dashboard:** resolve 'Provider not found' for free providers ([45a4d3b](45a4d3b)) * **db:** improve error handling and null checks ([e6ef852](e6ef852)) * **gemini:** improve base64 image data parsing ([5645d0a](5645d0a)) * **github:** Implement dynamic fallback for Codex models requiring /responses endpoint ([decolua#127](https://github.com/involvex/involvex-claude-router/issues/127)) ([6913129](6913129)) * improve code formatting and reduce auto-refresh interval ([7f71916](7f71916)) * improve cursor auto-import reliability on macOS ([decolua#161](https://github.com/involvex/involvex-claude-router/issues/161)) ([d7e06c3](d7e06c3)) * **login:** avoid infinite loading on settings fetch failure ([01c9410](01c9410)) * **open-sse:** emit [DONE] in passthrough SSE mode ([decolua#142](https://github.com/involvex/involvex-claude-router/issues/142)) ([b9a6979](b9a6979)) * prevent race conditions in sticky round-robin ([3ad2f8d](3ad2f8d)) * resolve SonarQube findings and Next.js Image warnings ([7058b06](7058b06)) * update Codex executor for gpt-5.3-codex support ([d7d5dc9](d7d5dc9)) ### Features * add /v1/embeddings endpoint (OpenAI-compatible) ([decolua#146](https://github.com/involvex/involvex-claude-router/issues/146)) ([e1b8361](e1b8361)), closes [decolua#117](https://github.com/involvex/involvex-claude-router/issues/117) * Add Anthropic Compatible provider support ([da5bdef](da5bdef)) * add API endpoint dimension to usage statistics dashboard ([decolua#152](https://github.com/involvex/involvex-claude-router/issues/152)) ([806bd4a](806bd4a)) * add Claude Opus 4.6 to GitHub Copilot provider ([decolua#97](https://github.com/involvex/involvex-claude-router/issues/97)) ([3d60597](3d60597)) * Add Claude Sonnet 4.6 to GitHub Copilot ([decolua#149](https://github.com/involvex/involvex-claude-router/issues/149)) ([4e2a3f8](4e2a3f8)) * add CLI entry points for claude-router commands ([62c632d](62c632d)) * add enable/disable toggle for provider connections ([ed796d2](ed796d2)) * add Gemini 3.1 Pro models to provider ([f2025cc](f2025cc)) * add Gemini embeddings support + Letta compatibility fixes ([a57a8ce](a57a8ce)), closes [decolua#148](decolua#148) * add GLM 5 and MiniMax M2.5 models to providerModels.js; add Claude Sonnet 4.6 to CLI tools ([e1e5a81](e1e5a81)) * add GLM Coding (China) provider and Usage by API Keys statistics ([1ae4e31](1ae4e31)) * add GPT 4o to GitHub Copilot provider ([decolua#98](https://github.com/involvex/involvex-claude-router/issues/98)) ([c090bb0](c090bb0)) * add GPT 5.3 Codex Spark model to pricing and provider models ([decolua#133](https://github.com/involvex/involvex-claude-router/issues/133)) ([c7d4410](c7d4410)) * Add GPT 5.3 Codex to GitHub Copilot ([decolua#150](https://github.com/involvex/involvex-claude-router/issues/150)) ([c4aa424](c4aa424)) * add GPT-3.5 Turbo to GitHub Copilot provider ([e3dbd44](e3dbd44)) * add GPT-4 to GitHub Copilot provider ([6ade8ef](6ade8ef)) * add GPT-4o mini to GitHub Copilot provider ([053e490](053e490)) * add models management and router process control commands ([c65fd42](c65fd42)) * Add OpenAI-compatible provider nodes ([0a28f9f](0a28f9f)) * add password change functionality and dependencies ([23cfb19](23cfb19)) * add pause/resume functionality for API keys ([decolua#158](https://github.com/involvex/involvex-claude-router/issues/158)) ([73388a0](73388a0)) * add Qwen3.5 Coder Model configuration ([decolua#156](https://github.com/involvex/involvex-claude-router/issues/156)) ([f933dd9](f933dd9)) * add request logging functionality and usage metrics display ([e476907](e476907)) * add round-robin routing strategy ([9ebd7d3](9ebd7d3)) * add sticky round-robin routing strategy ([4f292aa](4f292aa)) * add support for GLM 5 (if) ([decolua#123](https://github.com/involvex/involvex-claude-router/issues/123)) ([03ab554](03ab554)) * add URL-based tab state persistence in usage page ([decolua#129](https://github.com/involvex/involvex-claude-router/issues/129)) ([6caef7f](6caef7f)) * allow custom user data directory via DATA_DIR environment variable ([d83bd86](d83bd86)) * **antigravity:** initial steps for Antigravity anti-ban alignment ([a229d79](a229d79)), closes [decolua#141](decolua#141) * **antigravity:** integrate Antigravity tool with MITM support and update CLI tools ([2e854bd](2e854bd)) * **auth:** add model-level rate limit locking for multi-bucket providers ([decolua#120](https://github.com/involvex/involvex-claude-router/issues/120)) ([202fee7](202fee7)), closes [decolua#110](https://github.com/involvex/involvex-claude-router/issues/110) * **auth:** Enhance authentication flow and settings management ([249fc28](249fc28)) * **cli-tools:** update CLI tools and add new models ([a2122e3](a2122e3)) * **cli-tools:** update default local endpoint port to 20128 ([6c41573](6c41573)) * **cli:** add new CLI package with basic scaffolding ([9cf4628](9cf4628)) * **cloud:** harden sync/auth flow, SSE fallback, and update changelog ([3d43983](3d43983)) * **codex:** add GPT 5.3, fix API translation, add thinking levels ([127475d](127475d)) * **codex:** Cursor compatibility + Next.js 16 proxy migration ([1c6dd6d](1c6dd6d)) * **codex:** Cursor compatibility + Next.js 16 proxy migration ([7b864a9](7b864a9)) * **codex:** Cursor compatibility + Next.js 16 proxy migration ([e9b0a73](e9b0a73)) * **config:** add Cloudflare MCP server and account ID configuration ([67297ea](67297ea)) * **cursor:** Add cursor Provider ([0a026c7](0a026c7)) * **cursor:** Integrate Cursor IDE support with OAuth import token flow ([137f315](137f315)) * **docker:** add Docker setup, environment examples, and architecture docs ([5e4a15b](5e4a15b)) * enhance disconnect handling and request tracking in chatCore.js ([decolua#126](https://github.com/involvex/involvex-claude-router/issues/126)) ([3d29b86](3d29b86)) * enhance request handling and error management in chatCore and streamToJsonConverter ([e2db638](e2db638)) * enhance usage stats with sortable columns and improved data handling ([bf6e09b](bf6e09b)) * Enhance usage tracking across response handlers ([a33924b](a33924b)) * **executors:** Improved UI components for displaying provider limits and usage statistics in the dashboard. ([32aefe5](32aefe5)) * **iflow:** add IFlowExecutor with HMAC-SHA256 signature and enable models ([bd23ab4](bd23ab4)) * **iflow:** add kimi-k2.5 model support ([9e357a7](9e357a7)) * implement API key requirement toggle ([4cf25dc](4cf25dc)) * Implement buffer addition to usage tracking for improved context handling ([7881db8](7881db8)) * implement lazy loading for UsagePage with suspense fallback ([decolua#136](https://github.com/involvex/involvex-claude-router/issues/136)) ([05b09e6](05b09e6)) * implement provider connection reordering on create, update, and delete ([f2abcc6](f2abcc6)) * implement real project ID fetching for Antigravity ([decolua#170](https://github.com/involvex/involvex-claude-router/issues/170)) ([ea67742](ea67742)) * implement request tracking and enhance usage stats display ([e4f92cd](e4f92cd)) * implement usage tracking for AI requests ([9c3d6f4](9c3d6f4)) * Improve Antigravity quota monitoring and fix Droid CLI compatibility ([3c65e0c](3c65e0c)) * **open-sse:** add Claude Sonnet 4.6 ([b057c43](b057c43)) * OpenAI compatibility improvements & build fixes ([d9b8e48](d9b8e48)), closes [decolua#18](https://github.com/involvex/involvex-claude-router/issues/18) * **provider:** add free providers and enhance error handling ([bdbe816](bdbe816)) * **providers:** add Minimax Coding (China) provider ([7c609d7](7c609d7)) * **providers:** add provider icons to dashboard ([60bd686](60bd686)) * **providers:** auto-validate API keys on save ([b275dfd](b275dfd)) * rename app to involvex-claude-router and add lenient JSON parsing ([bec206d](bec206d)) * **responses:** respect client streaming preference + string input support ([decolua#121](https://github.com/involvex/involvex-claude-router/issues/121)) ([ac7cedd](ac7cedd)) * **translator:** add thinking parameter support in OpenAI → Claude ([54e01d6](54e01d6)) * **ui:** add cost tracking to usage dashboard and pricing settings ([f302c88](f302c88)) * **ui:** add model support for custom providers and improve UX ([a7a52be](a7a52be)) * Update response handling and logging for improved usage tracking ([df0e1d6](df0e1d6)) * **usage:** implement cost tracking backend and pricing configuration ([a36afaa](a36afaa))
…esponses endpoint (decolua#127) * fix(github): add dynamic fallback to /responses for Codex models * Refactor GithubExecutor: use config for URL detection
Summary
This PR resolves the
400 Bad Requesterror encountered when using newer GitHub Copilot Codex models (e.g.,gpt-5.1-codex,gpt-5.2-codex). The issue is addressed by implementing a dynamic, reactive fallback mechanism within theGithubExecutorthat automatically switches to the correct API endpoint.Root Cause Analysis
Certain GitHub Copilot models have deprecated the standard
/chat/completionsendpoint and now exclusively require requests to be sent to the/responsesendpoint, which also uses a different payload structure (inputarray instead ofmessages). The previous implementation was hardcoded to/chat/completions, causing API rejections for these models.Implementation Details
A resilient, self-correcting strategy has been implemented:
/chat/completions.400 Bad Requestis received with the specific error message"not accessible via the /chat/completions endpoint", the fallback logic is triggered.Set(this.knownCodexModels). All subsequent requests for this model will bypass the initial failed attempt and go directly to the correct/responsesendpoint, eliminating future latency./responsesendpoint. This involves two real-time translations:messagesarray is converted into theinputarray format./responsesis converted back into the standard OpenAI Chat Completion chunk format, ensuring client compatibility and correct tokenusagetracking.Considered Alternative: Proactive Service Discovery
An alternative, proactive "Service Discovery" approach was also investigated. This would involve making a
GET https://api.githubcopilot.com/modelsrequest at startup or periodically to build a map of which models support which endpoints, based on thesupported_endpointsfield present in the response.Pros of Service Discovery:
Cons / Reasons for Choosing the Fallback Approach:
/modelsendpoint and its schema are part of an internal, undocumented API. Relying on its structure is fragile, as it could change without warning and break the integration for all GitHub models./modelslist.Suggestion for the Future:
While the current reactive approach is safer, a hybrid model could be considered. We could implement Service Discovery as a "best-effort" optimization and still keep the reactive fallback as a safety net. This would provide the performance of Service Discovery with the resilience of the current implementation. I am happy to discuss or implement this if you think it's a worthwhile addition.
For now, this PR prioritizes robustness and future-proofing.
Related Issue
Fixes #102
