Skip to content

fix(github): Implement dynamic fallback for Codex models requiring /responses endpoint - #127

Merged
decolua merged 2 commits into
decolua:masterfrom
I3eka:fix/github-codex-endpoints
Feb 15, 2026
Merged

decolua merged 2 commits into
decolua:masterfrom
I3eka:fix/github-codex-endpoints

Conversation

@I3eka

@I3eka I3eka commented Feb 14, 2026

Copy link
Copy Markdown
Contributor

Summary

This PR resolves the 400 Bad Request error encountered when using newer GitHub Copilot Codex models (e.g., gpt-5.1-codex, gpt-5.2-codex). The issue is addressed by implementing a dynamic, reactive fallback mechanism within the GithubExecutor that automatically switches to the correct API endpoint.

Root Cause Analysis

Certain GitHub Copilot models have deprecated the standard /chat/completions endpoint and now exclusively require requests to be sent to the /responses endpoint, which also uses a different payload structure (input array instead of messages). The previous implementation was hardcoded to /chat/completions, causing API rejections for these models.

Implementation Details

A resilient, self-correcting strategy has been implemented:

  1. Initial Attempt: The executor first makes a standard request to /chat/completions.
  2. Error Detection: If a 400 Bad Request is received with the specific error message "not accessible via the /chat/completions endpoint", the fallback logic is triggered.
  3. Caching for Performance: The model's ID is immediately added to an in-memory Set (this.knownCodexModels). All subsequent requests for this model will bypass the initial failed attempt and go directly to the correct /responses endpoint, eliminating future latency.
  4. On-the-Fly Translation: The request is automatically retried on the /responses endpoint. This involves two real-time translations:
    • Request Body: The OpenAI messages array is converted into the input array format.
    • Response Stream: The SSE event stream from /responses is converted back into the standard OpenAI Chat Completion chunk format, ensuring client compatibility and correct token usage tracking.

Considered Alternative: Proactive Service Discovery

An alternative, proactive "Service Discovery" approach was also investigated. This would involve making a GET https://api.githubcopilot.com/models request at startup or periodically to build a map of which models support which endpoints, based on the supported_endpoints field present in the response.

Pros of Service Discovery:

  • Eliminates the latency of the initial failed request for new models.

Cons / Reasons for Choosing the Fallback Approach:

  • API Stability Risk: The /models endpoint and its schema are part of an internal, undocumented API. Relying on its structure is fragile, as it could change without warning and break the integration for all GitHub models.
  • Reliability: The reactive fallback relies on a more stable "contract"—the API's explicit error response. This makes the proxy more resilient to future changes and new models that may not yet appear in the /models list.

Suggestion for the Future:
While the current reactive approach is safer, a hybrid model could be considered. We could implement Service Discovery as a "best-effort" optimization and still keep the reactive fallback as a safety net. This would provide the performance of Service Discovery with the resilience of the current implementation. I am happy to discuss or implement this if you think it's a worthwhile addition.

For now, this PR prioritizes robustness and future-proofing.

Related Issue

Fixes #102
photo_2026-02-15_03-08-01

Copilot AI review requested due to automatic review settings February 14, 2026 22:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements a reactive fallback in the GitHub provider executor to handle newer GitHub Copilot Codex models that reject /chat/completions and require the /responses endpoint (Fixes #102).

Changes:

  • Adds dynamic detection of the specific 400 error and retries the request against /responses.
  • Caches models known to require /responses to avoid repeating the initial failing request.
  • Translates OpenAI Chat Completions request/response streaming to/from the OpenAI Responses API format.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +99 to +108
if (parsed.done) {
controller.enqueue(new TextEncoder().encode("data: [DONE]\n\n"));
continue;
}

const converted = openaiResponsesToOpenAIResponse(parsed, state);
if (converted) {
const sseString = formatSSE(converted, "openai");
controller.enqueue(new TextEncoder().encode(sseString));
}

Copilot AI Feb 14, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TransformStream creates a new TextEncoder() for every parsed SSE line (new TextEncoder().encode(...) inside the loop). This adds avoidable overhead in high-token streams. Instantiate a single TextEncoder once (outside the loop / outside the transform callback) and reuse it for all enqueues.

Copilot uses AI. Check for mistakes.
Comment on lines +60 to +72
async executeWithResponsesEndpoint({ model, body, stream, credentials, signal, log }) {
const url = "https://api.githubcopilot.com/responses";
const headers = this.buildHeaders(credentials, stream);

const transformedBody = openaiToOpenAIResponsesRequest(model, body, stream, credentials);

log?.debug("GITHUB", "Sending translated request to /responses");

const response = await fetch(url, {
method: "POST",
headers,
body: JSON.stringify(transformedBody),
signal

Copilot AI Feb 14, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

executeWithResponsesEndpoint assumes the upstream /responses call returns an SSE stream and immediately pipes response.body through an SSE transformer, but headers/body are still built from the caller-provided stream flag. If stream is false, buildHeaders() sends Accept: application/json, which can cause /responses to return JSON and make the streaming parser/pipeThrough path fail. Either force streaming when using /responses (set Accept to text/event-stream and ensure the request body has stream: true), or add a non-streaming code path that parses JSON and converts it to the expected OpenAI chat completion response format.

Copilot uses AI. Check for mistakes.
Comment on lines +112 to +120
if (buffer.trim()) {
const parsed = parseSSELine(buffer.trim());
if (parsed && !parsed.done) {
const converted = openaiResponsesToOpenAIResponse(parsed, state);
if (converted) {
controller.enqueue(new TextEncoder().encode(formatSSE(converted, "openai")));
}
}
}

Copilot AI Feb 14, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In the SSE transformer flush(), a trailing data: [DONE] line can be dropped if it arrives without a terminating newline: parseSSELine(buffer.trim()) returns {done:true}, but the code only forwards buffered data when !parsed.done. This can leave clients waiting indefinitely for [DONE]. Handle the parsed.done case in flush() (enqueue data: [DONE]\n\n), or flush any remaining buffered line through the same parsed.done branch used in transform().

Suggested change
if (buffer.trim()) {
const parsed = parseSSELine(buffer.trim());
if (parsed && !parsed.done) {
const converted = openaiResponsesToOpenAIResponse(parsed, state);
if (converted) {
controller.enqueue(new TextEncoder().encode(formatSSE(converted, "openai")));
}
}
}
const trimmed = buffer.trim();
if (!trimmed) {
return;
}
const parsed = parseSSELine(trimmed);
if (!parsed) {
return;
}
if (parsed.done) {
controller.enqueue(new TextEncoder().encode("data: [DONE]\n\n"));
return;
}
const converted = openaiResponsesToOpenAIResponse(parsed, state);
if (converted) {
controller.enqueue(new TextEncoder().encode(formatSSE(converted, "openai")));
}

Copilot uses AI. Check for mistakes.
Comment on lines +47 to +54
if (result.response.status === HTTP_STATUS.BAD_REQUEST) {
const errorBody = await result.response.clone().text();

if (errorBody.includes("not accessible via the /chat/completions endpoint")) {
log?.warn("GITHUB", `Model ${model} requires /responses. Switching...`);
this.knownCodexModels.add(model);
return this.executeWithResponsesEndpoint(options);
}

Copilot AI Feb 14, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fallback trigger checks errorBody.includes("not accessible via the /chat/completions endpoint") on the raw response text. Since GitHub returns structured JSON errors, it would be more robust to JSON.parse and inspect error.message (and guard parse failures). This avoids false positives and makes the fallback resilient to formatting changes (whitespace, wrapping, localization).

Copilot uses AI. Check for mistakes.
@I3eka

I3eka commented Feb 14, 2026

Copy link
Copy Markdown
Contributor Author

Hi @decolua,

First off, thanks for this awesome project!

I've just submitted PR #127 which fixes the 400 Bad Request error for newer GitHub Codex models. The fix uses a resilient, reactive fallback mechanism: it tries the standard endpoint, and if it gets the specific "endpoint not accessible" error, it retries on /responses and caches that model's preference to avoid future latency.

While working on this, I explored a more proactive architectural approach that I wanted to share for your consideration, even though I opted for the safer method in the PR itself.

Alternative Idea: Proactive Service Discovery

I confirmed that the GET https://api.githubcopilot.com/models endpoint returns a full list of available models, and most importantly, each model object contains a supported_endpoints array.

For example, for gpt-5.2-codex:

{
    "id": "gpt-5.2-codex",
    ...
    "supported_endpoints": [
        "/responses"
    ],
    "vendor": "OpenAI"
}

And for a standard model like claude-opus-4.6:

{
    "id": "claude-opus-4.6",
    ...
    "supported_endpoints": [
        "/v1/messages",
        "/chat/completions"
    ],
    "vendor": "Anthropic"
}

We could implement a "Service Discovery" mechanism that fetches this data on startup (or periodically) and builds an in-memory map (model -> endpoint).

Pros of this approach:

  • Eliminates the "cold start" latency of the initial failed request for new Codex models, as we would proactively know the correct endpoint.

Why I didn't implement it in the PR:

  • API Stability Risk: My main concern is that api.githubcopilot.com is an internal, undocumented API. Relying on its schema (like the supported_endpoints field) feels fragile. Any unannounced change from GitHub could break the entire integration, whereas relying on an explicit error message feels like a more stable contract.

Suggestion for the Future:
The current PR prioritizes resilience. However, a hybrid approach could potentially offer the best of both worlds:

  1. Try to use the "Service Discovery" map for optimal performance.
  2. If a request fails with the known 400 error (e.g., for a model not yet in the cache), fall back to the reactive mechanism and update the map.

Request for Review

Could you please review the current implementation?

I'd also love to hear your thoughts on the architectural direction:

  • Should we stick with this reactive fallback for maximum resilience?
  • Or would you prefer a proactive Service Discovery (or a hybrid) to completely eliminate the "cold start" latency on the very first request?

I am happy to adjust the code based on your preference.

Thanks for the great project!

@decolua
decolua merged commit 6913129 into decolua:master Feb 15, 2026
@I3eka
I3eka deleted the fix/github-codex-endpoints branch February 15, 2026 02:18
LinearSakana pushed a commit to LinearSakana/9router that referenced this pull request Feb 15, 2026
…esponses endpoint (decolua#127)

* fix(github): add dynamic fallback to /responses for Codex models

* Refactor GithubExecutor: use config for URL detection

(cherry picked from commit 6913129)
kwanLeeFrmVi pushed a commit to kwanLeeFrmVi/9router that referenced this pull request Feb 26, 2026
…esponses endpoint (decolua#127)

* fix(github): add dynamic fallback to /responses for Codex models

* Refactor GithubExecutor: use config for URL detection
involvex added a commit to involvex/involvex-claude-router that referenced this pull request Feb 28, 2026
### Bug Fixes

* remove duplicated opencode prefix from model IDs and handle it in executor ([180f884](180f884))
* update opencode endpoint and model IDs ([6be8a03](6be8a03))

### Features

* add opencode provider ([03c846e](03c846e))

## [0.2.96](5645d0a...v0.2.96) (2026-02-24)

### Bug Fixes

*  GitHub Copilot model ([95fd950](95fd950))
* **auth:** allow HTTP for local network ([0a394d0](0a394d0))
* **auth:** prevent auto-login after logout ([49df3dc](49df3dc))
* **codex:** use user-agent detection for Droid CLI compatibility ([8c6e3b8](8c6e3b8))
* Correct indentation for clarity in chatCore and claude-to-openai response handlers ([fa06226](fa06226))
* correct token extraction for Claude non-streaming responses ([decolua#131](https://github.com/involvex/involvex-claude-router/issues/131)) ([9fbd6e6](9fbd6e6))
* **dashboard:** resolve 'Provider not found' for free providers ([45a4d3b](45a4d3b))
* **db:** improve error handling and null checks ([e6ef852](e6ef852))
* **gemini:** improve base64 image data parsing ([5645d0a](5645d0a))
* **github:** Implement dynamic fallback for Codex models requiring /responses endpoint ([decolua#127](https://github.com/involvex/involvex-claude-router/issues/127)) ([6913129](6913129))
* improve code formatting and reduce auto-refresh interval ([7f71916](7f71916))
* improve cursor auto-import reliability on macOS ([decolua#161](https://github.com/involvex/involvex-claude-router/issues/161)) ([d7e06c3](d7e06c3))
* **login:** avoid infinite loading on settings fetch failure ([01c9410](01c9410))
* **open-sse:** emit [DONE] in passthrough SSE mode ([decolua#142](https://github.com/involvex/involvex-claude-router/issues/142)) ([b9a6979](b9a6979))
* prevent race conditions in sticky round-robin ([3ad2f8d](3ad2f8d))
* resolve SonarQube findings and Next.js Image warnings ([7058b06](7058b06))
* update Codex executor for gpt-5.3-codex support ([d7d5dc9](d7d5dc9))

### Features

* add /v1/embeddings endpoint (OpenAI-compatible) ([decolua#146](https://github.com/involvex/involvex-claude-router/issues/146)) ([e1b8361](e1b8361)), closes [decolua#117](https://github.com/involvex/involvex-claude-router/issues/117)
* Add Anthropic Compatible provider support ([da5bdef](da5bdef))
* add API endpoint dimension to usage statistics dashboard ([decolua#152](https://github.com/involvex/involvex-claude-router/issues/152)) ([806bd4a](806bd4a))
* add Claude Opus 4.6 to GitHub Copilot provider ([decolua#97](https://github.com/involvex/involvex-claude-router/issues/97)) ([3d60597](3d60597))
* Add Claude Sonnet 4.6 to GitHub Copilot ([decolua#149](https://github.com/involvex/involvex-claude-router/issues/149)) ([4e2a3f8](4e2a3f8))
* add CLI entry points for claude-router commands ([62c632d](62c632d))
* add enable/disable toggle for provider connections ([ed796d2](ed796d2))
* add Gemini 3.1 Pro models to provider ([f2025cc](f2025cc))
* add Gemini embeddings support + Letta compatibility fixes ([a57a8ce](a57a8ce)), closes [decolua#148](decolua#148)
* add GLM 5 and MiniMax M2.5 models to providerModels.js; add Claude Sonnet 4.6 to CLI tools ([e1e5a81](e1e5a81))
* add GLM Coding (China) provider and Usage by API Keys statistics ([1ae4e31](1ae4e31))
* add GPT 4o to GitHub Copilot provider ([decolua#98](https://github.com/involvex/involvex-claude-router/issues/98)) ([c090bb0](c090bb0))
* add GPT 5.3 Codex Spark model to pricing and provider models ([decolua#133](https://github.com/involvex/involvex-claude-router/issues/133)) ([c7d4410](c7d4410))
* Add GPT 5.3 Codex to GitHub Copilot ([decolua#150](https://github.com/involvex/involvex-claude-router/issues/150)) ([c4aa424](c4aa424))
* add GPT-3.5 Turbo to GitHub Copilot provider ([e3dbd44](e3dbd44))
* add GPT-4 to GitHub Copilot provider ([6ade8ef](6ade8ef))
* add GPT-4o mini to GitHub Copilot provider ([053e490](053e490))
* add models management and router process control commands ([c65fd42](c65fd42))
* Add OpenAI-compatible provider nodes ([0a28f9f](0a28f9f))
* add password change functionality and dependencies ([23cfb19](23cfb19))
* add pause/resume functionality for API keys ([decolua#158](https://github.com/involvex/involvex-claude-router/issues/158)) ([73388a0](73388a0))
* add Qwen3.5 Coder Model configuration ([decolua#156](https://github.com/involvex/involvex-claude-router/issues/156)) ([f933dd9](f933dd9))
* add request logging functionality and usage metrics display ([e476907](e476907))
* add round-robin routing strategy ([9ebd7d3](9ebd7d3))
* add sticky round-robin routing strategy ([4f292aa](4f292aa))
* add support for GLM 5 (if) ([decolua#123](https://github.com/involvex/involvex-claude-router/issues/123)) ([03ab554](03ab554))
* add URL-based tab state persistence in usage page ([decolua#129](https://github.com/involvex/involvex-claude-router/issues/129)) ([6caef7f](6caef7f))
* allow custom user data directory via DATA_DIR environment variable ([d83bd86](d83bd86))
* **antigravity:** initial steps for Antigravity anti-ban alignment ([a229d79](a229d79)), closes [decolua#141](decolua#141)
* **antigravity:** integrate Antigravity tool with MITM support and update CLI tools ([2e854bd](2e854bd))
* **auth:** add model-level rate limit locking for multi-bucket providers ([decolua#120](https://github.com/involvex/involvex-claude-router/issues/120)) ([202fee7](202fee7)), closes [decolua#110](https://github.com/involvex/involvex-claude-router/issues/110)
* **auth:** Enhance authentication flow and settings management ([249fc28](249fc28))
* **cli-tools:** update CLI tools and add new models ([a2122e3](a2122e3))
* **cli-tools:** update default local endpoint port to 20128 ([6c41573](6c41573))
* **cli:** add new CLI package with basic scaffolding ([9cf4628](9cf4628))
* **cloud:** harden sync/auth flow, SSE fallback, and update changelog ([3d43983](3d43983))
* **codex:** add GPT 5.3, fix API translation, add thinking levels ([127475d](127475d))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([1c6dd6d](1c6dd6d))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([7b864a9](7b864a9))
* **codex:** Cursor compatibility + Next.js 16 proxy migration ([e9b0a73](e9b0a73))
* **config:** add Cloudflare MCP server and account ID configuration ([67297ea](67297ea))
* **cursor:** Add cursor Provider ([0a026c7](0a026c7))
* **cursor:** Integrate Cursor IDE support with OAuth import token flow ([137f315](137f315))
* **docker:** add Docker setup, environment examples, and architecture docs ([5e4a15b](5e4a15b))
* enhance disconnect handling and request tracking in chatCore.js ([decolua#126](https://github.com/involvex/involvex-claude-router/issues/126)) ([3d29b86](3d29b86))
* enhance request handling and error management in chatCore and streamToJsonConverter ([e2db638](e2db638))
* enhance usage stats with sortable columns and improved data handling ([bf6e09b](bf6e09b))
* Enhance usage tracking across response handlers ([a33924b](a33924b))
* **executors:**  Improved UI components for displaying provider limits and usage statistics in the dashboard. ([32aefe5](32aefe5))
* **iflow:** add IFlowExecutor with HMAC-SHA256 signature and enable models ([bd23ab4](bd23ab4))
* **iflow:** add kimi-k2.5 model support ([9e357a7](9e357a7))
* implement API key requirement toggle ([4cf25dc](4cf25dc))
* Implement buffer addition to usage tracking for improved context handling ([7881db8](7881db8))
* implement lazy loading for UsagePage with suspense fallback ([decolua#136](https://github.com/involvex/involvex-claude-router/issues/136)) ([05b09e6](05b09e6))
* implement provider connection reordering on create, update, and delete ([f2abcc6](f2abcc6))
* implement real project ID fetching for Antigravity ([decolua#170](https://github.com/involvex/involvex-claude-router/issues/170)) ([ea67742](ea67742))
* implement request tracking and enhance usage stats display ([e4f92cd](e4f92cd))
* implement usage tracking for AI requests ([9c3d6f4](9c3d6f4))
* Improve Antigravity quota monitoring and fix Droid CLI compatibility ([3c65e0c](3c65e0c))
* **open-sse:** add Claude Sonnet 4.6 ([b057c43](b057c43))
* OpenAI compatibility improvements & build fixes ([d9b8e48](d9b8e48)), closes [decolua#18](https://github.com/involvex/involvex-claude-router/issues/18)
* **provider:** add free providers and enhance error handling ([bdbe816](bdbe816))
* **providers:** add Minimax Coding (China) provider ([7c609d7](7c609d7))
* **providers:** add provider icons to dashboard ([60bd686](60bd686))
* **providers:** auto-validate API keys on save ([b275dfd](b275dfd))
* rename app to involvex-claude-router and add lenient JSON parsing ([bec206d](bec206d))
* **responses:** respect client streaming preference + string input support ([decolua#121](https://github.com/involvex/involvex-claude-router/issues/121)) ([ac7cedd](ac7cedd))
* **translator:** add thinking parameter support in OpenAI → Claude ([54e01d6](54e01d6))
* **ui:** add cost tracking to usage dashboard and pricing settings ([f302c88](f302c88))
* **ui:** add model support for custom providers and improve UX ([a7a52be](a7a52be))
* Update response handling and logging for improved usage tracking ([df0e1d6](df0e1d6))
* **usage:** implement cost tracking backend and pricing configuration ([a36afaa](a36afaa))
idbara pushed a commit to idbara/9router that referenced this pull request Apr 23, 2026
…esponses endpoint (decolua#127)

* fix(github): add dynamic fallback to /responses for Codex models

* Refactor GithubExecutor: use config for URL detection
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: GitHub Copilot Codex models fail with 400 error (require /responses endpoint)

3 participants