Repository navigation
Conversation
GitHub Copilot's /chat/completions and /responses endpoints never surface prompt-cache token counts for Claude models. Copilot also exposes an Anthropic-native /v1/messages shim that does report cached_tokens, so Claude models (detected by name pattern — Copilot's live model catalog regularly exposes claude-* variants ahead of the static registry, see registry/github.js) are now routed through a new executeWithMessagesEndpoint(), translating OpenAI-shape requests to Claude natively (reusing translateRequest/translateResponse) instead of going through /chat/completions. This also drops _toolNameMap — an internal bookkeeping field from translateRequest() — before sending the body upstream; leaving it in made Anthropic's strict schema reject tool-call requests with a 400. The now-dead response_format Claude JSON-mode workaround in sanitizeMessagesForChatCompletions() is removed since Claude models no longer take that code path.
Owner
|
Thanks @yidecode for the contribution! Reviewed and merged into master. 🙏 |
7 tasks done
diegosouzapw
added a commit
to diegosouzapw/OmniRoute
that referenced
this pull request
Jul 17, 2026
…ages (#7223) * feat(sse): route GitHub Copilot Claude models through native /v1/messages GitHub Copilot's /chat/completions and /responses endpoints never surface prompt-cache token counts (cached_tokens) for Claude models, and round-tripping Claude tool_use/tool_result/thinking content blocks through the OpenAI shape is lossy. Copilot also exposes an Anthropic-native /v1/messages shim that reports cached_tokens correctly and accepts native content blocks as-is. Tag each github registry claude-* model with targetFormat: "claude" so chatCore.ts translates the request to Anthropic-native shape before the executor ever sees it (the same mechanism opencode/zen's Qwen entries and opencode/go already use), and teach the github executor's buildUrl() / buildHeaders() to dispatch those models at the new messagesUrl (api.githubcopilot.com/v1/messages) with the required anthropic-version header. transformRequest() now skips its /chat/completions-only quirks (content-part flattening, trailing-assistant-prefill drop, the response_format-as-system-prompt workaround) for the native path — the first would destroy native tool_use/tool_result blocks, the prefill drop is unnecessary because the real Anthropic API supports assistant prefill, and the response_format workaround is superseded by the generic openai-to-claude translator's own JSON-mode handling. Co-authored-by: luoyide <ydhome.code@gmail.com> Inspired-by: decolua/9router#2608 * chore(changelog): fragment for #7223 * fix(sse): green PR #7223 CI — complexity ratchet + stale test expectations - Extract applyChatCompletionsOnlyQuirks() and resolveInitiatorHeader() out of GithubExecutor.transformRequest()/buildHeaders() so the two methods drop back under the complexity/cognitive-complexity ratchets (2058/891 -> 2056/890, matching the frozen baseline). No behavior change — same guards, just relocated. - Update 4 pre-existing unit tests that hard-coded now-native claude-* Copilot ids (claude-sonnet-4.5/4.6) to exercise the /chat/completions legacy path via an unregistered id (claude-sonnet-4), matching the sibling test already using that pattern. These ids now intentionally route to the native /v1/messages shim added by this PR, which correctly skips the /chat/completions-only workarounds these tests were built to verify — the native path's own coverage lives in github-copilot-claude-native-messages.test.ts. - Split the routing invariant test (copilot-gemini-claude-route-no-responses.test.ts) into a Claude case (expects /v1/messages) and a Gemini case (still expects /chat/completions), reflecting the intentional routing change. --------- Co-authored-by: luoyide <ydhome.code@gmail.com>
HouMinXi
pushed a commit
to HouMinXi/OmniRoute
that referenced
this pull request
Aug 2, 2026
…ages (diegosouzapw#7223) * feat(sse): route GitHub Copilot Claude models through native /v1/messages GitHub Copilot's /chat/completions and /responses endpoints never surface prompt-cache token counts (cached_tokens) for Claude models, and round-tripping Claude tool_use/tool_result/thinking content blocks through the OpenAI shape is lossy. Copilot also exposes an Anthropic-native /v1/messages shim that reports cached_tokens correctly and accepts native content blocks as-is. Tag each github registry claude-* model with targetFormat: "claude" so chatCore.ts translates the request to Anthropic-native shape before the executor ever sees it (the same mechanism opencode/zen's Qwen entries and opencode/go already use), and teach the github executor's buildUrl() / buildHeaders() to dispatch those models at the new messagesUrl (api.githubcopilot.com/v1/messages) with the required anthropic-version header. transformRequest() now skips its /chat/completions-only quirks (content-part flattening, trailing-assistant-prefill drop, the response_format-as-system-prompt workaround) for the native path — the first would destroy native tool_use/tool_result blocks, the prefill drop is unnecessary because the real Anthropic API supports assistant prefill, and the response_format workaround is superseded by the generic openai-to-claude translator's own JSON-mode handling. Co-authored-by: luoyide <ydhome.code@gmail.com> Inspired-by: decolua/9router#2608 * chore(changelog): fragment for diegosouzapw#7223 * fix(sse): green PR diegosouzapw#7223 CI — complexity ratchet + stale test expectations - Extract applyChatCompletionsOnlyQuirks() and resolveInitiatorHeader() out of GithubExecutor.transformRequest()/buildHeaders() so the two methods drop back under the complexity/cognitive-complexity ratchets (2058/891 -> 2056/890, matching the frozen baseline). No behavior change — same guards, just relocated. - Update 4 pre-existing unit tests that hard-coded now-native claude-* Copilot ids (claude-sonnet-4.5/4.6) to exercise the /chat/completions legacy path via an unregistered id (claude-sonnet-4), matching the sibling test already using that pattern. These ids now intentionally route to the native /v1/messages shim added by this PR, which correctly skips the /chat/completions-only workarounds these tests were built to verify — the native path's own coverage lives in github-copilot-claude-native-messages.test.ts. - Split the routing invariant test (copilot-gemini-claude-route-no-responses.test.ts) into a Claude case (expects /v1/messages) and a Gemini case (still expects /chat/completions), reflecting the intentional routing change. --------- Co-authored-by: luoyide <ydhome.code@gmail.com>
golamrabbi696
pushed a commit
to golamrabbi696/EzRouter
that referenced
this pull request
Aug 12, 2026
decolua#2608 sent Claude models to Copilot's Anthropic-native /v1/messages shim by branching inside the executor: chatCore still resolved github's target format to "openai", and executeWithMessagesEndpoint translated openai -> claude on the way out and claude -> openai on the way back. That made targetFormat a lie, and a Claude-format client paid for it. A request from Claude Code / Zed (directly, or through a downstream Anthropic proxy) went claude -> openai -> claude, and request/claude-to-openai.js has no case for thinking blocks — its switch handles text, image, tool_use and tool_result only. The assistant turn's thinking block and signature were therefore dropped on the way in, so on every follow-up turn the model lost its own prior reasoning. It failed silently: Copilot accepts a history with no thinking block (verified), it only rejects a thinking block whose signature is missing. The provider now declares the routing decision where chatCore can see it: transport.resolveTargetFormat(model) -> "claude" | null resolved at request time from the model NAME, because Copilot's live catalog ships claude-* variants long before the static models[] list catches up (it currently serves claude-sonnet-5, claude-opus-4.8, claude-opus-4.8-fast and claude-fable-5, none of them listed). A static per-entry targetFormat would miss exactly those. Providers without the hook are unaffected — the helper returns null and the existing precedence chain is untouched. With the target format honest, a claude-source request short-circuits in translateRequest (sourceFormat === targetFormat) and reaches the shim untouched: thinking blocks and real signatures survive verbatim. prepareClaudeRequest still runs — it is gated on targetFormat === CLAUDE, outside the short-circuit — so decolua#2608's cache_control injection and its prompt-cache token counts are preserved. OpenAI-format clients still translate once, now in chatCore rather than the executor. executeWithMessagesEndpoint goes away entirely (-90 lines), along with its hand-rolled SSE transform and its duplicate _toolNameMap threading: the generic streaming handler already does both, keyed off the same target format. 由 Claude Code 辅助生成
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…ages (diegosouzapw#7223) * feat(sse): route GitHub Copilot Claude models through native /v1/messages GitHub Copilot's /chat/completions and /responses endpoints never surface prompt-cache token counts (cached_tokens) for Claude models, and round-tripping Claude tool_use/tool_result/thinking content blocks through the OpenAI shape is lossy. Copilot also exposes an Anthropic-native /v1/messages shim that reports cached_tokens correctly and accepts native content blocks as-is. Tag each github registry claude-* model with targetFormat: "claude" so chatCore.ts translates the request to Anthropic-native shape before the executor ever sees it (the same mechanism opencode/zen's Qwen entries and opencode/go already use), and teach the github executor's buildUrl() / buildHeaders() to dispatch those models at the new messagesUrl (api.githubcopilot.com/v1/messages) with the required anthropic-version header. transformRequest() now skips its /chat/completions-only quirks (content-part flattening, trailing-assistant-prefill drop, the response_format-as-system-prompt workaround) for the native path — the first would destroy native tool_use/tool_result blocks, the prefill drop is unnecessary because the real Anthropic API supports assistant prefill, and the response_format workaround is superseded by the generic openai-to-claude translator's own JSON-mode handling. Co-authored-by: luoyide <ydhome.code@gmail.com> Inspired-by: decolua/9router#2608 * chore(changelog): fragment for diegosouzapw#7223 * fix(sse): green PR diegosouzapw#7223 CI — complexity ratchet + stale test expectations - Extract applyChatCompletionsOnlyQuirks() and resolveInitiatorHeader() out of GithubExecutor.transformRequest()/buildHeaders() so the two methods drop back under the complexity/cognitive-complexity ratchets (2058/891 -> 2056/890, matching the frozen baseline). No behavior change — same guards, just relocated. - Update 4 pre-existing unit tests that hard-coded now-native claude-* Copilot ids (claude-sonnet-4.5/4.6) to exercise the /chat/completions legacy path via an unregistered id (claude-sonnet-4), matching the sibling test already using that pattern. These ids now intentionally route to the native /v1/messages shim added by this PR, which correctly skips the /chat/completions-only workarounds these tests were built to verify — the native path's own coverage lives in github-copilot-claude-native-messages.test.ts. - Split the routing invariant test (copilot-gemini-claude-route-no-responses.test.ts) into a Claude case (expects /v1/messages) and a Gemini case (still expects /chat/completions), reflecting the intentional routing change. --------- Co-authored-by: luoyide <ydhome.code@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
/chat/completionsand/responsesendpoints never surface prompt-cache token counts for Claude models. Copilot also exposes an Anthropic-native/v1/messagesshim that does reportcached_tokens.claude-*variants ahead of the static registry, so this is intentionally not a registrytargetFormatfield) are now routed through a newexecuteWithMessagesEndpoint(), translating OpenAI-shape requests to Claude natively (reusingtranslateRequest/translateResponse) instead of going through/chat/completions.translateRequest()'s internal bookkeeping field_toolNameMapwas being sent upstream as-is, which made Anthropic's strict schema reject tool-call requests with a 400. It's now stripped and threaded through response state instead, same as the rest of the codebase does.response_formatClaude JSON-mode workaround insanitizeMessagesForChatCompletions(), since Claude models no longer take that code path.Test plan
node --check open-sse/executors/github.js/open-sse/providers/registry/github.jsnpx eslinton both files (only a pre-existing, unrelated warning)npx vitest run unit/github-responses-routing.test.js— 5/5 passcached_tokenson the second turn;reasoning_effortmaps to Claudethinking.budget_tokens; streamingreasoning_contentcarries real thinking text; full tool_use → tool_result → final answer round-trip works