Skip to content

feat(github): route Claude models through Copilot's native /v1/messages - #2608

Closed
yidecode wants to merge 1 commit into
decolua:masterfrom
yidecode:feat/github-copilot-claude-messages-endpoint
Closed

yidecode wants to merge 1 commit into
decolua:masterfrom
yidecode:feat/github-copilot-claude-messages-endpoint

Conversation

@yidecode

Copy link
Copy Markdown
Contributor

Summary

  • GitHub Copilot's /chat/completions and /responses endpoints never surface prompt-cache token counts for Claude models. Copilot also exposes an Anthropic-native /v1/messages shim that does report cached_tokens.
  • Claude models (detected by name pattern — Copilot's live model catalog regularly exposes claude-* variants ahead of the static registry, so this is intentionally not a registry targetFormat field) are now routed through a new executeWithMessagesEndpoint(), translating OpenAI-shape requests to Claude natively (reusing translateRequest/translateResponse) instead of going through /chat/completions.
  • Fixes a real bug found along the way: translateRequest()'s internal bookkeeping field _toolNameMap was being sent upstream as-is, which made Anthropic's strict schema reject tool-call requests with a 400. It's now stripped and threaded through response state instead, same as the rest of the codebase does.
  • Removes the now-dead response_format Claude JSON-mode workaround in sanitizeMessagesForChatCompletions(), since Claude models no longer take that code path.

Test plan

  • node --check open-sse/executors/github.js / open-sse/providers/registry/github.js
  • npx eslint on both files (only a pre-existing, unrelated warning)
  • npx vitest run unit/github-responses-routing.test.js — 5/5 pass
  • Manually verified against a real Copilot account: streaming and non-streaming Claude requests both correctly hit cached_tokens on the second turn; reasoning_effort maps to Claude thinking.budget_tokens; streaming reasoning_content carries real thinking text; full tool_use → tool_result → final answer round-trip works
  • Audited GPT-5.x models — no changes needed, cache/reasoning already worked via the existing endpoints

GitHub Copilot's /chat/completions and /responses endpoints never
surface prompt-cache token counts for Claude models. Copilot also
exposes an Anthropic-native /v1/messages shim that does report
cached_tokens, so Claude models (detected by name pattern — Copilot's
live model catalog regularly exposes claude-* variants ahead of the
static registry, see registry/github.js) are now routed through a new
executeWithMessagesEndpoint(), translating OpenAI-shape requests to
Claude natively (reusing translateRequest/translateResponse) instead
of going through /chat/completions.

This also drops _toolNameMap — an internal bookkeeping field from
translateRequest() — before sending the body upstream; leaving it in
made Anthropic's strict schema reject tool-call requests with a 400.

The now-dead response_format Claude JSON-mode workaround in
sanitizeMessagesForChatCompletions() is removed since Claude models no
longer take that code path.
@decolua

decolua commented Jul 15, 2026

Copy link
Copy Markdown
Owner

Thanks @yidecode for the contribution! Reviewed and merged into master. 🙏

@decolua decolua closed this Jul 15, 2026
diegosouzapw added a commit to diegosouzapw/OmniRoute that referenced this pull request Jul 17, 2026
…ages (#7223)

* feat(sse): route GitHub Copilot Claude models through native /v1/messages

GitHub Copilot's /chat/completions and /responses endpoints never surface
prompt-cache token counts (cached_tokens) for Claude models, and round-tripping
Claude tool_use/tool_result/thinking content blocks through the OpenAI shape
is lossy. Copilot also exposes an Anthropic-native /v1/messages shim that
reports cached_tokens correctly and accepts native content blocks as-is.

Tag each github registry claude-* model with targetFormat: "claude" so
chatCore.ts translates the request to Anthropic-native shape before the
executor ever sees it (the same mechanism opencode/zen's Qwen entries and
opencode/go already use), and teach the github executor's buildUrl() /
buildHeaders() to dispatch those models at the new messagesUrl
(api.githubcopilot.com/v1/messages) with the required anthropic-version
header. transformRequest() now skips its /chat/completions-only quirks
(content-part flattening, trailing-assistant-prefill drop, the
response_format-as-system-prompt workaround) for the native path — the first
would destroy native tool_use/tool_result blocks, the prefill drop is
unnecessary because the real Anthropic API supports assistant prefill, and
the response_format workaround is superseded by the generic openai-to-claude
translator's own JSON-mode handling.

Co-authored-by: luoyide <ydhome.code@gmail.com>
Inspired-by: decolua/9router#2608

* chore(changelog): fragment for #7223

* fix(sse): green PR #7223 CI — complexity ratchet + stale test expectations

- Extract applyChatCompletionsOnlyQuirks() and resolveInitiatorHeader()
  out of GithubExecutor.transformRequest()/buildHeaders() so the two
  methods drop back under the complexity/cognitive-complexity ratchets
  (2058/891 -> 2056/890, matching the frozen baseline). No behavior
  change — same guards, just relocated.
- Update 4 pre-existing unit tests that hard-coded now-native claude-*
  Copilot ids (claude-sonnet-4.5/4.6) to exercise the /chat/completions
  legacy path via an unregistered id (claude-sonnet-4), matching the
  sibling test already using that pattern. These ids now intentionally
  route to the native /v1/messages shim added by this PR, which
  correctly skips the /chat/completions-only workarounds these tests
  were built to verify — the native path's own coverage lives in
  github-copilot-claude-native-messages.test.ts.
- Split the routing invariant test (copilot-gemini-claude-route-no-responses.test.ts)
  into a Claude case (expects /v1/messages) and a Gemini case (still
  expects /chat/completions), reflecting the intentional routing change.

---------

Co-authored-by: luoyide <ydhome.code@gmail.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…ages (diegosouzapw#7223)

* feat(sse): route GitHub Copilot Claude models through native /v1/messages

GitHub Copilot's /chat/completions and /responses endpoints never surface
prompt-cache token counts (cached_tokens) for Claude models, and round-tripping
Claude tool_use/tool_result/thinking content blocks through the OpenAI shape
is lossy. Copilot also exposes an Anthropic-native /v1/messages shim that
reports cached_tokens correctly and accepts native content blocks as-is.

Tag each github registry claude-* model with targetFormat: "claude" so
chatCore.ts translates the request to Anthropic-native shape before the
executor ever sees it (the same mechanism opencode/zen's Qwen entries and
opencode/go already use), and teach the github executor's buildUrl() /
buildHeaders() to dispatch those models at the new messagesUrl
(api.githubcopilot.com/v1/messages) with the required anthropic-version
header. transformRequest() now skips its /chat/completions-only quirks
(content-part flattening, trailing-assistant-prefill drop, the
response_format-as-system-prompt workaround) for the native path — the first
would destroy native tool_use/tool_result blocks, the prefill drop is
unnecessary because the real Anthropic API supports assistant prefill, and
the response_format workaround is superseded by the generic openai-to-claude
translator's own JSON-mode handling.

Co-authored-by: luoyide <ydhome.code@gmail.com>
Inspired-by: decolua/9router#2608

* chore(changelog): fragment for diegosouzapw#7223

* fix(sse): green PR diegosouzapw#7223 CI — complexity ratchet + stale test expectations

- Extract applyChatCompletionsOnlyQuirks() and resolveInitiatorHeader()
  out of GithubExecutor.transformRequest()/buildHeaders() so the two
  methods drop back under the complexity/cognitive-complexity ratchets
  (2058/891 -> 2056/890, matching the frozen baseline). No behavior
  change — same guards, just relocated.
- Update 4 pre-existing unit tests that hard-coded now-native claude-*
  Copilot ids (claude-sonnet-4.5/4.6) to exercise the /chat/completions
  legacy path via an unregistered id (claude-sonnet-4), matching the
  sibling test already using that pattern. These ids now intentionally
  route to the native /v1/messages shim added by this PR, which
  correctly skips the /chat/completions-only workarounds these tests
  were built to verify — the native path's own coverage lives in
  github-copilot-claude-native-messages.test.ts.
- Split the routing invariant test (copilot-gemini-claude-route-no-responses.test.ts)
  into a Claude case (expects /v1/messages) and a Gemini case (still
  expects /chat/completions), reflecting the intentional routing change.

---------

Co-authored-by: luoyide <ydhome.code@gmail.com>
golamrabbi696 pushed a commit to golamrabbi696/EzRouter that referenced this pull request Aug 12, 2026
decolua#2608 sent Claude models to Copilot's Anthropic-native /v1/messages shim by
branching inside the executor: chatCore still resolved github's target format
to "openai", and executeWithMessagesEndpoint translated openai -> claude on
the way out and claude -> openai on the way back.

That made targetFormat a lie, and a Claude-format client paid for it. A request
from Claude Code / Zed (directly, or through a downstream Anthropic proxy) went
claude -> openai -> claude, and request/claude-to-openai.js has no case for
thinking blocks — its switch handles text, image, tool_use and tool_result only.
The assistant turn's thinking block and signature were therefore dropped on the
way in, so on every follow-up turn the model lost its own prior reasoning. It
failed silently: Copilot accepts a history with no thinking block (verified), it
only rejects a thinking block whose signature is missing.

The provider now declares the routing decision where chatCore can see it:

  transport.resolveTargetFormat(model) -> "claude" | null

resolved at request time from the model NAME, because Copilot's live catalog
ships claude-* variants long before the static models[] list catches up (it
currently serves claude-sonnet-5, claude-opus-4.8, claude-opus-4.8-fast and
claude-fable-5, none of them listed). A static per-entry targetFormat would miss
exactly those. Providers without the hook are unaffected — the helper returns
null and the existing precedence chain is untouched.

With the target format honest, a claude-source request short-circuits in
translateRequest (sourceFormat === targetFormat) and reaches the shim untouched:
thinking blocks and real signatures survive verbatim. prepareClaudeRequest still
runs — it is gated on targetFormat === CLAUDE, outside the short-circuit — so
decolua#2608's cache_control injection and its prompt-cache token counts are preserved.
OpenAI-format clients still translate once, now in chatCore rather than the
executor.

executeWithMessagesEndpoint goes away entirely (-90 lines), along with its
hand-rolled SSE transform and its duplicate _toolNameMap threading: the generic
streaming handler already does both, keyed off the same target format.

由 Claude Code 辅助生成
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…ages (diegosouzapw#7223)

* feat(sse): route GitHub Copilot Claude models through native /v1/messages

GitHub Copilot's /chat/completions and /responses endpoints never surface
prompt-cache token counts (cached_tokens) for Claude models, and round-tripping
Claude tool_use/tool_result/thinking content blocks through the OpenAI shape
is lossy. Copilot also exposes an Anthropic-native /v1/messages shim that
reports cached_tokens correctly and accepts native content blocks as-is.

Tag each github registry claude-* model with targetFormat: "claude" so
chatCore.ts translates the request to Anthropic-native shape before the
executor ever sees it (the same mechanism opencode/zen's Qwen entries and
opencode/go already use), and teach the github executor's buildUrl() /
buildHeaders() to dispatch those models at the new messagesUrl
(api.githubcopilot.com/v1/messages) with the required anthropic-version
header. transformRequest() now skips its /chat/completions-only quirks
(content-part flattening, trailing-assistant-prefill drop, the
response_format-as-system-prompt workaround) for the native path — the first
would destroy native tool_use/tool_result blocks, the prefill drop is
unnecessary because the real Anthropic API supports assistant prefill, and
the response_format workaround is superseded by the generic openai-to-claude
translator's own JSON-mode handling.

Co-authored-by: luoyide <ydhome.code@gmail.com>
Inspired-by: decolua/9router#2608

* chore(changelog): fragment for diegosouzapw#7223

* fix(sse): green PR diegosouzapw#7223 CI — complexity ratchet + stale test expectations

- Extract applyChatCompletionsOnlyQuirks() and resolveInitiatorHeader()
  out of GithubExecutor.transformRequest()/buildHeaders() so the two
  methods drop back under the complexity/cognitive-complexity ratchets
  (2058/891 -> 2056/890, matching the frozen baseline). No behavior
  change — same guards, just relocated.
- Update 4 pre-existing unit tests that hard-coded now-native claude-*
  Copilot ids (claude-sonnet-4.5/4.6) to exercise the /chat/completions
  legacy path via an unregistered id (claude-sonnet-4), matching the
  sibling test already using that pattern. These ids now intentionally
  route to the native /v1/messages shim added by this PR, which
  correctly skips the /chat/completions-only workarounds these tests
  were built to verify — the native path's own coverage lives in
  github-copilot-claude-native-messages.test.ts.
- Split the routing invariant test (copilot-gemini-claude-route-no-responses.test.ts)
  into a Claude case (expects /v1/messages) and a Gemini case (still
  expects /chat/completions), reflecting the intentional routing change.

---------

Co-authored-by: luoyide <ydhome.code@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants