Skip to content

fix(handlers): honor the client format for non-streaming binary-transport replies - #3247

Closed
ntdatt812 wants to merge 1 commit into
decolua:masterfrom
ntdatt812:fix/nonstream-claude-binary-transport-3199
Closed

ntdatt812 wants to merge 1 commit into
decolua:masterfrom
ntdatt812:fix/nonstream-claude-binary-transport-3199

Conversation

@ntdatt812

@ntdatt812 ntdatt812 commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #3199

Problem

curl -s http://localhost:20128/v1/messages -H "anthropic-version: 2023-06-01" \
  -d '{"model":"kr/claude-haiku-4.5","max_tokens":20,"stream":false,"messages":[{"role":"user","content":"Say hi"}]}'

returns {"id":"chatcmpl-…","object":"chat.completion","choices":[…]} instead of an Anthropic message. Streaming against the same route is correct.

Providers on a binary or proprietary transport decode their upstream reply to OpenAI Chat Completions inside their own executor — the kiro executor emits chat.completion.chunk SSE, which handleNonStreamingResponse collapses via parseSSEToOpenAIResponse(). By the time translateNonStreamingResponse() runs, targetFormat is kiro and the body is already OpenAI-shaped.

translateNonStreamingResponse() branches on openai, gemini/vertex/antigravity, claude and ollama. kiro (and cursor, commandcode) match none of them, so control reached the trailing return responseBody and the OpenAI shape went straight to a Claude client. The existing targetFormat === OPENAI && sourceFormat === CLAUDE branch was the right conversion — it just was not reachable for these targets.

Downstream impact: LiteLLM's anthropic/ prefix expects an Anthropic body from /v1/messages and fails with KeyError: 'content'.

Fix

The "body is OpenAI Chat Completions — what does the client speak?" dispatch existed once already, keyed on targetFormat === OPENAI. Rather than write it a second time, it is extracted into fromOpenAICompletion() and called from both places: the existing targetFormat === OPENAI branch, and a new trailing guard for bodies that carry a choices array under a target format with no branch of its own.

Net result is one list of client formats instead of two. The Responses arm covers the identical gap for a Responses client on the same targets. An OpenAI client is unaffected, and a body without choices is returned untouched.

Scope note — the deeper seam, deliberately not taken

The codebase already has the right long-term mechanism for this: result.responseFormat, honored in chatCore.js ("Most executors return their registry format. Cursor AgentService is an exception: it is decoded by the executor into OpenAI-compatible output."). Today only executors/cursor.js sets it.

executors/kiro.js and executors/commandcode.js both qualify — kiro's execute() always runs transformEventStreamToSSE, and commandcode's always runs wrapNdjsonAsOpenAISse(), so both unconditionally return OpenAI Chat Completions. Declaring responseFormat: FORMATS.OPENAI on those two executors would fix this issue with no new code at all, since the existing targetFormat === OPENAI branch would then handle them.

I did not do that here because providerResponseFormat also feeds handleStreamingResponse, so it would reroute kiro's streaming path off the direct kiro:claude route onto openai:claude and make translator/response/kiro-to-claude.js dead. That should be behavior-neutral, but it is a much larger change on a heavily used provider and belongs in its own PR. Flagging it so the direction is on record.

Tests

New: tests/unit/nonstream-binary-transport-source-format-3199.test.js — a kiro-decoded body converts to an Anthropic message (text, stop_reason, usage), tool calls become tool_use blocks, a Responses client gets object:"response", an OpenAI client's body is returned by identity, and a body with no choices is untouched.

Full suite (npx vitest run unit translator) against master as the baseline: 1670 → 1675 passed, 91 failed on both sides — no pass→fail, the 5 new passes are this PR's tests.

…port replies

Providers on a binary or proprietary transport — kiro EventStream, cursor
protobuf, commandcode NDJSON — decode their upstream reply to OpenAI Chat
Completions inside their own executor. translateNonStreamingResponse has no
branch for those target formats, so it fell through and returned the raw
chat.completion body regardless of what the client asked for: a stream:false
request to /v1/messages handed a Claude client `object:"chat.completion"`, and
LiteLLM's anthropic/ prefix died on KeyError: 'content'. The streaming path was
already correct.

Convert on the way out when the body carries a choices array and the client
speaks Claude or Responses.

Fixes decolua#3199
@ntdatt812
ntdatt812 force-pushed the fix/nonstream-claude-binary-transport-3199 branch from a046a22 to 02d402c Compare August 15, 2026 02:07
@momotseuk

Copy link
Copy Markdown

will this fix be incorporated into next version?

@ntdatt812

Copy link
Copy Markdown
Contributor Author

I wrote the patch, so I cannot answer for the release schedule — that is @decolua's call. What I can give you is the current state and a workaround.

State. Open since 12 Aug, MERGEABLE / CLEAN against master (699edac3), waiting on review. Nothing is blocking it from my side.

I re-checked it just now rather than assuming it had kept: merged the current master into the branch — no conflicts — and ran its tests on the result.

tests/unit/nonstream-binary-transport-source-format-3199.test.js
Test Files  1 passed (1)
     Tests  5 passed (5)

Workaround until it lands. Only the non-streaming path is affected; streaming through the same route already returns the right shape. So on a binary-transport provider (kr/*, cursor/*) sending "stream": true and collapsing the SSE client-side gives you a correct Anthropic message today.

@ntdatt812

Copy link
Copy Markdown
Contributor Author

Superseded by #3878, which carries this commit unchanged, rebased onto current master and grouped with the other fixes in the same area. Closing here so the two do not sit in the queue as duplicates — reopen if you would rather review it on its own.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: Non-streaming /v1/messages returns OpenAI format instead of Anthropic format (Kiro provider)

2 participants