Skip to content

[WRONG BRANCH] fix(providers): give Kimi the Responses tool-result adjacency repair (#4726) - #4770

Merged
lidge-jun merged 11 commits into
codex/pw2-deepseek-reasoning-replayfrom
codex/pw3-kimi-tool-adjacency
Sep 16, 2026
Merged

lidge-jun merged 11 commits into
codex/pw2-deepseek-reasoning-replayfrom
codex/pw3-kimi-tool-adjacency

Conversation

@lidge-jun

Copy link
Copy Markdown
Owner

Summary

Kimi's Code Plan Responses endpoint (https://api.kimi.com/coding/v1) requires a tool result to follow its call immediately. When Codex Desktop's LSP hook injects a developer message between a code-mode exec call and its output, Kimi rejects the whole request with HTTP 400, naming the tool_call_id that "did not have response messages". Because the affected row replays its full history every turn, the session then fails on every subsequent turn rather than once.

opencodex already implements exactly this repair — normalizeResponsesToolResultAdjacency() in src/adapters/openai-responses/tool-output-recovery.ts — but passthrough.ts gates it on provider.requiresAdjacentResponsesToolResults === true, and only DeepSeek's registry entry seeded that flag. Kimi inherited no normalization, so a user who configures Kimi onto the Responses wire gets a permanently broken session.

This seeds the flag on both kimi (entries-core.ts) and kimi-code (entries-extended.ts). The existing fill-only derivation in providerConfigSeed, enrichProviderFromRegistry and routedProviderConfig carries it into new and already-persisted rows without overriding an explicit user value, and the flag is inert while these presets use the Chat wire.

The repair reorders; it does not delete. That distinction is the point of the change. The intervening developer message is preserved and moves after the batch, so this cannot be mistaken for silencing the 400 by discarding hook context. I read normalizeResponsesToolResultAdjacency line by line to confirm it before relying on it: it keeps every intervening non-tool item in original relative order, preserves function/custom tool-type pairing, and deliberately leaves duplicate, missing, or backwards pairs for the upstream to reject rather than guessing.

Two things the issue asserted that current source does not support, checked rather than assumed:

  • Neither Kimi entry seeds statelessResponses; that belongs to the reporter's own configuration. The defect and the fix are unaffected.
  • No upstream specification documents this adjacency requirement. The evidence is the reported 400 plus DeepSeek's identical failure shape under [Provider compatibility] DeepSeek V4 Flash returns 400 when developer message is interleaved between function_call and function_call_output #1292. Upstream Codex deliberately leaves an intervening developer message where it is and only synthesizes a missing result, so this stays a per-provider capability rather than becoming a wire-wide default. Moonshot's own tool documentation independently states that the messages following tool_calls must be exactly the matching role=tool messages, which is consistent with the observed behaviour.

Closes #4726

Verification

Static source review only, plus hosted CI. No local suite, typecheck, or build was run — the repository owner prohibits local suite execution in this lane after a past local run deleted real ~/.opencodex data.

Static checks performed:

  • Traced the seeding path end to end so the flag actually reaches the adapter: src/providers/derive.ts (providerConfigSeed, the registry enrichment merge, and the stale-row backfill), src/router.ts, then the gate in src/adapters/openai-responses/passthrough.ts.
  • Read normalizeResponsesToolResultAdjacency in full to confirm it preserves rather than drops intervening items, and that it refuses ambiguous duplicate or reversed pairs instead of reordering them.
  • Confirmed the new test file is registered in both scripts/test-layout/layout.json (explicit) and tests/fixtures/test-layout-expected.json, which tests/test-layout.test.ts and tests/test-layout-tooling.test.ts enforce.
  • Confirmed no existing test was deleted or weakened in this layer.
  • git diff --check clean.

Regression coverage in tests/providers/kimi-responses-adjacency.test.ts, mirroring the assertion style of tests/providers/deepseek-inbound-wire.test.ts:

  • both Kimi registry entries seed the adjacency capability
  • a stale persisted Kimi row is backfilled and activates adjacency repair on replay
  • moves a result next to its call while preserving an intervening developer message
  • keeps call_id pairing and all interleaved history with two outstanding replayed calls
  • leaves an already-adjacent call and result input untouched

structure/providers/chat-compat.md is updated: the adjacency pass is documented as gated by the capability rather than by provider name, with the Kimi evidence and the reason it is not a wire-wide default.

Hosted CI: this is a non-tip layer of a stacked lane and carries [skip ci] under the maintainer-approved DEV-STACK-08 tip-only policy. The lane's CI gate runs on the tip branch, which contains this commit.

Checklist

  • Scope stays focused and avoids unrelated cleanup.
  • Docs or release notes were updated when needed. (structure/providers/chat-compat.md; no user-facing configuration changed — the capability is seeded, not user-set)
  • Security-sensitive changes were reviewed for secrets, auth, and unsafe defaults. (registry metadata only; no auth, credential or token surface touched)

@lidge-jun
lidge-jun requested a review from Ingwannu as a code owner September 16, 2026 02:32
@coderabbitai

coderabbitai Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^dev$
  • ^preview$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 38fc2eb7-e19a-4be9-97a4-01f7a4a379b7

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-16T02:36:47.635487Z 68b5c73 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@lidge-jun

Copy link
Copy Markdown
Owner Author

리뷰 · 우선순위 77 / 80

이 PR은 Kimi를 Responses 와이어에 올렸을 때 나는 고질병을 고친다. Codex Desktop의 LSP 훅이 exec 같은 툴 호출과 그 결과 사이에 developer 메시지를 끼워 넣으면, Kimi Code Plan의 Responses 엔드포인트(https://api.kimi.com/coding/v1)가 “이 tool_call_id에 대한 응답이 없다”며 HTTP 400으로 통째로 거절한다. 문제는 한 번으로 끝나지 않는다. 그 행은 매 턴마다 전체 히스토리를 다시 보내기 때문에, 한 번 끼어든 간격이 있으면 그 세션은 이후 매 턴이 같은 400으로 죽는다. 이슈 #4726이 정확히 그 증상을 적었다.

opencodex 쪽에는 이미 같은 모양을 고치는 수리기가 있다. src/adapters/openai-responses/tool-output-recovery.ts의 normalizeResponsesToolResultAdjacency()가 호출과 결과를 바로 붙여 주고, 사이에 있던 메시지는 배치 뒤로 옮긴다(지우지 않는다). 그런데 passthrough.ts는 이 수리를 provider.requiresAdjacentResponsesToolResults === true일 때만 켠다. 현재 dev(HEAD 3070d64d8, 패키지 2.57.0)에서는 DeepSeek 레지스트리 항목만 그 플래그를 심어 두었다(#1292). Kimi 프리셋(kimi, kimi-code)에는 플래그가 없어서, 사용자가 Kimi를 Responses 어댑터로 바꾼 순간부터 수리기가 한 번도 돌지 않았다.

이 PR이 하는 일은 단순하다. entries-core.ts의 kimi와 entries-extended.ts의 kimi-code에 requiresAdjacentResponsesToolResults: true를 심는다. providerConfigSeed / enrichProviderFromRegistry / routedProviderConfig의 fill-only 경로가 새 행과 이미 저장된 행 모두에 플래그를 채운다(사용자가 이미 false로 박아 둔 값은 건드리지 않는다). 기본 어댑터는 여전히 openai-chat이라 Chat 와이어에서는 플래그가 잠들어 있고, Responses로 올린 뒤에만 살아난다. 그게 이슈가 재현한 설정과 같다.

테스트 tests/providers/kimi-responses-adjacency.test.ts는 DeepSeek inbound 테스트 톤을 그대로 따른다. 레지스트리 시드, stale 행 backfill + 실제 passthrough 재정렬, intervening developer 보존, 두 개 미해결 호출의 call_id 페어링, 이미 붙어 있는 입력은 그대로 두기까지 다섯 케이스가 있다. scripts/test-layout/layout.json과 tests/fixtures/test-layout-expected.json에도 등록되어 layout 가드에 걸린다. structure/providers/chat-compat.md는 “이름 하드코딩이 아니라 capability 게이트”라는 점을 Kimi 증거와 함께 적었다. 와이어 전역 기본값으로 올리지 않은 이유도 같다. 업스트림 Codex는 intervening developer를 그대로 두고, 명세가 아니라 관측된 400(+ DeepSeek #1292)만 근거이기 때문이다.

스택은 dev ← #4752(pw1) ← #4769(pw2) ← #4770(pw3, 이 PR) ← #4773(pw4 tip). 베이스는 dev가 아니라 codex/pw2-deepseek-reasoning-replay다. 헤드 커밋은 [skip ci](DEV-STACK-08 tip-only). MERGEABLE UNSTABLE, CI는 방금 큐에 들어간 상태다. types.ts/config.ts 스플릿과 충돌하는 모놀리스 경로는 건드리지 않았다. Closes #4726.

라인 - 없음(치명 버그). 아래는 운영·스택 관찰이다.

헤드 커밋 [skip ci] - 이 레이어 단독으로는 호스티드 스위트가 안 돈다. tip #4773에 CI가 실려야 한다.
base codex/pw2-deepseek-reasoning-replay - dev에 단독 머지하면 스택이 깨진다. 부모 #4752→#4769 먼저.
src/providers/registry/entries-core.ts kimi 시드 - 기본 어댑터는 여전히 openai-chat이라 Chat만 쓰는 행에는 효과가 없다(의도된 inert). Responses로 올린 뒤에만 수리기가 켜진다.
tests/providers/kimi-responses-adjacency.test.ts stale 픽스처 - statelessResponses: true는 이슈 리포터 설정 재현용이지, 레지스트리가 새로 시드하는 값이 아니다. 본문 주장과 맞다.
#4726 - 아직 OPEN. 이 PR이 랜딩되면 closes로 같이 닫혀야 한다.

메인테이너의 판단이 필요한 지점

너의 추천
스택 유지. 부모 #4752·#4769 랜딩 후 이 PR을 이어서 머지하고, 머지 시 #4726을 닫는다. tip #4773 CI가 이 커밋을 포함해 통과하는지 확인한 뒤 랜딩 열차에 태운다. 지금 당장 close/rebase 할 이유 없음.

이 댓글은 grok-bot이 작성했습니다

@github-actions

Copy link
Copy Markdown
Contributor

✅ Deterministic PR hygiene checks passed.

…4726) [skip ci]

Kimi's Code Plan Responses endpoint requires a tool result to follow its call
immediately. When the desktop LSP hook injects a developer message between a
code-mode exec call and its output, Kimi rejects the whole request with HTTP 400
naming the unanswered tool_call_id. Because the row replays full history every
turn, the session then fails permanently rather than once.

opencodex already implements exactly this repair in
normalizeResponsesToolResultAdjacency, but passthrough gates it on
requiresAdjacentResponsesToolResults, which only DeepSeek's entry seeded. Kimi
inherited no normalization, so a user who configures Kimi onto the Responses
wire hits the 400 on every affected turn.

Seed the flag on both kimi and kimi-code. The existing fill-only derivation in
providerConfigSeed, enrichProviderFromRegistry and routedProviderConfig carries
it into new and already-persisted rows without overriding an explicit user
value, and the flag is inert while these presets use the Chat wire.

The repair reorders; it does not delete. The intervening developer message is
preserved and moves after the batch, so the fix cannot be mistaken for silencing
the 400 by dropping hook context. Coverage pins that, plus call_id pairing with
two outstanding calls and interleaved noise, and an already-adjacent input being
left untouched.

No upstream specification documents the requirement; the evidence is the
reported 400 and DeepSeek's identical failure shape under #1292. Upstream Codex
deliberately leaves an intervening developer message where it is, so this stays
a per-provider capability rather than a wire-wide default.
@lidge-jun
lidge-jun force-pushed the codex/pw3-kimi-tool-adjacency branch from 68b5c73 to 70129ac Compare September 16, 2026 03:11
@lidge-jun
lidge-jun force-pushed the codex/pw2-deepseek-reasoning-replay branch from 862eef4 to 1eccd0f Compare September 16, 2026 03:11
lidge-jun and others added 10 commits September 16, 2026 14:01
…) [skip ci]

A custom provider whose baseUrl ends in a slash produced a doubled discovery
path: https://gateway.example.com/v1//models. Gateways that route the doubled
path as a distinct route reject it — the report observed HTTP 403 — so discovery
failed and the catalog silently fell back to the configured models.

The send paths were normalized already, by openaiChatCompletionsUrl and
openaiResponsesUrl, but discovery was not. buildModelsRequest appended the
endpoint verbatim, and resolveProviderModelDiscoveryUrl returns that default
unchanged for a provider with no registry spec, which is exactly the custom
case. Registry providers escaped it because new URL(spec.path, base) collapses
the doubled slash.

providerModelsUrl mirrors openaiChatCompletionsUrl rather than inventing a
second policy: trim outer whitespace and trailing slashes, drop an already
pasted /models, then append exactly one. Both production callers of the default
discovery URL use it — catalog discovery and API-key validation.

An existing path prefix is preserved, so /api/openai/v1 is not collapsed to the
origin. Registry spec.path, absolute endpoint overrides and relative endpoint
overrides all resolve exactly as before; a baseUrl already written without a
trailing slash is byte-identical to its previous output.
The CodeBuddy route launches the vendor CLI with --tools "" and
--strict-mcp-config, so the routed model has no native tool channel and writes
its call as prose. The shared coding-agent projection forwards text_delta
unrepaired, so that markup reached the client as an ordinary assistant answer.

Qoder's guard does not match it. The leaked tags are wrapped in FULLWIDTH
VERTICAL LINE (U+FF5C), which none of the shipped UNREPAIRABLE_MARKERS cover, so
this needed a signature of its own rather than a port.

Refusal requires the observed two-line grammar: a calls control line at column
zero, outside a Markdown fence, immediately followed by an invoke line naming a
functions.* tool. A lone tag, a quoted or inline-code literal, a fenced example,
a blockquote, indented source, or prose discussing the markup all carry extra
syntax before the tag and are forwarded untouched. Matching the marker alone
would refuse a legitimate answer that merely explains this protocol, which is
why the detector is narrower than the marker spelling.

A detected leak preserves the answer text already proven safe, emits one
non-retryable vendor_scaffold_detected error, and suppresses the vendor's later
success terminal so the client never sees a completed turn. Markers split across
streamed deltas are caught by holding only a bounded suffix that could still
complete a control sequence or a fence; unrelated pending text is released at
the next mismatch or terminal. The reasoning channel is guarded independently.

Leaked prose is never promoted into a real tool call. The text channel carries
no authenticated call envelope and no validated arguments, so converting it
would manufacture execution authority out of model output.

Kept CodeBuddy-owned rather than lifted into the shared coding-agent path, the
same containment #4234 chose for Qoder: the contract observed here is this
vendor's, and #4190's lane packet asked for a report rather than symmetry.

Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
…#4679) [skip ci]

Command Code's gateway rejects a request outright with 400 name must be at most
64 characters, got 66. Codex Desktop built-in app tools flatten to
<namespace>__<name> past that bound, a user cannot exclude them, and
Responses-Lite catalogs bundle every declared tool, so the surface cannot be
shrunk from configuration.

The bound belongs to the adapter, not to the shared name helper. Three adapters
already solve this for themselves: Kiro normalizes to its own charset with a
deterministic 8-hex suffix, Google compiles and restores names in its wire
compiler, and Meta Muse aliases names on api.meta.ai. The translated
openai-chat path is the only one with no answer, and it is the path Command Code
uses.

A request-scoped registry now owns one collision domain per translated Chat
Completions request, following Kiro's shape. A namespaced name whose flattened
spelling exceeds 64 characters becomes a charset-safe alias derived purely from
the native identity, so it is stable across processes, catalog order and catalog
membership. Declarations, replayed assistant tool calls and tool_choice all pass
through the same registry, and both the streaming and buffered parsers restore
the echoed alias before tool_call_start, so the existing bridge map still hands
the client its native {namespace, name}.

The registry is seeded from the union of the current catalog and the structured
tool calls still present in replay history, because a historical call can keep
its namespace without being redeclared; seeding from the catalog alone would let
exactly the reported over-limit name reach the gateway again on a later turn.

Nothing else changes. Names at or under 64 characters and bare names are
byte-identical on the wire, and Kiro, Google and Muse still receive the raw
flattened name and run their own normalization.

64 is the Chat Completions function-name limit and a strict-gateway
compatibility concern, not OpenAI Responses parity: upstream Codex raised its own
MCP ceiling to 128 bytes in openai/codex#39594 because native Responses accepts
128. Applying it on this wire is correct for that wire alone.

Carried from #4715. That PR placed the bound in the shared namespacedToolName
helper and was provisionally accepted there. Hosted CI then showed twice that the
shared point intercepts adapters which already had an answer: it broke Google's
wire-compiler restore, and after that was narrowed it broke Kiro's normalizer.
The problem statement and issue analysis are the original author's; only the
placement changed.

Co-authored-by: Hulian Buligon <205309211+HulianBuligon@users.noreply.github.com>
…guard

fix(codebuddy): refuse leaked vendor tool-call scaffolding (#4596)
…ames

fix(openai-chat): bound flattened tool wire names for strict gateways (#4679)
fix(catalog): normalize the custom-provider model-discovery join (#4724)
@lidge-jun

Copy link
Copy Markdown
Owner Author

Cascading downward. Kimi declares the adjacency capability so a hook-injected developer message no longer separates a tool call from its result. The middle message is preserved and the result is repositioned, rather than dropped to silence the 400.

Evidence at the verified tip 49f815d (tree cea66da56f10b8d1aee2290fa185d321442eab98), from run 35061163092:

  • test 1-4/4 and macos 1-2/2 all completed with conclusion success, confirmed through the check-runs API rather than the check rollup, so the heavy jobs actually executed and were not path-filtered. gates, changes, storage policy, api usage, docker smoke, keyring and npm-global on three platforms, the three service-lifecycle jobs, and the aggregate ci check all succeeded.
  • The ci failure at this commit belongs to run 35061161660, which this push superseded; run 35061163092 is the live one and it concluded success.
  • The lane absorbed dev at cf6e939 from the bottom layer upward, so each pull request keeps its own layer diff (4 / 9 / 6 / 7 / 7 / 5 files) and no dev commit appears in any layer's diff.
  • The single conflict was structure/transports/responses.md, where both sides appended a new section to the same empty base. It was resolved by keeping both: the section count goes from 14 on dev to 15 here, and every dev section name is still present. That loss is the kind CI cannot detect, so the names were compared directly rather than trusting the count.
  • The file-size ratchet reports no offenders after the absorption; openai-chat.ts sits exactly at its 822 cap.
  • git merge-tree --write-tree origin/dev <tip> reports a clean merge, and origin/dev is itself an ancestor of this tip.
  • Ancestry verified so each layer closes as MERGED: pw1 through pw5 are all ancestors of this tip.

Maintainer integration decision under MAINTAINERS.md / AGENTS.md: a maintainer with maintain or admin access may integrate into dev without a second maintainer approval, recording the decision and exact-head CI evidence.

@lidge-jun
lidge-jun merged commit aa31ed4 into codex/pw2-deepseek-reasoning-replay Sep 16, 2026
7 checks passed
@lidge-jun
lidge-jun deleted the codex/pw3-kimi-tool-adjacency branch September 16, 2026 06:07
@github-actions github-actions Bot changed the title fix(providers): give Kimi the Responses tool-result adjacency repair (#4726) [WRONG BRANCH] fix(providers): give Kimi the Responses tool-result adjacency repair (#4726) Sep 16, 2026
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

⏳ DRAFT

  • wrong target branch (codex/pw2-deepseek-reasoning-replay); retarget to dev.

What to do

  • Retarget this PR to dev — all contributions go to dev.

Its title has been prefixed with [WRONG BRANCH].
Automatic draft conversion failed (token cannot change draft status). Please convert this pull request to a draft manually. The required enforce-target check will keep failing until every issue above is resolved.

agentHits pushed a commit to agentHits/opencodex that referenced this pull request Sep 17, 2026
…adjacency

fix(providers): give Kimi the Responses tool-result adjacency repair (lidge-jun#4726)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant