Skip to content

fix(responses): unique output_index for message + tool call - #2908

Merged
steebchen merged 6 commits into
theopenco:mainfrom
serhiizghama:fix/streaming-responses-message-tool-output-index
Jul 23, 2026
Merged

steebchen merged 6 commits into
theopenco:mainfrom
serhiizghama:fix/streaming-responses-message-tool-output-index

Conversation

@serhiizghama

@serhiizghama serhiizghama commented Jul 3, 2026 •

Copy link
Copy Markdown
Contributor

When the chat-completions stream emits assistant text before a tool call — common when a model says something like "let me check" and then calls a function — the streamed Responses events gave the message and the function_call the same output_index.

The message output item opens at state.outputItemIndex but never advances it (unlike reasoning and tool items, which claim-and-increment), so the next tool call reuses index 0. On top of that the message's output_text.done/content_part.done/output_item.done events read the counter after the tool call bumped it, so the message's own added and done events disagreed on the index. And the final output array was always built as [reasoning, tools, message], which contradicts the streamed order once text comes first.

Reserve a dedicated index for the message when it opens (and record the reasoning index too), use it consistently across its added/delta/done events, and sort the final output array by output_index so it matches the stream. Content-only and tool-only streams are unaffected — their indices are unchanged.

Added a unit test for the text-then-tool_call case asserting the two items get distinct indices, the message keeps one index across added/done, and the completed output is ordered message-before-tool.

Summary by CodeRabbit

  • Bug Fixes
    • Improved streamed response item ordering with stable output_index values, keeping reasoning, assistant messages, and function/tool calls in the correct final sequence.
    • Updated completion payload generation to use index-based ordering so message content is consistently positioned relative to tool calls.
    • Ensured message-related streaming events (including annotation events) reuse the same output_index throughout the stream.
  • Tests
    • Expanded streaming conversion test coverage for mixed reasoning/content/tool-call flows, validating output_index stability and final output ordering.

When assistant text streamed before a tool call, the message output item
opened at outputItemIndex without advancing it, so the following tool call
reused the same output_index. The message's done events also read the
mutated counter, drifting from its added index. Reserve a dedicated index
for the message and order the final output array by output_index.
@coderabbitai

coderabbitai Bot commented Jul 3, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 74d921b7-06ef-4bf6-9186-26fd8d6ef326

📥 Commits

Reviewing files that changed from the base of the PR and between f97bc6d and a74a3e4.

📒 Files selected for processing (2)
  • apps/gateway/src/responses/responses.spec.ts
  • apps/gateway/src/responses/tools/convert-streaming-to-responses.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • apps/gateway/src/responses/responses.spec.ts

Walkthrough

This change gives streamed reasoning, assistant messages, and function calls distinct output indices. Completion events reuse the assistant message index, and response.completed sorts output items by index. Tests cover event index stability, annotation indices, and final output ordering.

Changes

Streaming Output Index Tracking

Layer / File(s) Summary
Streaming index assignment
apps/gateway/src/responses/tools/convert-streaming-to-responses.ts
Tracks dedicated reasoning and message indices and applies them to streamed message, text, and annotation events.
Indexed completion output assembly
apps/gateway/src/responses/tools/convert-streaming-to-responses.ts
Reuses the message index for completion events and sorts reasoning, tool-call, and message outputs by index.
Streaming output ordering validation
apps/gateway/src/responses/responses.spec.ts
Tests distinct and stable indices, annotation indices, and ordering across reasoning, content, and multi-chunk tool-call streams.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
    participant Client
    participant Converter as processStreamChunk
    participant State as StreamingState
    participant Completion as createCompletionEvents

    Client->>Converter: reasoning or content delta
    Converter->>State: assign output index
    Converter-->>Client: streamed output events
    Client->>Converter: tool_calls delta
    Converter-->>Client: function_call output event
    Client->>Completion: create completion events
    Completion->>State: read stored indices
    Completion->>Completion: sort output items by index
    Completion-->>Client: response.completed
Loading

Possibly related PRs

Suggested reviewers: ratchaw, steebchen, steebchen

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is relevant and clearly summarizes the core fix: assigning unique output_index values for message and tool call streaming.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
apps/gateway/src/responses/responses.spec.ts (1)

636-696: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add coverage for reasoning followed by streamed tool-call chunks.

This test covers message → tool. Since this PR also changes reasoningOutputIndex, add a case with reasoning, multiple tool-call argument chunks, and a later second tool call to assert emitted indices stay stable and align with final output order.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/gateway/src/responses/responses.spec.ts` around lines 636 - 696, Extend
the existing streaming response test around
processStreamChunk/createCompletionEvents to cover reasoning followed by
multiple tool-call chunks and a later second tool call. Add assertions that
reasoningOutputIndex stays stable across added/done events, that each streamed
function_call keeps a consistent output_index as arguments arrive in chunks, and
that the final response.completed output order matches the assigned output_index
sequence.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/gateway/src/responses/tools/convert-streaming-to-responses.ts`:
- Around line 263-267: The reasoning output slot is only being recorded in
convert-streaming-to-responses’s reasoning-start path, but state.outputItemIndex
is not advanced there, so later tool-call chunks can reuse a stale output_index.
Update the reasoning-start handling in convert-streaming-to-responses so that it
claims the current slot immediately by advancing state.outputItemIndex when
state.reasoningOutputIndex is set, and remove the deferred increments in the
later close-reasoning branch so the output_item.added / response.output
positions stay aligned.

---

Nitpick comments:
In `@apps/gateway/src/responses/responses.spec.ts`:
- Around line 636-696: Extend the existing streaming response test around
processStreamChunk/createCompletionEvents to cover reasoning followed by
multiple tool-call chunks and a later second tool call. Add assertions that
reasoningOutputIndex stays stable across added/done events, that each streamed
function_call keeps a consistent output_index as arguments arrive in chunks, and
that the final response.completed output order matches the assigned output_index
sequence.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: 6a0e03e1-ea75-43fd-9cd9-17568e667291

📥 Commits

Reviewing files that changed from the base of the PR and between 3301bfe and 2476be4.

📒 Files selected for processing (2)
  • apps/gateway/src/responses/responses.spec.ts
  • apps/gateway/src/responses/tools/convert-streaming-to-responses.ts

Comment thread apps/gateway/src/responses/tools/convert-streaming-to-responses.ts Outdated
@steebchen

Copy link
Copy Markdown
Member

thanks @serhiizghama! wondering if the PR coderabbit feedback is relevant before merging this? a specific proof of concept or payload to test would be also helpful, although not necessary

…nses-message-tool-output-index

# Conflicts:
#	apps/gateway/src/responses/tools/convert-streaming-to-responses.ts
@serhiizghama

Copy link
Copy Markdown
Contributor Author

Rebased on main — the conflict was just the message output item picking up the new state.annotations field; kept that alongside the output_index change. Spec green (39/39).

@steebchen

Copy link
Copy Markdown
Member

@serhiizghama please read my actual comment

The close-reasoning branch ran on every tool-call chunk while no message had started, so reasoning followed by multi-chunk tool calls inflated the shared index and a later tool call got an output_index past its final response.output position. Claim the reasoning slot once when reasoning starts and drop the deferred increments.
@serhiizghama

Copy link
Copy Markdown
Contributor Author

Sorry, you're right — I answered the wrong thing last time and skipped both of your actual questions.

On the coderabbit feedback: it was a real gap, and it was in my own fix. When reasoning is streamed and then tool calls arrive with no assistant message, the "close reasoning" branch ran on every tool-call chunk, because outputItemStarted only flips on content. So a multi-chunk tool call kept bumping the shared outputItemIndex, and a later tool call got an output_index past its real position in the final response.output. Same class of collision as the original bug, just in the reasoning path. Fixed it the way coderabbit suggested — claim the reasoning slot once when reasoning starts (reasoningOutputIndex = outputItemIndex++) and drop the two deferred increments.

For the proof of concept, here's the exact chunk sequence — reasoning, then a two-chunk get_weather call, then a second get_time call:

processStreamChunk({ choices: [{ delta: { reasoning: "thinking" } }] }, state);
processStreamChunk({ choices: [{ delta: { tool_calls: [{ index: 0, id: "call_a", function: { name: "get_weather", arguments: "" } }] } }] }, state);
processStreamChunk({ choices: [{ delta: { tool_calls: [{ index: 0, function: { arguments: '{"city":"NYC"}' } }] } }] }, state);
processStreamChunk({ choices: [{ delta: { tool_calls: [{ index: 1, id: "call_b", function: { name: "get_time", arguments: "{}" } }] } }] }, state);
createCompletionEvents(state);

Before the fix, the streamed events tell the client get_time is at output_index: 4, but it ends up at array position 2 in response.completed:

streamed output_index: get_weather -> 1, get_time -> 4
final response.output:  [reasoning(0), get_weather(1), get_time(2)]

After the fix they line up — get_time streams at output_index: 2, matching position 2:

streamed output_index: get_weather -> 1, get_time -> 2
final response.output:  [reasoning(0), get_weather(1), get_time(2)]

To be clear about how I verified this: I ran that sequence through the streaming-conversion functions directly (processStreamChunk → createCompletionEvents) and printed the indices — not a full request against a running gateway. It's covered by two new regression tests in responses.spec.ts (reasoning-then-message, and reasoning-then-multi-chunk-tool-calls), and the whole spec is green (41/41). Rebased on main too. Latest coderabbit review came back with no actionable comments.

@steebchen

Copy link
Copy Markdown
Member

Thanks for the thorough follow-up — verified your PoC and the fix, everything checks out. One thing before merge: response.output_text.annotation.added (convert-streaming-to-responses.ts, annotations delta block) still uses state.outputItemIndex, which now always points past the message's claimed slot — so web-search annotation events stream with the wrong output_index (this is a regression vs main, where the counter happened to still equal the message index there). Please switch it to state.messageOutputIndex (and ideally add a small test asserting the annotation event's output_index matches the message's). Then this is good to merge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@steebchen

Copy link
Copy Markdown
Member

Pushed the annotation fix directly (a74a3e4): response.output_text.annotation.added now uses state.messageOutputIndex, plus a regression test asserting the annotation event's output_index matches the message's. Spec is 42/42 locally.

@serhiizghama

Copy link
Copy Markdown
Contributor Author

Nice catch on the annotations event, thanks for fixing it directly — pulled it in, 42/42 still green on my end.

@steebchen
steebchen merged commit 0fa3a42 into theopenco:main Jul 23, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants