Skip to content

Surface tool calls and outputs in OpenAI Responses API - #5010

Merged
ilblackdragon merged 4 commits into
mainfrom
feat/openai-responses-tool-outputs
Jun 17, 2026
Merged

ilblackdragon merged 4 commits into
mainfrom
feat/openai-responses-tool-outputs

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

What

The non-streaming Responses API (POST /v1/responses create-and-wait and GET /v1/responses/{id}) previously returned only the assistant's final text Message. Tool activity in the run was invisible to clients.

This projects each run's tool calls and their raw outputs into the response output array:

  • a function_call item (call_id / name / arguments), paired with
  • its function_call_output (the raw model_observation tool result),
  • followed by the assistant Message.

How

All changes are in the composition projection reader OpenAiResponsesThreadProjectionReader (crates/ironclaw_reborn_composition/src/openai_compat_serve.rs). The OpenAI-compat DTOs already supported these item types, so no change was needed in ironclaw_reborn_openai_compat.

Neither thread-service read alone carries everything:

  • the history projection keeps turn_run_id (run attribution) + the tool-result envelope content (the raw model_observation) but strips tool_result_provider_call;
  • the context projection (load_context_messages) preserves the provider call (function name/arguments/call id) but drops turn_run_id.

So the reader joins list_thread_history and load_context_messages by message_id. It reuses the existing load_context_messages read (which already legitimately exposes the provider call to model-context consumers) rather than widening the guarded ironclaw_threads trait.

Ordering: tool items are always emitted before the assistant message. A run's single assistant draft reserves its sequence at turn start, which can predate the tool results it later produces, so transcript order alone would mis-order them.

Polling cost: wait_for_response_completion checks a cheap completion gate (finalized_assistant_message_by_run) on each tick and builds the full projection once, rather than re-reading the whole transcript every 100ms.

Tests

Three integration-tier tests drive the production reader through a real InMemorySessionThreadService (per the "test through the caller" rule):

  • paired function_call + function_call_output with raw model_observation output,
  • in-progress run surfaces tool output without a final message,
  • tool items ordered before the assistant message even when the assistant draft's sequence predates the tool result.

cargo fmt + clippy --tests clean; module tests pass.

Notes / caveats

  • Streaming not covered. Only the non-streaming Responses path changes; stream: true runs through a different projection-streamer mechanism and still emits text only.
  • Policy exception (intentional). This deliberately exposes the raw model_observation and provider tool name/arguments through the API surface — a documented exception to the crate's narrow-DTO / no-raw-provider-diagnostics policy. The sanitized safe-summary alternative would stay within policy if preferred.

The non-streaming Responses API (`POST /v1/responses` create-and-wait and
`GET /v1/responses/{id}`) previously returned only the assistant's final
text message. Tool activity in the run was invisible to clients.

This projects each run's tool calls and their raw outputs into the
response `output` array: a `function_call` item (call_id / name /
arguments) paired with its `function_call_output` (the raw
`model_observation` tool result), followed by the assistant `Message`.

Implementation lives in the composition projection reader
(`OpenAiResponsesThreadProjectionReader`). Neither thread-service read
alone carries everything — the history projection keeps `turn_run_id`
plus the tool-result envelope content but strips
`tool_result_provider_call`, while the context projection preserves the
provider call but drops run attribution — so the reader joins
`list_thread_history` and `load_context_messages` by `message_id`.

Tool items are always emitted before the assistant message: a run's
single assistant draft reserves its sequence at turn start, which can
predate the tool results it produces, so transcript order alone would
mis-order them. The wait poll loop checks a cheap completion gate
(`finalized_assistant_message_by_run`) and builds the full projection
once, rather than re-reading the whole transcript every tick.

Notes:
- Only the non-streaming Responses path is covered; `stream: true` still
  emits text only (separate projection-streamer mechanism).
- This deliberately exposes the raw `model_observation` and provider tool
  name/arguments through the API surface, a documented exception to the
  crate's narrow-DTO policy.
Copilot AI review requested due to automatic review settings June 17, 2026 03:51
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules labels Jun 17, 2026
@coderabbitai

coderabbitai Bot commented Jun 17, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4799e50d-02c2-475c-a33e-977476f86f00

📥 Commits

Reviewing files that changed from the base of the PR and between 31eacd4 and bb33b3c.

📒 Files selected for processing (1)
  • crates/ironclaw_reborn_composition/src/openai_compat_serve.rs

📝 Walkthrough

Summary by CodeRabbit

  • Bug Fixes
    • Improved ordering of tool calls and results in API responses to ensure correct sequencing
    • Fixed assistant message finalization in response streams
    • Enhanced error handling for session-related failures with consistent response formatting

Walkthrough

The OpenAI-compatible "Responses" projection in openai_compat_serve.rs is reworked to emit full output item streams. New helpers (read_run_output, load_provider_calls, run_completed) join thread history with provider-call metadata, enforce tool-before-assistant ordering, and replace the prior single-string content assembly. Callers and response_object are updated accordingly, and unit tests cover the new projection behavior.

Changes

OpenAI Responses Output Projection Pipeline

Layer / File(s) Summary
Core projection helpers
crates/ironclaw_reborn_composition/src/openai_compat_serve.rs
Adds RunResponseProjection, read_run_output (joins tool-result messages with provider-call metadata loaded from context, enforces tool-items-before-assistant ordering, parses ToolResultReferenceEnvelope), load_provider_calls (keyed by message id), run_completed polling gate, and map_thread_read_error for OpenAI-compat error mapping. Import groupings are adjusted alongside.
Updated callers and response_object signature
crates/ironclaw_reborn_composition/src/openai_compat_serve.rs
wait_for_response_completion polls via run_completed then assembles OpenAiResponseProjection from read_run_output. read_response derives OpenAiResponseStatus from assistant_finalized and passes projection.items to response_object. response_object changed from content: Option<String> to output: Vec<OpenAiResponseOutputItem>.
Unit tests
crates/ironclaw_reborn_composition/src/openai_compat_serve.rs
Tests cover: paired function_call/function_call_output emission with model_observation passthrough, tool items ordered before assistant message, and in-progress runs returning only tool items with assistant_finalized = false.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Poem

🐇 Hop, hop through the thread history lane,
Tool calls and results, neatly arranged in a chain!
Provider metadata retrieved from the store,
Assistant finalized? Then we poll no more.
Output items ordered, the rabbit declares—
A tidy projection assembled with care! 🌿

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 inconclusive)

Check name Status Explanation Resolution
Description check ❓ Inconclusive The description covers the 'What' and 'How' sections with clear motivation and implementation details. However, required template sections like Change Type, Linked Issue, and Validation checklist are not filled out according to the template structure. Complete the missing required template sections: mark applicable Change Type boxes, provide Linked Issue reference, check relevant Validation items, and fill in Security Impact and Database Impact sections.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately and concisely summarizes the main change: surfacing tool calls and outputs in the OpenAI Responses API, which matches the core objective of making previously invisible tool activity visible to clients.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/openai-responses-tool-outputs
⚔️ Resolve merge conflicts
  • Resolve merge conflict in branch feat/openai-responses-tool-outputs

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added the contributor: core 20+ merged PRs label Jun 17, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5010 June 17, 2026 03:51 Destroyed
@railway-app

railway-app Bot commented Jun 17, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5010 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jun 17, 2026 at 6:33 pm

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds tool-call visibility to the non-streaming OpenAI-compatible Responses API projection by emitting function_call + function_call_output items (with raw tool outputs) ahead of the assistant message, and optimizes the wait/poll loop to avoid repeatedly re-reading the full transcript.

Changes:

  • Project each run’s tool calls and raw tool outputs into the Responses output array (paired function_call / function_call_output), then the finalized assistant message.
  • Join thread history and context projections by message_id to recover provider call metadata that history intentionally strips.
  • Add integration-tier tests exercising the production projection reader via InMemorySessionThreadService.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +490 to +496
match ToolResultReferenceEnvelope::from_json_str(content) {
Ok(envelope) => envelope.model_observation.unwrap_or_else(|| {
serde_json::Value::String(envelope.safe_summary.as_str().to_string())
}),
Err(_) => serde_json::Value::String(content.to_string()),
}
}

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 71eb66e. tool_result_output now deserializes the envelope shape directly (best-effort) instead of using strict from_json_str, so a version mismatch or a legacy model_observation/result_ref shape degrades to the intended model_observation/safe_summary rather than leaking the raw envelope string. safe_summary is still validated on deserialize, so genuinely non-envelope content still falls through to the raw-string branch.

id, role, content, ..
} => {
assert!(matches!(role, OpenAiResponsesMessageRole::Assistant));
// Message ids are sequence-keyed so multi-step runs stay unique.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 71eb66e — the comment now reads "Message ids are response-id-keyed (msg_{response_id})" to match the implementation.

Comment thread crates/ironclaw_reborn_composition/src/openai_compat_serve.rs
@zetyquickly

Copy link
Copy Markdown
Contributor

this is for Reborn. the same functionality for Legacy is going to be in a separate PR

- tool_result_output: best-effort envelope parse instead of strict
  from_json_str, so a version/model_observation-shape mismatch falls
  back to safe_summary rather than leaking the raw envelope string
- read_run_output: filter tool results to MessageStatus::Finalized to
  avoid surfacing redacted/deleted tool activity
- fix misleading test comment (message ids are response-id-keyed)
…-tool-outputs

# Conflicts:
#	crates/ironclaw_reborn_composition/src/openai_compat_serve.rs
The merge left two duplicate `tests` modules (inline + file-based) and
my read_run_output tests called the old single-arg reader constructor.
Combine the wait/read paths so projected run status (failed/cancelled +
error object, from main) and the full tool-call/output projection (from
this branch) both flow through, and fold the three read_run_output
tests into the file-based tests module against the 2-arg constructor.

All 10 openai_compat_serve tests pass; clippy clean with
--features openai-compat-beta.
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5010 June 17, 2026 18:23 Destroyed
@ilblackdragon
ilblackdragon merged commit 46c0262 into main Jun 17, 2026
50 of 70 checks passed
@ilblackdragon
ilblackdragon deleted the feat/openai-responses-tool-outputs branch June 17, 2026 18:44
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
* Surface tool calls and outputs in OpenAI Responses API

The non-streaming Responses API (`POST /v1/responses` create-and-wait and
`GET /v1/responses/{id}`) previously returned only the assistant's final
text message. Tool activity in the run was invisible to clients.

This projects each run's tool calls and their raw outputs into the
response `output` array: a `function_call` item (call_id / name /
arguments) paired with its `function_call_output` (the raw
`model_observation` tool result), followed by the assistant `Message`.

Implementation lives in the composition projection reader
(`OpenAiResponsesThreadProjectionReader`). Neither thread-service read
alone carries everything — the history projection keeps `turn_run_id`
plus the tool-result envelope content but strips
`tool_result_provider_call`, while the context projection preserves the
provider call but drops run attribution — so the reader joins
`list_thread_history` and `load_context_messages` by `message_id`.

Tool items are always emitted before the assistant message: a run's
single assistant draft reserves its sequence at turn start, which can
predate the tool results it produces, so transcript order alone would
mis-order them. The wait poll loop checks a cheap completion gate
(`finalized_assistant_message_by_run`) and builds the full projection
once, rather than re-reading the whole transcript every tick.

Notes:
- Only the non-streaming Responses path is covered; `stream: true` still
  emits text only (separate projection-streamer mechanism).
- This deliberately exposes the raw `model_observation` and provider tool
  name/arguments through the API surface, a documented exception to the
  crate's narrow-DTO policy.

* Address PR nearai#5010 review: harden tool-output projection

- tool_result_output: best-effort envelope parse instead of strict
  from_json_str, so a version/model_observation-shape mismatch falls
  back to safe_summary rather than leaking the raw envelope string
- read_run_output: filter tool results to MessageStatus::Finalized to
  avoid surfacing redacted/deleted tool activity
- fix misleading test comment (message ids are response-id-keyed)

* Reconcile openai_compat_serve after origin/main merge

The merge left two duplicate `tests` modules (inline + file-based) and
my read_run_output tests called the old single-arg reader constructor.
Combine the wait/read paths so projected run status (failed/cancelled +
error object, from main) and the full tool-call/output projection (from
this branch) both flow through, and fold the three read_run_output
tests into the file-based tests module against the 2-arg constructor.

All 10 openai_compat_serve tests pass; clippy clean with
--features openai-compat-beta.

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5010 — b033765b Deployed Jun 17, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants