Repository navigation
Surface tool calls and outputs in OpenAI Responses API - #5010
Conversation
The non-streaming Responses API (`POST /v1/responses` create-and-wait and
`GET /v1/responses/{id}`) previously returned only the assistant's final
text message. Tool activity in the run was invisible to clients.
This projects each run's tool calls and their raw outputs into the
response `output` array: a `function_call` item (call_id / name /
arguments) paired with its `function_call_output` (the raw
`model_observation` tool result), followed by the assistant `Message`.
Implementation lives in the composition projection reader
(`OpenAiResponsesThreadProjectionReader`). Neither thread-service read
alone carries everything — the history projection keeps `turn_run_id`
plus the tool-result envelope content but strips
`tool_result_provider_call`, while the context projection preserves the
provider call but drops run attribution — so the reader joins
`list_thread_history` and `load_context_messages` by `message_id`.
Tool items are always emitted before the assistant message: a run's
single assistant draft reserves its sequence at turn start, which can
predate the tool results it produces, so transcript order alone would
mis-order them. The wait poll loop checks a cheap completion gate
(`finalized_assistant_message_by_run`) and builds the full projection
once, rather than re-reading the whole transcript every tick.
Notes:
- Only the non-streaming Responses path is covered; `stream: true` still
emits text only (separate projection-streamer mechanism).
- This deliberately exposes the raw `model_observation` and provider tool
name/arguments through the API surface, a documented exception to the
crate's narrow-DTO policy.
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe OpenAI-compatible "Responses" projection in ChangesOpenAI Responses Output Projection Pipeline
Estimated code review effort🎯 4 (Complex) | ⏱️ ~50 minutes Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 inconclusive)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
⚔️ Resolve merge conflicts
Comment |
|
🚅 Deployed to the ironclaw-pr-5010 environment in ironclaw-ci-preview
|
There was a problem hiding this comment.
Pull request overview
Adds tool-call visibility to the non-streaming OpenAI-compatible Responses API projection by emitting function_call + function_call_output items (with raw tool outputs) ahead of the assistant message, and optimizes the wait/poll loop to avoid repeatedly re-reading the full transcript.
Changes:
- Project each run’s tool calls and raw tool outputs into the Responses
outputarray (pairedfunction_call/function_call_output), then the finalized assistant message. - Join thread history and context projections by
message_idto recover provider call metadata that history intentionally strips. - Add integration-tier tests exercising the production projection reader via
InMemorySessionThreadService.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| match ToolResultReferenceEnvelope::from_json_str(content) { | ||
| Ok(envelope) => envelope.model_observation.unwrap_or_else(|| { | ||
| serde_json::Value::String(envelope.safe_summary.as_str().to_string()) | ||
| }), | ||
| Err(_) => serde_json::Value::String(content.to_string()), | ||
| } | ||
| } |
There was a problem hiding this comment.
Fixed in 71eb66e. tool_result_output now deserializes the envelope shape directly (best-effort) instead of using strict from_json_str, so a version mismatch or a legacy model_observation/result_ref shape degrades to the intended model_observation/safe_summary rather than leaking the raw envelope string. safe_summary is still validated on deserialize, so genuinely non-envelope content still falls through to the raw-string branch.
| id, role, content, .. | ||
| } => { | ||
| assert!(matches!(role, OpenAiResponsesMessageRole::Assistant)); | ||
| // Message ids are sequence-keyed so multi-step runs stay unique. |
There was a problem hiding this comment.
Fixed in 71eb66e — the comment now reads "Message ids are response-id-keyed (msg_{response_id})" to match the implementation.
|
this is for Reborn. the same functionality for Legacy is going to be in a separate PR |
- tool_result_output: best-effort envelope parse instead of strict from_json_str, so a version/model_observation-shape mismatch falls back to safe_summary rather than leaking the raw envelope string - read_run_output: filter tool results to MessageStatus::Finalized to avoid surfacing redacted/deleted tool activity - fix misleading test comment (message ids are response-id-keyed)
…-tool-outputs # Conflicts: # crates/ironclaw_reborn_composition/src/openai_compat_serve.rs
The merge left two duplicate `tests` modules (inline + file-based) and my read_run_output tests called the old single-arg reader constructor. Combine the wait/read paths so projected run status (failed/cancelled + error object, from main) and the full tool-call/output projection (from this branch) both flow through, and fold the three read_run_output tests into the file-based tests module against the 2-arg constructor. All 10 openai_compat_serve tests pass; clippy clean with --features openai-compat-beta.
* Surface tool calls and outputs in OpenAI Responses API
The non-streaming Responses API (`POST /v1/responses` create-and-wait and
`GET /v1/responses/{id}`) previously returned only the assistant's final
text message. Tool activity in the run was invisible to clients.
This projects each run's tool calls and their raw outputs into the
response `output` array: a `function_call` item (call_id / name /
arguments) paired with its `function_call_output` (the raw
`model_observation` tool result), followed by the assistant `Message`.
Implementation lives in the composition projection reader
(`OpenAiResponsesThreadProjectionReader`). Neither thread-service read
alone carries everything — the history projection keeps `turn_run_id`
plus the tool-result envelope content but strips
`tool_result_provider_call`, while the context projection preserves the
provider call but drops run attribution — so the reader joins
`list_thread_history` and `load_context_messages` by `message_id`.
Tool items are always emitted before the assistant message: a run's
single assistant draft reserves its sequence at turn start, which can
predate the tool results it produces, so transcript order alone would
mis-order them. The wait poll loop checks a cheap completion gate
(`finalized_assistant_message_by_run`) and builds the full projection
once, rather than re-reading the whole transcript every tick.
Notes:
- Only the non-streaming Responses path is covered; `stream: true` still
emits text only (separate projection-streamer mechanism).
- This deliberately exposes the raw `model_observation` and provider tool
name/arguments through the API surface, a documented exception to the
crate's narrow-DTO policy.
* Address PR nearai#5010 review: harden tool-output projection
- tool_result_output: best-effort envelope parse instead of strict
from_json_str, so a version/model_observation-shape mismatch falls
back to safe_summary rather than leaking the raw envelope string
- read_run_output: filter tool results to MessageStatus::Finalized to
avoid surfacing redacted/deleted tool activity
- fix misleading test comment (message ids are response-id-keyed)
* Reconcile openai_compat_serve after origin/main merge
The merge left two duplicate `tests` modules (inline + file-based) and
my read_run_output tests called the old single-arg reader constructor.
Combine the wait/read paths so projected run status (failed/cancelled +
error object, from main) and the full tool-call/output projection (from
this branch) both flow through, and fold the three read_run_output
tests into the file-based tests module against the 2-arg constructor.
All 10 openai_compat_serve tests pass; clippy clean with
--features openai-compat-beta.
What
The non-streaming Responses API (
POST /v1/responsescreate-and-wait andGET /v1/responses/{id}) previously returned only the assistant's final textMessage. Tool activity in the run was invisible to clients.This projects each run's tool calls and their raw outputs into the response
outputarray:function_callitem (call_id/name/arguments), paired withfunction_call_output(the rawmodel_observationtool result),Message.How
All changes are in the composition projection reader
OpenAiResponsesThreadProjectionReader(crates/ironclaw_reborn_composition/src/openai_compat_serve.rs). The OpenAI-compat DTOs already supported these item types, so no change was needed inironclaw_reborn_openai_compat.Neither thread-service read alone carries everything:
turn_run_id(run attribution) + the tool-result envelope content (the rawmodel_observation) but stripstool_result_provider_call;load_context_messages) preserves the provider call (function name/arguments/call id) but dropsturn_run_id.So the reader joins
list_thread_historyandload_context_messagesbymessage_id. It reuses the existingload_context_messagesread (which already legitimately exposes the provider call to model-context consumers) rather than widening the guardedironclaw_threadstrait.Ordering: tool items are always emitted before the assistant message. A run's single assistant draft reserves its sequence at turn start, which can predate the tool results it later produces, so transcript order alone would mis-order them.
Polling cost:
wait_for_response_completionchecks a cheap completion gate (finalized_assistant_message_by_run) on each tick and builds the full projection once, rather than re-reading the whole transcript every 100ms.Tests
Three integration-tier tests drive the production reader through a real
InMemorySessionThreadService(per the "test through the caller" rule):function_call+function_call_outputwith rawmodel_observationoutput,cargo fmt+clippy --testsclean; module tests pass.Notes / caveats
stream: trueruns through a different projection-streamer mechanism and still emits text only.model_observationand provider tool name/arguments through the API surface — a documented exception to the crate's narrow-DTO / no-raw-provider-diagnostics policy. The sanitized safe-summary alternative would stay within policy if preferred.