Skip to content

[codex] Port Reborn Responses API input handling - #5347

Merged
ilblackdragon merged 7 commits into
mainfrom
codex/reborn-openai-responses
Jun 27, 2026
Merged

ilblackdragon merged 7 commits into
mainfrom
codex/reborn-openai-responses

Conversation

@ilblackdragon

Copy link
Copy Markdown
Member

Summary

  • Ports the Reborn OpenAI-compatible Responses API input handling and workflow contracts from the legacy coverage effort.
  • Adds the matching Reborn WebUI legacy Responses API browser scenario while keeping frontend/static fixes out of this PR.

Validation

  • tests/e2e/.venv/bin/python -m py_compile tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py
  • cargo test -p ironclaw_reborn_openai_compat --features openai-compat-beta --test responses_workflow_handlers_contract

Stacked after #5346.

@coderabbitai

coderabbitai Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 7ed38fa0-549f-416a-9dd0-60b844002b8c

📥 Commits

Reviewing files that changed from the base of the PR and between ed45079 and 97738c2.

📒 Files selected for processing (1)
  • tests/e2e/scenarios/test_reborn_webui_v2_smoke.py

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added optional response context support (x_context) to legacy Responses API flows, injected into generated outputs.
    • Added end-to-end coverage for legacy Responses API, including streaming, “continue and retrieve”, and context-injection scenarios.
  • Bug Fixes

    • Rejects invalid inputs early, including empty/whitespace-only text.
    • Enforces a maximum size for x_context.
    • Sanitizes injected/aliased context content and applies consistent external tool-name handling for capability resolution.
  • Tests

    • Expanded OpenAI-compat contract and E2E scenarios, including new validation and error-reporting checks.

Walkthrough

Updates OpenAI Responses compatibility for x_context, manual input-item decoding, validation, and handler/e2e coverage. Adjusts one external-tool sanitization test and one WebUI smoke assertion.

Changes

Local-dev external tool capabilities

Layer / File(s) Summary
Client tool name expectations
crates/ironclaw_reborn_composition/src/runtime/local_dev/external_tool_capability.rs
The test now uses client_tool and asserts the sanitized capability id, safe name, provider tool name, and call mapping.

OpenAI Responses compatibility

Layer / File(s) Summary
Request schema and input decoding
crates/ironclaw_reborn_openai_compat/src/responses.rs, crates/ironclaw_reborn_openai_compat/tests/dto_contract.rs
OpenAiResponsesCreateRequest adds x_context, OpenAiResponsesInputItem uses manual deserialization, and explicit message items without role are rejected.
Context validation and text rendering
crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs
x_context gets byte-limit validation, whitespace-only text is rejected, and context data is flattened into product text.
Handler contract coverage
crates/ironclaw_reborn_openai_compat/tests/responses_workflow_handlers_contract.rs
Contract tests cover x_context injection, legacy message input normalization, context alias sanitization, empty-input rejection, and oversized x_context rejection.
E2E harness and request helper
tests/e2e/helpers.py, tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py
The SSE helper accepts method/body overrides, and the Reborn v2 scenario builds the compat binary, starts the server, and creates the authenticated client.
Legacy Responses scenarios
tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py
The scenario covers non-streaming text, legacy message aliases, continuation and retrieval, streaming, context injection, auth, and invalid-input paths.

Reborn WebUI smoke test

Layer / File(s) Summary
Composer disabled state
tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
The in-flight reply scenario now checks data-send-disabled="true" on the composer element.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • nearai/ironclaw#5235 — Updates the same WebUI composer behavior exercised by the added smoke assertion.
  • nearai/ironclaw#5303 — Changes the same external-tool capability path and ProviderToolName sanitization test.
  • nearai/ironclaw#5309 — Also adjusts external-tool name normalization into capability IDs and provider-call mappings.

Suggested reviewers

  • hanakannzashi

Poem

client_tool hums in tidy code,
x_context walks the request road.
Legacy streams and inputs align,
while smoke-test buttons hold the line.
Small changes, clear paths, all set in place.

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description only includes Summary and Validation; most required template sections are missing or unfilled. Add the missing template sections: Change Type, Linked Issue, Security Impact, Trust-Boundary Checklist, Database Impact, Blast Radius, Rollback Plan, Review Follow-Through, and Review track.
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title is related to the main change, though it does not use the preferred Conventional Commits style.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5347 June 26, 2026 16:17 Destroyed
@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jun 26, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces structured context injection (x_context) to the OpenAI-compatible Responses API, including size validation, custom deserialization for input items, and comprehensive integration tests. The review feedback suggests fixing a deserialization edge case where a missing role field on a message item produces a misleading error, and recommends performance optimizations to avoid unnecessary heap allocations during context size validation and string formatting.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread crates/ironclaw_reborn_openai_compat/src/responses.rs Outdated
Comment thread crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs
Comment thread crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs
@railway-app

railway-app Bot commented Jun 26, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-5347 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jun 27, 2026 at 4:05 am

Base automatically changed from codex/reborn-runtime-tool-surface to main June 27, 2026 03:24
@ilblackdragon
ilblackdragon marked this pull request as ready for review June 27, 2026 03:25
Copilot AI review requested due to automatic review settings June 27, 2026 03:25
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_reborn_composition/src/runtime/local_dev/external_tool_capability.rs (1)

366-381: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Reject external tools that reuse an existing provider tool name.

This only checks safe_name/capability_id. If an external tool’s ProviderToolName matches an existing host/extension provider tool name, capability_id_for_tool_name() resolves the external map first and can route a model call to the external-tool parking path instead of the intended host capability.

Fail closed on provider tool-name collisions
+        let existing_provider_tool_names = self
+            .inner_tool_definitions()?
+            .into_iter()
+            .map(|definition| definition.name)
+            .collect::<Vec<_>>();
         for spec in specs {
@@
             let provider_tool_name = provider_tool_name_for_external_tool(spec.name())?;
-            let capability_id = external_tool_capability_id(provider_tool_name.as_str())?;
+            if existing_provider_tool_names
+                .iter()
+                .any(|name| name == &provider_tool_name)
+                || capability_ids_by_tool_name.contains_key(&provider_tool_name)
+            {
+                return Err(AgentLoopHostError::new(
+                    AgentLoopHostErrorKind::InvalidInvocation,
+                    "external tool conflicts with another provider tool name",
+                ));
+            }
+            let capability_id = spec.capability_id().clone();

As per coding guidelines, local capability routing must fail closed for trust/approval-sensitive surfaces.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/external_tool_capability.rs`
around lines 366 - 381, The collision check in external_tool_capability handling
only compares capability_id/safe_name, so an external tool can still reuse an
existing ProviderToolName and hijack routing. Update the conflict validation
near provider_tool_name_for_external_tool, external_tool_capability_id, and
capability_ids_by_tool_name to also reject any external ProviderToolName that
already exists in the host/extension provider-tool registry, and fail closed
with InvalidInvocation before inserting the new ToolSpec.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_agent_loop/src/families/mod.rs`:
- Around line 55-64: The `default_with_iteration_limit` helper in
`families::mod` overrides the planner budget via `DefaultPlanner::with_budget()`
but still reuses the original `id()` and `version()`, so this variant publishes
the default `LoopFamily` identity while behaving differently. Recompute or
explicitly override the `ComponentIdentity` after applying the
`DefaultBudgetStrategy`, and ensure `LoopFamily::new` is given an
identity/version that matches the iteration-limited variant instead of the
default planner’s values.

In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/external_tool_capability.rs`:
- Around line 453-457: Preserve the specific ExternalToolCatalogError kind in
catalog_error instead of mapping every case to
AgentLoopHostErrorKind::Unavailable. Update catalog_error in
external_tool_capability.rs to match on ironclaw_turns::ExternalToolCatalogError
and keep InvalidRegistration distinct from true catalog outages, while still
constructing an AgentLoopHostError with the appropriate kind and message. Use
the existing catalog_error function and AgentLoopHostError::new so the
underlying cause is not collapsed into a generic infrastructure failure.
- Around line 116-127: The capability id generation in
external_tool_capability_id is duplicating the naming logic instead of using the
catalog-owned source of truth. Update the code path that builds the external
tool capability to read the canonical capability_id already validated and stored
on ExternalToolSpec, and remove the local sanitizer-based recomputation so the
runtime stays aligned with ironclaw_turns and any future catalog changes.

In `@crates/ironclaw_reborn_openai_compat/src/responses.rs`:
- Around line 63-108: The deserialization match in responses.rs is
misclassifying typed message items because the `if object.contains_key("role")`
guard applies to both `Some("message")` and `None`, so a `type: "message"`
object without `role` falls through to the unsupported-type branch. Update the
match around `Self::deserialize` so `Some("message")` always deserializes
through `MessageWire`, letting field validation report the missing
`role`/`content` error, while keeping the `None` fallback for untyped role-based
messages.

In `@tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py`:
- Around line 157-173: The raw SSE test is using httpx.stream and manually
parsing event frames, but the e2e harness requires aiohttp for SSE streaming.
Update test_reborn_legacy_responses_streaming_raw_sse to use the shared
sse_stream() helper from helpers.py instead of response.aiter_lines() so raw
SSE/keepalive behavior matches the rest of the suite. Keep the REST client usage
on httpx.AsyncClient, and preserve the existing assertions for response.created
and response.completed.

---

Outside diff comments:
In
`@crates/ironclaw_reborn_composition/src/runtime/local_dev/external_tool_capability.rs`:
- Around line 366-381: The collision check in external_tool_capability handling
only compares capability_id/safe_name, so an external tool can still reuse an
existing ProviderToolName and hijack routing. Update the conflict validation
near provider_tool_name_for_external_tool, external_tool_capability_id, and
capability_ids_by_tool_name to also reject any external ProviderToolName that
already exists in the host/extension provider-tool registry, and fail closed
with InvalidInvocation before inserting the new ToolSpec.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 894b73f8-72a1-4e83-8903-92cd312c1d36

📥 Commits

Reviewing files that changed from the base of the PR and between c8a51ad and 0a31589.

📒 Files selected for processing (13)
  • crates/ironclaw_agent_loop/src/families/mod.rs
  • crates/ironclaw_reborn/src/app_loop_family.rs
  • crates/ironclaw_reborn/src/runtime.rs
  • crates/ironclaw_reborn_composition/src/runtime.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/external_tool_capability.rs
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/tests.rs
  • crates/ironclaw_reborn_openai_compat/src/responses.rs
  • crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs
  • crates/ironclaw_reborn_openai_compat/tests/responses_workflow_handlers_contract.rs
  • tests/e2e/helpers.py
  • tests/e2e/mock_llm.py
  • tests/e2e/reborn_webui_harness.py
  • tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py

Comment thread crates/ironclaw_agent_loop/src/families/mod.rs
Comment thread crates/ironclaw_reborn_openai_compat/src/responses.rs
Comment thread tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py Outdated
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5347 June 27, 2026 03:37 Destroyed
Copilot AI review requested due to automatic review settings June 27, 2026 03:39
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5347 June 27, 2026 03:39 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs (1)

891-896: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Return the canonical wire name context in error.param.

This validator sits on the public /api/v1/responses boundary, but it reports "x_context" when the request field is exposed as context. That leaks the Rust-only field name to clients and points callers at a parameter they never sent. Please return the canonical wire name here and cover it in the handler contract test. As per coding guidelines, At channel boundaries to users, map internal errors to sanitized messages, and the line-range change details state that this field serializes under the context alias.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs` around lines
891 - 896, The Responses request validator is returning the Rust-only parameter
name instead of the public wire name. Update the `responses_workflow` validation
path that builds `OpenAiCompatHttpError::invalid_request` for
`request.x_context` so `error.param` uses the canonical `context` name, and make
sure the corresponding handler contract test asserts the exposed parameter name
is `context` rather than `x_context`.

Source: Coding guidelines

tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py (1)

167-183: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Assert a streaming-only event, not just terminal lifecycle events.

This now passes even if /v1/responses emits only response.created and response.completed. The compat streaming contract also emits text events, so this caller-level E2E should require at least response.output_text.delta/response.output_text.done and keep the completion ordering check.

Suggested tightening
     assert events
     assert "response.created" in events
+    assert "response.output_text.delta" in events
+    assert "response.output_text.done" in events
     assert "response.completed" in events
+    assert events.index("response.created") < events.index("response.completed")
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py` around
lines 167 - 183, The E2E check in the response-stream parsing loop is too weak
because it only asserts terminal lifecycle events; tighten it in the streaming
assertion around the events collection so it also requires a streaming-only text
event such as response.output_text.delta or response.output_text.done, while
still keeping the existing response.created and response.completed ordering
checks.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_reborn_openai_compat/src/responses.rs`:
- Around line 63-76: The deserialization in responses.rs is treating a present
non-string "type" the same as an absent one, which lets malformed payloads fall
back to the legacy role-based Message path. Update the matching logic in the
OpenAiResponsesInputItem deserializer to inspect object.get("type") directly so
only a truly missing type can use the legacy branch, and reject any non-string
type value before reaching the fallback. Keep the existing MessageWire /
OpenAiResponsesInputItem::Message handling for valid "message" values, and add a
regression test covering a payload with "type": 1 to ensure it fails at the
boundary.

---

Outside diff comments:
In `@crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs`:
- Around line 891-896: The Responses request validator is returning the
Rust-only parameter name instead of the public wire name. Update the
`responses_workflow` validation path that builds
`OpenAiCompatHttpError::invalid_request` for `request.x_context` so
`error.param` uses the canonical `context` name, and make sure the corresponding
handler contract test asserts the exposed parameter name is `context` rather
than `x_context`.

In `@tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py`:
- Around line 167-183: The E2E check in the response-stream parsing loop is too
weak because it only asserts terminal lifecycle events; tighten it in the
streaming assertion around the events collection so it also requires a
streaming-only text event such as response.output_text.delta or
response.output_text.done, while still keeping the existing response.created and
response.completed ordering checks.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 162f54ba-b0b4-4173-ac0c-76e51e42cca5

📥 Commits

Reviewing files that changed from the base of the PR and between 0a31589 and 7c34c80.

📒 Files selected for processing (7)
  • crates/ironclaw_reborn_composition/src/runtime/local_dev/external_tool_capability.rs
  • crates/ironclaw_reborn_openai_compat/src/responses.rs
  • crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs
  • crates/ironclaw_reborn_openai_compat/tests/dto_contract.rs
  • crates/ironclaw_reborn_openai_compat/tests/responses_workflow_handlers_contract.rs
  • tests/e2e/helpers.py
  • tests/e2e/scenarios/test_reborn_webui_v2_legacy_responses_api.py

Comment on lines +63 to +76
match object.get("type").and_then(serde_json::Value::as_str) {
Some("message") => {
#[derive(Deserialize)]
struct MessageWire {
role: OpenAiResponsesMessageRole,
content: serde_json::Value,
}
let wire = MessageWire::deserialize(value).map_err(de::Error::custom)?;
Ok(Self::Message {
role: wire.role,
content: wire.content,
})
}
None if object.contains_key("role") => {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Reject non-string type values instead of treating them as legacy messages.

object.get("type").and_then(serde_json::Value::as_str) makes "type"-missing and "type"-present-but-not-a-string indistinguishable. A malformed payload like {"type": 1, "role": "user", "content": "hi"} will therefore hit the legacy role branch and be accepted as OpenAiResponsesInputItem::Message instead of failing at the request boundary. Match on object.get("type") directly so only an actually absent type can use the legacy fallback, and add a DTO regression test for the non-string case. As per path instructions, Fail loud at request boundaries, and as per coding guidelines, Use enums with #[serde(rename_all = "snake_case")] or explicit #[serde(rename = "...")] for fixed small sets instead of string comparisons.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_openai_compat/src/responses.rs` around lines 63 - 76,
The deserialization in responses.rs is treating a present non-string "type" the
same as an absent one, which lets malformed payloads fall back to the legacy
role-based Message path. Update the matching logic in the
OpenAiResponsesInputItem deserializer to inspect object.get("type") directly so
only a truly missing type can use the legacy branch, and reject any non-string
type value before reaching the fallback. Keep the existing MessageWire /
OpenAiResponsesInputItem::Message handling for valid "message" values, and add a
regression test covering a payload with "type": 1 to ensure it fails at the
boundary.

Sources: Coding guidelines, Path instructions

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5347 June 27, 2026 03:50 Destroyed
Copilot AI review requested due to automatic review settings June 27, 2026 03:59
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-5347 June 27, 2026 03:59 Destroyed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@ilblackdragon
ilblackdragon merged commit f0f46a5 into main Jun 27, 2026
106 checks passed
@ilblackdragon
ilblackdragon deleted the codex/reborn-openai-responses branch June 27, 2026 06:36

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-5347 — 97738c2c Deployed Jun 27, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants