Repository navigation
fix(responses-api): thread creation, GET by ID, streaming delta, context injection #2167
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
eb11d45
aa64999
d79e590
9a845bf
f3a81e1
8279990
0ac4896
5143848
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -197,10 +197,31 @@ impl SessionManager { | |
| } | ||
| } | ||
|
|
||
| // Create new thread (always create a new one for a new key) | ||
| // Create new thread (always create a new one for a new key). | ||
| // If the external_thread_id is a valid UUID AND it isn't already | ||
| // mapped to a different ThreadKey, adopt it as the internal thread ID | ||
| // so callers (e.g. the Responses API) can look up conversations by | ||
| // the same UUID they encoded in the response ID. | ||
| let thread_id = { | ||
| // Check under read lock: only adopt ext_uuid if no other key | ||
| // maps to it (prevents aliasing two keys to the same thread). | ||
| let safe_ext_uuid = if let Some(uuid) = ext_uuid { | ||
| let thread_map = self.thread_map.read().await; | ||
| if thread_map.values().any(|&v| v == uuid) { | ||
| None // Already mapped elsewhere — generate a new UUID | ||
| } else { | ||
|
Comment on lines
+208
to
+212
|
||
| Some(uuid) | ||
| } | ||
| } else { | ||
| None | ||
| }; | ||
|
|
||
| let mut sess = session.lock().await; | ||
| let thread = sess.create_thread(Some(channel)); | ||
| let thread = if let Some(uuid) = safe_ext_uuid { | ||
| sess.create_thread_with_id(uuid, Some(channel)) | ||
| } else { | ||
| sess.create_thread(Some(channel)) | ||
| }; | ||
|
Comment on lines
+200
to
+224
|
||
| thread.id | ||
| }; | ||
|
|
||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -263,6 +263,23 @@ impl Agent { | |
| }; | ||
|
|
||
| if requires_preexisting_uuid_thread(&message.channel) { | ||
| // Allow new thread creation only from the Responses API. | ||
| // Both checks are required: | ||
| // - channel == "gateway": server-set, unforgeable by WASM | ||
| // - metadata.source == "responses_api": set server-side in | ||
| // create_response_handler, not controllable by the web UI | ||
| // chat which also uses the gateway channel | ||
| let is_responses_api = message.channel == "gateway" | ||
| && message.metadata.get("source").and_then(|v| v.as_str()) | ||
| == Some("responses_api"); | ||
| if !exists && is_responses_api { | ||
| tracing::debug!( | ||
| user = %message.user_id, | ||
| thread_id = %thread_uuid, | ||
| "Allowing new thread from gateway (Responses API)" | ||
| ); | ||
| return None; | ||
| } | ||
|
Comment on lines
265
to
+282
|
||
| tracing::warn!( | ||
| user = %message.user_id, | ||
| channel = %message.channel, | ||
|
|
||
| Original file line number | Diff line number | Diff line change | ||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
@@ -64,6 +64,17 @@ pub struct ResponsesRequest { | |||||||||||||||||||||||||||||||||||||||||||||||||||||
| pub tools: Option<Vec<ResponsesTool>>, | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| #[serde(default)] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| pub tool_choice: Option<serde_json::Value>, | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// IronClaw extension: structured context injected into the agent's conversation. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// NOT part of the OpenAI Responses API spec — IronClaw extension. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// The `context` alias is kept for convenience but may collide with | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// a future OpenAI field; prefer `x_context`. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// Used by integrations to pass structured data (notification responses, | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// approval status). Should be a flat `{key: {flat_object}}` structure; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// nested objects are serialized as raw JSON. Max 10 KB. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| #[serde(default, alias = "context")] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| pub x_context: Option<serde_json::Value>, | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| } | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||
| fn default_model() -> String { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
@@ -299,6 +310,35 @@ fn make_item_id() -> String { | |||||||||||||||||||||||||||||||||||||||||||||||||||||
| format!("item_{}", Uuid::new_v4().simple()) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| } | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// Format structured context as a human-readable prefix for the agent. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| fn format_context(ctx: &serde_json::Value) -> String { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let obj = match ctx.as_object() { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Some(o) => o, | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| None => return format!("[Context: {}]", ctx), | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| }; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let mut parts = Vec::new(); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| for (key, value) in obj { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let detail = match value.as_object() { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Some(inner) => { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let fields: Vec<String> = inner | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| .iter() | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| .map(|(k, v)| { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let s = match v { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| serde_json::Value::String(s) => s.clone(), | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| other => other.to_string(), | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| }; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| format!("{k}: {s}") | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| }) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| .collect(); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| format!("[Context: {key} \u{2014} {}]", fields.join(", ")) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| } | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| None => format!("[Context: {key}: {value}]"), | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| }; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| parts.push(detail); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| } | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| parts.join("\n") | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| } | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||
| /// Extract the user message text from the input. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| fn extract_user_content(input: &ResponsesInput) -> Result<String, String> { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| match input { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
@@ -622,9 +662,24 @@ pub async fn create_response_handler( | |||||||||||||||||||||||||||||||||||||||||||||||||||||
| )); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| } | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let content = extract_user_content(&req.input) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let mut content = extract_user_content(&req.input) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| .map_err(|e| api_error(StatusCode::BAD_REQUEST, e, "invalid_request_error"))?; | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||
| // Prepend structured context (e.g. notification approval/rejection). | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| // Enforce a 10 KB size limit to prevent context window exhaustion. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| if let Some(ref ctx) = req.x_context { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let ctx_bytes = serde_json::to_string(ctx).map(|s| s.len()).unwrap_or(0); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| if ctx_bytes > 10 * 1024 { | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| return Err(api_error( | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| StatusCode::BAD_REQUEST, | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| format!("x_context exceeds 10 KB limit ({ctx_bytes} bytes)"), | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| "invalid_request_error", | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| )); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| } | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| let prefix = format_context(ctx); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
| content = format!("<user-context>\n{prefix}\n</user-context>\n\n{content}"); | ||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Comment on lines
+669
to
+680
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||
| // Enforce a 10 KB size limit to prevent context window exhaustion. | |
| if let Some(ref ctx) = req.x_context { | |
| let ctx_bytes = serde_json::to_string(ctx).map(|s| s.len()).unwrap_or(0); | |
| if ctx_bytes > 10 * 1024 { | |
| return Err(api_error( | |
| StatusCode::BAD_REQUEST, | |
| format!("x_context exceeds 10 KB limit ({ctx_bytes} bytes)"), | |
| "invalid_request_error", | |
| )); | |
| } | |
| let prefix = format_context(ctx); | |
| content = format!("<user-context>\n{prefix}\n</user-context>\n\n{content}"); | |
| // Enforce a 10 KB size limit on the exact bytes injected into the prompt | |
| // to prevent context window exhaustion. | |
| if let Some(ref ctx) = req.x_context { | |
| let prefix = format_context(ctx); | |
| let wrapped_prefix = format!("<user-context>\n{prefix}\n</user-context>\n\n"); | |
| let prefix_bytes = wrapped_prefix.len(); | |
| if prefix_bytes > 10 * 1024 { | |
| return Err(api_error( | |
| StatusCode::BAD_REQUEST, | |
| format!("x_context exceeds 10 KB limit ({prefix_bytes} bytes)"), | |
| "invalid_request_error", | |
| )); | |
| } | |
| content = format!("{wrapped_prefix}{content}"); |
Copilot
AI
Apr 14, 2026
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The new streaming behavior conditionally skips emitting response.output_text.delta when StreamChunk events already delivered content (if acc.text_chunks.is_empty()). There isn’t currently a unit/integration test asserting that no duplicate OutputTextDelta is emitted in the “chunks + terminal Response” path. Adding a test around streaming_worker (or extracting the emission decision into a testable helper) would prevent regressions where clients see duplicated text again.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The guard if acc.text_chunks.is_empty() ensures: if StreamChunk events delivered content incrementally, the final Response event does NOT re-emit the full text as a duplicate delta. If no StreamChunks arrived (non-streaming LLM), the fallback single-delta path fires. The Response event always updates acc.output[idx] with the authoritative final text regardless — only the delta emission is conditional.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The id of the ResponseOutputItem::Message is being regenerated here using make_item_id(). However, a placeholder for this item was already created with a different ID, either in the StreamChunk handler or earlier in this Response handler. This leads to an inconsistency where the id of the item in the response.output_item.added event is different from the id in the response.output_item.done event. The item ID should remain stable throughout its lifecycle.
To fix this, you should reuse the ID from the placeholder item that already exists in acc.output[idx].
let item_id = if let Some(ResponseOutputItem::Message { id, .. }) = acc.output.get(idx) {
id.clone()
} else {
// This path should not be taken if a placeholder was correctly inserted.
make_item_id()
};
let item = ResponseOutputItem::Message {
id: item_id,
role: "assistant".to_string(),
content: vec![MessageContent::OutputText { text }],
};There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Fixed in d79e590 — now reuses the placeholder ID from acc.output[idx] instead of calling make_item_id() again.
Copilot
AI
Apr 14, 2026
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The updated streaming finalization logic (placeholder output_item.added, conditional output_text.delta emission based on acc.text_chunks.is_empty(), and ID reuse for output_item.done) isn’t covered by unit tests. Since this behavior is client-visible and previously regressed (duplicate text / ID mismatch), add a focused test that drives streaming_worker (or an extracted helper) through both paths: (1) StreamChunk-delivered text (no terminal full-text delta) and (2) no StreamChunks (single full-text delta), asserting the emitted SSE event sequence and stable item IDs.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Medium — Non-blocking: Missing regression tests for 3 of 4 fixes
These two format_context tests are good, but the other three fixes lack automated regression coverage:
- Thread creation bypass (
thread_ops.rs): No test verifying gateway messages with non-existent UUIDs returnNonefrommaybe_hydrate_thread - UUID adoption (
session_manager.rs): No test exercisingcreate_thread_with_idthroughresolve_thread - Streaming delta dedup: No test verifying the
text_chunks.is_empty()guard prevents duplicate deltas
Per project rules, bug fixes should include tests that would have caught the bug. Could be a follow-up PR.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Medium — Non-blocking: TOCTOU between read lock and write lock
The read lock at line 209 checks that
uuidisn't already mapped, then drops the lock. The thread is created and the mapping inserted later under a separate write lock. Between the two, another task could map the same UUID.The existing adoption path (lines 178-198 above) handles this correctly with a double-check under the write lock (
if !thread_map.values().any(|&v| v == ext_uuid)at line 186). This new creation path doesn't follow the same pattern.Practical risk is very low since the Responses API generates
Uuid::new_v4(), but it's inconsistent with the codebase's own concurrency discipline. Consider adding a re-check under the write lock when inserting the mapping, matching the pattern at line 186.