Skip to content

fix(openai-compat): stop emitting [non_text_content] for non-text parts (#4644) - #4680

Merged
ilblackdragon merged 2 commits into
mainfrom
fix/4644-openai-compat-content-parts
Jun 14, 2026
Merged

ilblackdragon merged 2 commits into
mainfrom
fix/4644-openai-compat-content-parts

Conversation

@ilblackdragon

@ilblackdragon ilblackdragon commented Jun 10, 2026 •

Copy link
Copy Markdown
Member

Summary

Kills the [non_text_content] canary the issue calls out: the Chat Completions and Responses inbound paths both collapsed any non-text content part (image_url / input_audio / file) to that opaque literal, which then reached the model as user text. They were also byte-identical duplicate parsers (the duplicate-pipeline smell in architecture.md).

What changed

  • New crate-private module content_parts owns the single copy of sanitize_product_text_fragment, content_array_item_text, and a new non_text_part_marker. Both chat_workflow.rs and responses_workflow.rs delegate to it; their duplicate local copies are deleted.
  • non_text_part_marker maps a part type to a bounded, static marker ([image omitted] / [audio omitted] / [file omitted] / [unsupported content omitted]). It returns &'static str and never echoes the attacker-controlled type string, so a crafted type cannot inject transcript content — and the legacy [non_text_content] token is gone from both paths.

This route surface still cannot carry image/audio bytes to the model (the product envelope is bytes-free by design); the marker is the honest signal that a non-text part was provided but omitted. Multimodal delivery remains a separate follow-up (needs the byte-readback path).

Tests

content_parts unit tests: text-part sanitization, every non-text marker, the no-echo guarantee for crafted types, non-object items dropped. Existing chat/responses workflow tests still pass.

cargo clippy -p ironclaw_reborn_openai_compat --features openai-compat-beta --tests   # clean
cargo test  -p ironclaw_reborn_openai_compat --features openai-compat-beta            # all pass

Stacking

Based on fix/4644-model-visible-attachments (#4677).

Summary by CodeRabbit

  • Bug Fixes
    • Improved sanitization for OpenAI-compatible chat content to prevent injected line breaks and separator characters from affecting transcript formatting.
    • Standardized non-text rendering (images, audio, files, and unknown types) using bounded placeholder markers, avoiding legacy hard-coded text.
    • Enhanced handling of malformed or unexpected content parts (including bare object-shaped values) to render a safe marker instead of dropping or echoing attacker-controlled fields.

@github-actions github-actions Bot added size: M 50-199 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jun 10, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@ilblackdragon
ilblackdragon force-pushed the fix/4644-model-visible-attachments branch from 3691655 to 919ec99 Compare June 13, 2026 05:19
@ilblackdragon
ilblackdragon marked this pull request as ready for review June 13, 2026 05:21
@ilblackdragon
ilblackdragon force-pushed the fix/4644-openai-compat-content-parts branch from 383f72d to aecd2ad Compare June 13, 2026 05:21
@coderabbitai

coderabbitai Bot commented Jun 13, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: ecfd02ed-78ff-4f17-ab25-5ac88e07c322

📥 Commits

Reviewing files that changed from the base of the PR and between b68eaa3 and 4d465f8.

📒 Files selected for processing (3)
  • crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs
  • crates/ironclaw_reborn_openai_compat/src/content_parts.rs
  • crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs

📝 Walkthrough

Walkthrough

A new content_parts module centralizes OpenAI-compat content normalization: text sanitization (newline/CR/Unicode separator → space), type-keyed non-text fixed markers, and content-array item dispatch. Both chat_workflow and responses_workflow drop their inline/hardcoded equivalents and delegate to it.

Changes

Content-part normalization extraction

Layer / File(s) Summary
New content_parts module: sanitization and non-text markers
crates/ironclaw_reborn_openai_compat/src/content_parts.rs
Introduces sanitize_product_text_fragment (newline/CR/Unicode → space), non_text_part_marker (type-keyed fixed markers without echoing attacker input), and content_array_item_text (dispatch to text or marker). Unit tests enforce no legacy literal, no type-string echo, and non-object drops.
Module registration
crates/ironclaw_reborn_openai_compat/src/lib.rs
Adds mod content_parts; behind #[cfg(feature = "openai-compat-beta")].
Workflow migration
crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs, crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs
Both files import the shared helpers; responses_workflow removes its local copy; both replace hardcoded "[non_text_content]" with non_text_part_marker(None).to_string() and add object-dispatch via content_array_item_text.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

A marker was hardcoded, brittle and bare,
[non_text_content] lurked everywhere.
Now one tidy module stands guard at the gate,
Sanitizing separators, fixing the fate.
No injected newlines shall pass through unchecked —
The transcript is safe, the invariants checked. 🦀

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Title follows Conventional Commits style (type(scope): summary) and accurately describes the core fix: removing the [non_text_content] literal that was leaking to models.
Description check ✅ Passed Description covers summary, what changed, tests, and stacking; missing explicit Change Type checkbox and Security Impact section completeness, but substantive content is present.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@ilblackdragon
ilblackdragon force-pushed the fix/4644-model-visible-attachments branch from 919ec99 to 10a6165 Compare June 14, 2026 04:42
Base automatically changed from fix/4644-model-visible-attachments to main June 14, 2026 05:36
…ts (#4644)

The Chat Completions and Responses inbound paths both collapsed any non-text
content part (image_url / input_audio / file) to the opaque literal
"[non_text_content]", which then reached the model as user text - the canary
the issue calls out. They were also byte-identical duplicate parsers (the
duplicate-pipeline smell in architecture.md).

- New shared crate-private module content_parts owns the one copy of
  sanitize_product_text_fragment, content_array_item_text, and a new
  non_text_part_marker. Both chat_workflow.rs and responses_workflow.rs now
  delegate to it; their duplicate local copies are deleted.
- non_text_part_marker maps a part type to a bounded, static marker
  ([image omitted] / [audio omitted] / [file omitted] / [unsupported content
  omitted]). It returns &'static str and never echoes the attacker-controlled
  part `type` string, so a crafted type cannot inject transcript content - and
  the legacy [non_text_content] token is gone from both paths.

This route surface still cannot carry image/audio bytes to the model (the
product envelope is bytes-free by design); the marker is the honest signal that
a non-text part was provided but omitted. Multimodal delivery is a separate
follow-up.

Tests: content_parts unit tests cover text-part sanitization, every non-text
marker, the no-echo guarantee for crafted types, and non-object items dropped;
existing chat/responses workflow tests still pass.
Copilot AI review requested due to automatic review settings June 14, 2026 05:44
@ilblackdragon
ilblackdragon force-pushed the fix/4644-openai-compat-content-parts branch from aecd2ad to b68eaa3 Compare June 14, 2026 05:44
@ilblackdragon
ilblackdragon enabled auto-merge June 14, 2026 05:45

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR removes the legacy [non_text_content] literal from the OpenAI-compat Chat Completions and Responses inbound parsing paths by centralizing content-part normalization in a shared module and emitting bounded, static non-text markers instead.

Changes:

  • Introduced content_parts module to share sanitization + content-part-to-text normalization logic.
  • Replaced [non_text_content] with static markers like [image omitted], [audio omitted], [file omitted], or [unsupported content omitted].
  • Deleted duplicate per-workflow implementations and routed both workflows through the shared helpers.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.

File Description
crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs Switches inbound content parsing to shared helpers and removes the legacy placeholder.
crates/ironclaw_reborn_openai_compat/src/lib.rs Adds the new content_parts module behind the openai-compat-beta feature gate.
crates/ironclaw_reborn_openai_compat/src/content_parts.rs New shared parsing + sanitization module with unit tests and bounded non-text markers.
crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs Switches inbound content parsing to shared helpers and removes the legacy placeholder.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

.collect::<Vec<_>>()
.join(" "),
Some(value) if !value.is_null() => "[non_text_content]".to_string(),
Some(value) if !value.is_null() => non_text_part_marker(None).to_string(),

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 4d465f869. A bare object-form top-level content now routes through the same per-part normalizer (content_array_item_text) as array items, so a typed object like {"type":"image_url",...} renders its specific marker ([image omitted]) — or, for a text-typed object, its text — instead of discarding the type for the generic marker. The type is still never echoed (the marker is a fixed &'static str).

.collect::<Vec<_>>()
.join(" "),
value if !value.is_null() => "[non_text_content]".to_string(),
value if !value.is_null() => non_text_part_marker(None).to_string(),

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 4d465f869. A bare object-form top-level content now routes through the same per-part normalizer (content_array_item_text) as array items, so a typed object like {"type":"image_url",...} renders its specific marker ([image omitted]) — or, for a text-typed object, its text — instead of discarding the type for the generic marker. The type is still never echoed (the marker is a fixed &'static str).

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs (1)

6-7: ⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Update stale module header after the shared-module migration.

Line 6-7 says mirroring continues “until” a shared normalization module exists, but this file now imports and uses crate::content_parts (Line 13-15). Please update the header to match current behavior.

As per coding guidelines, “When changing behavior in a function, re-read its docstring and adjacent comments; update or delete them in the same change to keep documentation in sync with code.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs` around lines
6 - 7, The module header comment at lines 6-7 contains outdated documentation.
It states that the ack and text helpers mirror the chat slice "until" a shared
normalization module exists, but the code now imports and uses
crate::content_parts (visible in lines 13-15). Update the header comment to
accurately reflect that the code is currently using the shared
crate::content_parts module instead of implying that mirroring continues until
such a module is created.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/ironclaw_reborn_openai_compat/src/content_parts.rs`:
- Around line 37-41: The match arm for text types ("text" | "input_text" |
"output_text") in content_array_item_text returns None when the "text" field is
missing or not a string, causing these malformed parts to be silently dropped by
downstream filter_map. Instead, return a marker indicating unsupported content
(using non_text_part_marker) to flag the issue loudly. Additionally, add a
regression test using #[test] that verifies malformed text parts where "text" is
missing or non-string are marked with "[unsupported content omitted]" rather
than silently filtered out.

---

Outside diff comments:
In `@crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs`:
- Around line 6-7: The module header comment at lines 6-7 contains outdated
documentation. It states that the ack and text helpers mirror the chat slice
"until" a shared normalization module exists, but the code now imports and uses
crate::content_parts (visible in lines 13-15). Update the header comment to
accurately reflect that the code is currently using the shared
crate::content_parts module instead of implying that mirroring continues until
such a module is created.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 6eb7fb80-f318-41f0-9585-141075eded42

📥 Commits

Reviewing files that changed from the base of the PR and between 9f22a08 and b68eaa3.

📒 Files selected for processing (4)
  • crates/ironclaw_reborn_openai_compat/src/chat_workflow.rs
  • crates/ironclaw_reborn_openai_compat/src/content_parts.rs
  • crates/ironclaw_reborn_openai_compat/src/lib.rs
  • crates/ironclaw_reborn_openai_compat/src/responses_workflow.rs

Comment thread crates/ironclaw_reborn_openai_compat/src/content_parts.rs Outdated
…ping them

Address review on the shared content-part normalizer:

- A text-typed array part (`text`/`input_text`/`output_text`) whose `text` is
  missing or non-string was returning None and getting silently dropped by the
  downstream filter_map. It now emits a bounded marker so the part is observable
  rather than vanishing (and the doc's "None only when not an object" holds).
- Bare object-form top-level `content` (non-standard but tolerated) was routed
  to the generic marker, discarding any `type`. It now runs through the same
  per-part logic, so `{"type":"image_url"}` renders `[image omitted]` etc.

Regression test for the malformed-text-part case.
@ilblackdragon
ilblackdragon disabled auto-merge June 14, 2026 06:39
@ilblackdragon
ilblackdragon merged commit 2392420 into main Jun 14, 2026
68 checks passed
@ilblackdragon
ilblackdragon deleted the fix/4644-openai-compat-content-parts branch June 14, 2026 06:39
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…ts (nearai#4644) (nearai#4680)

* fix(openai-compat): stop emitting [non_text_content] for non-text parts (nearai#4644)

The Chat Completions and Responses inbound paths both collapsed any non-text
content part (image_url / input_audio / file) to the opaque literal
"[non_text_content]", which then reached the model as user text - the canary
the issue calls out. They were also byte-identical duplicate parsers (the
duplicate-pipeline smell in architecture.md).

- New shared crate-private module content_parts owns the one copy of
  sanitize_product_text_fragment, content_array_item_text, and a new
  non_text_part_marker. Both chat_workflow.rs and responses_workflow.rs now
  delegate to it; their duplicate local copies are deleted.
- non_text_part_marker maps a part type to a bounded, static marker
  ([image omitted] / [audio omitted] / [file omitted] / [unsupported content
  omitted]). It returns &'static str and never echoes the attacker-controlled
  part `type` string, so a crafted type cannot inject transcript content - and
  the legacy [non_text_content] token is gone from both paths.

This route surface still cannot carry image/audio bytes to the model (the
product envelope is bytes-free by design); the marker is the honest signal that
a non-text part was provided but omitted. Multimodal delivery is a separate
follow-up.

Tests: content_parts unit tests cover text-part sanitization, every non-text
marker, the no-echo guarantee for crafted types, and non-object items dropped;
existing chat/responses workflow tests still pass.

* fix(openai-compat): render typed/object content parts instead of dropping them

Address review on the shared content-part normalizer:

- A text-typed array part (`text`/`input_text`/`output_text`) whose `text` is
  missing or non-string was returning None and getting silently dropped by the
  downstream filter_map. It now emits a bounded marker so the part is observable
  rather than vanishing (and the doc's "None only when not an object" holds).
- Bare object-form top-level `content` (non-standard but tolerated) was routed
  to the generic marker, discarding any `type`. It now runs through the same
  per-part logic, so `{"type":"image_url"}` renders `[image omitted]` etc.

Regression test for the malformed-text-part case.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants