Skip to content

fix(llm): default missing OpenAI image detail to auto - #1940

Merged
ilblackdragon merged 2 commits into
nearai:stagingfrom
protocol-AMG:fix/openai-image-detail-auto
Apr 19, 2026
Merged

ilblackdragon merged 2 commits into
nearai:stagingfrom
protocol-AMG:fix/openai-image-detail-auto

Conversation

@protocol-AMG

Copy link
Copy Markdown
Contributor

Summary

  • fix OpenAI-compatible image conversion so missing image_url.detail is normalized to "auto" instead of failing before request send
  • patch the RigAdapter boundary used by the generic OpenAI-compatible provider, which was passing detail: None into rig-core for data:image/... inputs
  • preserve explicit low and high detail values and keep missing image_url.url rejected
  • align sibling OpenAI-style serializers and add regression tests for the image detail cases
  • update FEATURE_PARITY.md note for OpenAI-compatible provider behavior

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

None

Validation

  • cargo fmt
  • cargo clippy --all --benches --tests --examples --all-features
  • Relevant tests pass: cargo test --manifest-path /Users/1954d/projects/IronClaw/Cargo.toml image_detail, cargo test --manifest-path /Users/1954d/projects/IronClaw/Cargo.toml without_detail, cargo test --manifest-path /Users/1954d/projects/IronClaw/Cargo.toml without_url
  • Manual testing: traced IncomingAttachment -> ContentPart::ImageUrl -> RigAdapter/OpenAI-compatible payload and verified missing image detail now defaults to auto

Security Impact

None

Database Impact

None

Blast Radius

Touches OpenAI-style image message conversion in RigAdapter, nearai_chat, and github_copilot. Main risk would be provider-specific image payload regressions, covered by focused conversion tests.

Rollback Plan

Revert commit a6a2c63c to restore previous behavior.


Review track: B

@github-actions github-actions Bot added scope: llm LLM integration scope: docs Documentation size: L 200-499 changed lines risk: low Changes to docs, tests, or low-risk modules contributor: new First-time contributor labels Apr 3, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces standardized normalization for OpenAI-style image input 'detail' fields across multiple providers, including GitHub Copilot, NearAI, and the Rig adapter, ensuring they default to 'auto' when missing or invalid. The changes include new utility functions for normalization, updated message conversion logic, and comprehensive unit tests. Review feedback suggests adopting a more functional, iterator-based approach in the NearAI implementation to avoid in-place mutation and optimize for WASM environments. Additionally, it is recommended to explicitly use ImageDetail::Auto instead of unwrap_or_default() in the Rig adapter to ensure robustness against future changes in the underlying library.

Comment thread src/llm/nearai_chat.rs
Comment on lines +992 to +998
parts.extend(msg.content_parts.into_iter().map(|part| match part {
crate::llm::ContentPart::ImageUrl { mut image_url } => {
image_url.detail = Some(image_url.normalized_openai_detail());
crate::llm::ContentPart::ImageUrl { image_url }
}
other => other,
}));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current implementation of From<ChatMessage> for ChatCompletionMessage performs an in-place mutation of image_url.detail. To maintain consistency with other providers in this PR and optimize for WASM by avoiding unnecessary heap allocations, consider using a functional approach with iterators for the transformation. This avoids side effects and aligns with repository guidelines on performance and consistency.

References
  1. To improve performance in WASM, avoid unnecessary heap allocations by using iterators directly instead of collecting them into a Vec.
  2. Prioritize consistency with existing code patterns over refactoring, especially when adding new components like providers.

Comment thread src/llm/rig_adapter.rs
Comment on lines +312 to +313
ImageDetail::from_str(&image_url.normalized_openai_detail())
.unwrap_or_default();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using unwrap_or_default() on ImageDetail::from_str is safe but potentially fragile. If the rig crate's ImageDetail enum changes its default variant in a future version, it could lead to silent behavior changes. Using unwrap_or(ImageDetail::Auto) is more robust and follows the principle of avoiding implicit defaults for critical fields to prevent silent failures.

Suggested change
ImageDetail::from_str(&image_url.normalized_openai_detail())
.unwrap_or_default();
ImageDetail::from_str(&image_url.normalized_openai_detail())
.unwrap_or(ImageDetail::Auto);
References
  1. For critical fields, avoid using default values to prevent silent failures; it is safer for logic to be explicit or fail if data is missing.

@serrrfirat serrrfirat left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings:

  1. The fix is incomplete: the ChatGPT Responses-provider path still sends image inputs without a detail field. In src/llm/codex_chatgpt.rs, input_image is serialized with only image_url, so image-bearing requests through CodexChatGptProvider still bypass the new normalization added elsewhere. That path is reachable from the normal attachment flow in src/agent/session.rs, which builds image turns with user_with_parts(..., turn.image_content_parts.clone()). If the bug here is “missing OpenAI image detail breaks requests”, this provider remains exposed.

  2. The added coverage misses the live attachment message shape and would not catch the gap above. The existing codex_chatgpt test builds a synthetic message with embedded ContentPart::Text and no msg.content, while the runtime path uses msg.content plus image-only content_parts. A regression test for that real shape is still missing.

Residual risks:

  • Validation is still mostly serializer-level unit coverage; there is no end-to-end/provider-request test for the actual attachment flow across the affected providers.
  • I did not complete a local cargo test run during review, so this review is source-based rather than backed by a finished test pass.

@ilblackdragon

Copy link
Copy Markdown
Member

Reviewed. Clean, correct, and well-tested. Approving with one minor architectural observation.

What it does right

  • normalize_openai_image_detail() in src/llm/provider.rs is the single source of truth: lowercase normalization, trim, allowlist gate (auto/low/high), fall back to "auto" otherwise. Centralizing this is exactly the right call.
  • The actual rig-core bug is confirmed: rig-0.30.0/src/providers/openai/completion/mod.rs:436-438 returns MessageError::ConversionError("OpenAI image URI must have image detail") when detail is None for a base64 data: URL. Previously attachments.rs emitted ImageUrl { detail: None } for every inline image, so every base64 image through RigAdapter was failing conversion before request send. Setting detail: Some(ImageDetail::Auto) for both the URL and base64 branches fixes the real failure.
  • Default is "auto" (not "high") — correct cost posture.
  • Three sibling serializers (rig_adapter, nearai_chat, github_copilot) each updated in the right way (mutating the ImageUrl before reuse vs. constructing a provider-specific struct). codex_chatgpt.rs and openai_codex_provider.rs correctly untouched (Responses API doesn't take detail, and the latter drops images anyway).
  • Regression tests at each layer (helper-level + serializer-level) plus a deserialization guard for missing url. Symmetric coverage of None, empty, unknown, mixed-case, and explicit low/high.

Minor

  • src/llm/rig_adapter.rs:312 uses ImageDetail::from_str(&image_url.normalized_openai_detail()).unwrap_or_default(). normalized_openai_detail() already guarantees one of three values that FromStr for ImageDetail accepts, so the unwrap_or_default() branch is structurally unreachable. Not a .unwrap()/.expect() violation and it correctly falls back to Auto, but a .map_err(…)? or a direct match would be slightly more honest about the contract. Non-blocking.
  • normalize_openai_image_detail silently maps unknown values (e.g. "ultra") to "auto". Per the CLAUDE.md "LLM data is never deleted" principle this is fine (we're normalizing outbound request shape, not discarding user data), but a tracing::debug! on the unknown-value branch would make future provider-API changes easier to notice. Non-blocking.

No Critical/High/Medium issues. LGTM.

@ilblackdragon
ilblackdragon merged commit 3d51423 into nearai:staging Apr 19, 2026
17 checks passed
ilblackdragon added a commit that referenced this pull request Apr 19, 2026
Resolve conflicts in two files:

- src/bridge/effect_adapter.rs: keep both HEAD's engine_store /
  skill_registry fields (for v1→v2 skill sync) and staging's new
  workspace_mounts field (per-project sandbox from #2211). Merge
  the ironclaw_engine import list accordingly.
- src/llm/rig_adapter.rs: adopt staging's detail-normalized image
  handling (#1940) — the upstream `ImageDetail::from_str(...
  normalized_openai_detail())` helper already defaults missing
  values to auto, so drop the HEAD-only parse_image_detail helper
  and its two call sites. Keep staging's expanded test coverage
  (data-url auto + explicit low/high preservation).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This was referenced Apr 19, 2026
This was referenced Apr 22, 2026
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
Co-authored-by: Edward Ji <26658037+edwardji@users.noreply.github.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: new First-time contributor risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation scope: llm LLM integration size: L 200-499 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants