Skip to content

feat(models): add gpt-image-2 from openai - #2059

Merged
smakosh merged 8 commits into
mainfrom
feat/gpt-image-2
Apr 23, 2026
Merged

smakosh merged 8 commits into
mainfrom
feat/gpt-image-2

Conversation

@smakosh

@smakosh smakosh commented Apr 21, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Adds `gpt-image-2` to the OpenAI model definitions with image generation routing enabled (`imageGenerations: true`).
  • Token-based pricing per OpenAI's model page: text $5/1M input, $10/1M output, cached text $1.25/1M, image input $8/1M, image output $30/1M.
  • 128K context, 4K max output, vision enabled, no streaming / tools / JSON output.

Test plan

  • `pnpm --filter @llmgateway/models build` passes
  • `pnpm --filter @llmgateway/models lint` passes
  • Model appears in the playground image generation picker and routes through the OpenAI image endpoint

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Support for a new image-generation model and uploads, including single or multiple input images.
    • Automatic mapping of image size/aspect presets and improved image request handling.
  • Bug Fixes

    • OpenAI-style image responses now display correctly.
    • Gateway routes image requests to the correct image endpoint.
    • Playground URL sync refined to avoid redundant navigation and flicker.

Add GPT Image 2 with token-based pricing:
text $5/$10 per 1M in/out, cached text $1.25/1M, image input $8/1M,
image output $30/1M. Routes through the image generations pipeline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Apr 21, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Adds OpenAI image-generation support: new gpt-image-2 model entry, prepareRequestBody multipart/FormData and image payload handling, endpoint selection for image generations/edits, gateway request/response rewrites and parsing for OpenAI image outputs, ProviderContext type widening, and playground URL-sync fixes.

Changes

Cohort / File(s) Summary
Model config
packages/models/src/models/openai.ts
Added gpt-image-2 model with metadata, image capability, pricing, limits, and feature flags.
Request preparation
packages/actions/src/prepare-request-body.ts
Added OpenAI image-generation path: extracts prompt and image URLs from last user message, normalizes image_config to OpenAI size, returns either OpenAIImageRequest JSON or multipart FormData (fetches/decodes image URLs into Blobs). Return type widened to `ProviderRequestBody
Endpoint routing
packages/actions/src/get-provider-endpoint.ts
For openai, short-circuits to /v1/images/generations when model/provider imageGenerations is enabled.
Gateway request routing
apps/gateway/src/chat/chat.ts
Upstream prepared requestBody can be FormData; avoids max_tokens validation for multipart; passes FormData through without JSON-stringifying and conditionally sets Content-Type; rewrites /v1/images/generations → /v1/images/edits for edit flows when multipart.
Provider response parsing
apps/gateway/src/chat/tools/parse-provider-response.ts
Adds OpenAI image response handling: maps json.data image payloads (b64_json/url) into images entries, sets response content to image label, forces finishReason: "stop", populates token usage (with fallbacks), and exits early for these cases.
Response transformation
apps/gateway/src/chat/tools/transform-response-to-openai.ts
New branch to detect OpenAI image-generation payloads (data array, no choices/output) and transform them into a chat.completion shape with choices[0].message, usage, and metadata, bypassing responses-endpoint logic.
Provider context type
apps/gateway/src/chat/tools/resolve-provider-context.ts
Widened ProviderContext.requestBody from ProviderRequestBody to `ProviderRequestBody
Playground URL sync
apps/playground/src/components/playground/image-page-client.tsx, apps/playground/src/components/playground/video-page-client.tsx
URL-sync effects now derive query params from window.location.search, compute nextUrl with pathname, call router.replace only when URL differs, use { scroll: false }, and adjust effect dependencies.

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant Gateway
  participant Actions as "Actions\n(prepare-request-body / get-provider-endpoint)"
  participant OpenAI as "OpenAI API"
  participant Parser as "Gateway Parser\n(parse-provider-response)"

  Client->>Gateway: POST image generation request
  Gateway->>Actions: prepare-request-body (extract prompt, images, image_config)
  Actions-->>Gateway: requestBody (JSON or FormData with images)
  Gateway->>Actions: get-provider-endpoint (resolve endpoint)
  Actions-->>Gateway: /v1/images/generations
  Gateway->>Gateway: if multipart and edit flow -> rewrite to /v1/images/edits
  Gateway->>OpenAI: forward request (JSON or multipart)
  OpenAI-->>Gateway: response (json.data with b64_json or url)
  Gateway->>Parser: parse-provider-response (map images, tokens, finishReason)
  Parser-->>Gateway: normalized response with images
  Gateway-->>Client: return normalized response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 75.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title accurately describes the primary change: adding the gpt-image-2 model from OpenAI. While the changeset extends beyond just model definitions to include routing, transformation, and UI fixes, the title correctly captures the main stated objective and the most visible addition to the repository.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/gpt-image-2

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@packages/models/src/models/openai.ts`:
- Around line 1825-1835: The model config currently enables imageGenerations:
true while leaving imageOutputTokensByResolution unset and cachedInputPrice
incorrect; update the model object (the block containing cachedInputPrice,
imageInputPrice, imageOutputPrice, imageGenerations) to (1) set cachedInputPrice
to OpenAI's $2/M equivalent (2 / 1e6) instead of 1.25 / 1e6, (2) add a complete
imageOutputTokensByResolution mapping keyed by quality/size (populate token
estimates for low/medium/high quality and small/medium/large sizes consistent
with OpenAI guidance to avoid falling back to LEGACY_DEFAULT_TOKENS_PER_IMAGE),
and (3) ensure any gateway cost logic that reads imageOutputTokensByResolution
(references: LEGACY_DEFAULT_TOKENS_PER_IMAGE and imageOutputTokensByResolution)
will find the new map before allowing imageGenerations: true to be exposed in
production.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: a1b32d2d-6a86-460f-9127-039f26935a38

📥 Commits

Reviewing files that changed from the base of the PR and between f09953e and 338a87c.

📒 Files selected for processing (1)
  • packages/models/src/models/openai.ts

Comment on lines +1825 to +1835
cachedInputPrice: 1.25 / 1e6,
imageInputPrice: 8 / 1e6,
imageOutputPrice: 30 / 1e6,
requestPrice: 0,
contextSize: 128000,
maxOutput: 4096,
streaming: false,
vision: true,
tools: false,
jsonOutput: false,
imageGenerations: true,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Description: Inspect image-generation cost paths for GPT Image 2-specific
# size/quality token handling or actual API usage-token accounting.
# Expected: cost code should not rely only on LEGACY_DEFAULT_TOKENS_PER_IMAGE
# for gpt-image-2, and should account for separate cached image input pricing.

rg -nP -C3 '\b(gpt-image-2|imageOutputTokensByResolution|LEGACY_DEFAULT_TOKENS_PER_IMAGE|imageOutputPrice|imageInputPrice|cachedInputPrice|maxOutput|quality|imageSize|usage)\b' \
  --glob '!**/node_modules/**' \
  --glob '!**/dist/**' \
  --glob '!**/build/**'

Repository: theopenco/llmgateway

Length of output: 50377


🏁 Script executed:

find . -name "costs.ts" -o -name "cost.ts" | grep -E '(apps|packages)' | head -20

Repository: theopenco/llmgateway

Length of output: 95


🏁 Script executed:

fd -H "costs.ts|cost.ts" --type f

Repository: theopenco/llmgateway

Length of output: 93


🏁 Script executed:

rg -l "LEGACY_DEFAULT_TOKENS_PER_IMAGE|imageOutputTokensByResolution" --type ts --type tsx --max-count=20

Repository: theopenco/llmgateway

Length of output: 91


🏁 Script executed:

sed -n '1810,1838p' packages/models/src/models/openai.ts

Repository: theopenco/llmgateway

Length of output: 758


🏁 Script executed:

cat -n apps/gateway/src/lib/costs.ts

Repository: theopenco/llmgateway

Length of output: 17022


🏁 Script executed:

rg -n "imageGenerations|imageOutputTokensByResolution|imageOutputPrice" apps/gateway/src/lib/costs.ts -A 5 -B 5

Repository: theopenco/llmgateway

Length of output: 1380


Add GPT Image 2 quality/size-dependent output token handling before enabling imageGenerations: true.

The cost calculation in apps/gateway/src/lib/costs.ts (lines 363–371) resolves imageOutputTokensByResolution for per-resolution token counts. When absent, it falls back to LEGACY_DEFAULT_TOKENS_PER_IMAGE = 1120 tokens per image. GPT Image 2 output tokens vary significantly by quality and size parameters—OpenAI's guide documents costs ranging from ~$0.006 (low quality, small size) to ~$0.211 (high quality, large size), which corresponds to roughly 200–7000 output tokens at the configured $30/M price. The static 1120-token fallback will systematically misbill these requests.

Additionally, the cachedInputPrice: 1.25 / 1e6 in the model definition does not align with OpenAI's separate $2/M cached image input pricing, and there is no corresponding cached image input cost handler in the cost path.

Populate imageOutputTokensByResolution with quality/size-specific token values (from OpenAI's documentation or API usage data) and ensure cached image input pricing is correctly represented before marking this model selectable for production traffic.

Sources: OpenAI pricing, OpenAI image generation guide.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/models/src/models/openai.ts` around lines 1825 - 1835, The model
config currently enables imageGenerations: true while leaving
imageOutputTokensByResolution unset and cachedInputPrice incorrect; update the
model object (the block containing cachedInputPrice, imageInputPrice,
imageOutputPrice, imageGenerations) to (1) set cachedInputPrice to OpenAI's $2/M
equivalent (2 / 1e6) instead of 1.25 / 1e6, (2) add a complete
imageOutputTokensByResolution mapping keyed by quality/size (populate token
estimates for low/medium/high quality and small/medium/large sizes consistent
with OpenAI guidance to avoid falling back to LEGACY_DEFAULT_TOKENS_PER_IMAGE),
and (3) ensure any gateway cost logic that reads imageOutputTokensByResolution
(references: LEGACY_DEFAULT_TOKENS_PER_IMAGE and imageOutputTokensByResolution)
will find the new map before allowing imageGenerations: true to be exposed in
production.

@smakosh smakosh self-assigned this Apr 21, 2026
smakosh added 2 commits April 22, 2026 23:59
- Route openai imageGenerations to /v1/images/generations
- Transform chat-shape payload to images API shape and normalize
  "1K"/aspect-ratio into 1024x1024/1024x1536/1536x1024
- Parse { data:[{b64_json|url}] } responses back into images
- Swap to /v1/images/edits when an input image is present

Also fix infinite /orgs refetch loop: image/video page URL-sync
effects depended on searchParams, causing router.replace to
produce a new ref that retriggered the effect and refetched the
RSC forever. Match chat-page-client pattern (read
window.location.search, compare before replace, scroll:false).

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 3900-3908: In the OpenAI image-edit branch inside
apps/gateway/src/chat/chat.ts (the block that rewrites url from
"/v1/images/generations" to "/v1/images/edits"), change the payload presence
check to look for either "image" or "images" (e.g., ("image" in requestBody ||
"images" in requestBody)) so edits are selected for both singular and plural
image fields; also consider updating the prepareRequestBody logic that builds
OpenAI image payloads to emit an "images" array for multi-image requests to
match xAI's pattern and OpenAI's API expectations.

In `@apps/gateway/src/chat/tools/parse-provider-response.ts`:
- Around line 538-550: The mapping over imageData can produce undefined urls
because only json.data[0] was checked; update the images construction so each
item is validated: for each item in imageData check item.b64_json and item.url,
build url = `data:image/png;base64,...` when b64_json exists, otherwise use
item.url only if defined, and skip (or log and continue) any item lacking both
values instead of returning an ImageObject with an undefined url; adjust the
images assignment that produces ImageObject entries to filter out invalid items
and ensure returned objects always have a defined image_url.url.

In `@packages/actions/src/prepare-request-body.ts`:
- Around line 646-659: The code returns an object named openaiImageRequest typed
as any containing model/prompt/size/n and image fields, but when routing to the
OpenAI /v1/images/edits endpoint you must send multipart/form-data with binary
files rather than JSON; change the return type from any to an explicit interface
(e.g., ImageRequest { model: string; prompt?: string; size?: string; n?: number;
image?: string | string[] }) and detect the edits flow (the URL swap in chat.ts)
to build and return a FormData instead of a JSON object: decode data URLs in
imageUrls into binary Blobs/Uint8Arrays and append each as image[] with
filenames, plus append model/prompt/size/n as form fields, otherwise continue
returning the typed JSON object; ensure no use of any.
- Around line 610-644: The code in prepare-request-body.ts incorrectly restricts
allowed image sizes by mapping any non-whitelisted rawSize to "auto" (affecting
the openaiSize assigned from image_config?.image_size) and so breaks valid
gpt-image-2 dimensions; change the logic in the block that computes openaiSize
to accept numeric or arbitrary "WIDTHxHEIGHT" strings that meet OpenAI
constraints (multiples of 16, max edge ≤3840, aspect ratio ≤3:1, pixel count
between 655,360 and 8,294,400) instead of forcing "auto", and preserve the
existing aspectRatio-based fallback behavior for presets; also replace the
untyped openaiImageRequest (currently declared as any) with a proper
interface/type (e.g., OpenAIImageRequest) or a typed cast reflecting the OpenAI
image generation request shape so the request object is strongly typed rather
than any.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: bc5c6909-34e4-4867-862b-65412e52c0a4

📥 Commits

Reviewing files that changed from the base of the PR and between 338a87c and a186396.

📒 Files selected for processing (6)
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/chat/tools/parse-provider-response.ts
  • apps/playground/src/components/playground/image-page-client.tsx
  • apps/playground/src/components/playground/video-page-client.tsx
  • packages/actions/src/get-provider-endpoint.ts
  • packages/actions/src/prepare-request-body.ts

Comment thread apps/gateway/src/chat/chat.ts Outdated
Comment on lines +538 to +550
const imageData = json.data;
images = imageData.map((item: any): ImageObject => {
let url: string;
if (item.b64_json) {
url = `data:image/png;base64,${item.b64_json}`;
} else {
url = item.url;
}
return {
type: "image_url",
image_url: { url },
};
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Per-item fallback may produce an undefined URL.

The outer guard on Line 536 only checks json.data[0], but the map iterates every item. If a later element has neither b64_json nor url, the else branch assigns url = item.url (undefined), producing a malformed ImageObject. Consider guarding per item and skipping/erroring on invalid entries.

🛡️ Suggested hardening
-				images = imageData.map((item: any): ImageObject => {
-					let url: string;
-					if (item.b64_json) {
-						url = `data:image/png;base64,${item.b64_json}`;
-					} else {
-						url = item.url;
-					}
-					return {
-						type: "image_url",
-						image_url: { url },
-					};
-				});
+				images = imageData
+					.filter((item: any) => item?.b64_json || item?.url)
+					.map((item: any): ImageObject => ({
+						type: "image_url",
+						image_url: {
+							url: item.b64_json
+								? `data:image/png;base64,${item.b64_json}`
+								: item.url,
+						},
+					}));
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
const imageData = json.data;
images = imageData.map((item: any): ImageObject => {
let url: string;
if (item.b64_json) {
url = `data:image/png;base64,${item.b64_json}`;
} else {
url = item.url;
}
return {
type: "image_url",
image_url: { url },
};
});
const imageData = json.data;
images = imageData
.filter((item: any) => item?.b64_json || item?.url)
.map((item: any): ImageObject => ({
type: "image_url",
image_url: {
url: item.b64_json
? `data:image/png;base64,${item.b64_json}`
: item.url,
},
}));
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/tools/parse-provider-response.ts` around lines 538 -
550, The mapping over imageData can produce undefined urls because only
json.data[0] was checked; update the images construction so each item is
validated: for each item in imageData check item.b64_json and item.url, build
url = `data:image/png;base64,...` when b64_json exists, otherwise use item.url
only if defined, and skip (or log and continue) any item lacking both values
instead of returning an ImageObject with an undefined url; adjust the images
assignment that produces ImageObject entries to filter out invalid items and
ensure returned objects always have a defined image_url.url.

Comment thread packages/actions/src/prepare-request-body.ts
Comment thread packages/actions/src/prepare-request-body.ts Outdated
smakosh and others added 3 commits April 23, 2026 22:23
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
apps/gateway/src/chat/tools/resolve-provider-context.ts (2)

396-411: ⚠️ Potential issue | 🔴 Critical

Missing FormData type guard before hasMaxTokens.

requestBody is now ProviderRequestBody | FormData, but hasMaxTokens is typed to accept ProviderRequestBody only (packages/models/src/types.ts). Passing a potential FormData here is both a TS type error and semantically wrong — "max_tokens" in formData is always false, but it shouldn't be calling this branch at all for image-edit multipart bodies.

The sibling callsite in apps/gateway/src/chat/chat.ts (3866-3870) already uses the correct pattern; mirror it here:

🩹 Proposed fix
 	// Post-validation of max_tokens in request body
 	if (
-		hasMaxTokens(requestBody) &&
+		!(requestBody instanceof FormData) &&
+		hasMaxTokens(requestBody) &&
 		requestBody.max_tokens !== undefined &&
 		providerMappingForSelected
 	) {
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/tools/resolve-provider-context.ts` around lines 396 -
411, The current block calls hasMaxTokens(requestBody) but requestBody can be
ProviderRequestBody | FormData; add a FormData type guard before calling
hasMaxTokens so multipart image-edit FormData bodies are skipped (mirror the
pattern used in chat.ts). Specifically, check that requestBody is not an
instance of FormData (or otherwise not FormData) before invoking
hasMaxTokens(requestBody), then proceed with the existing checks against
providerMappingForSelected.maxOutput and usedModel to throw the HTTPException
when requestBody.max_tokens exceeds providerMappingForSelected.maxOutput.

413-425: ⚠️ Potential issue | 🔴 Critical

Fix unconditional Content-Type: application/json header to allow FormData multipart boundaries.

When requestBody is FormData, the unconditional headers["Content-Type"] = "application/json" at line 418 blocks fetch from auto-generating the required multipart/form-data; boundary=… header, causing OpenAI's /v1/images/edits and similar endpoints to reject with 400. The same pattern is already correctly implemented elsewhere in the codebase (see chat.ts:7827–7829).

Conditionally set the header only when the body is not multipart:

const headers = getProviderHeaders(usedProvider as Provider, usedToken, {
    requestId: options.requestId,
    webSearchEnabled: options.webSearchEnabled,
});
-headers["Content-Type"] = "application/json";
+if (!(requestBody instanceof FormData)) {
+    headers["Content-Type"] = "application/json";
+}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@apps/gateway/src/chat/tools/resolve-provider-context.ts` around lines 413 -
425, The code unconditionally sets headers["Content-Type"] = "application/json"
which prevents fetch from auto-generating multipart boundaries for FormData;
instead, in resolve-provider-context.ts where headers are built via
getProviderHeaders and assigned to the headers variable, only set the
Content-Type when the outgoing requestBody is not multipart (e.g., guard with a
check like "requestBody is not FormData" or use a helper that detects multipart
bodies) so that FormData requests allow fetch to create the correct
"multipart/form-data; boundary=…" header; keep the existing anthropic effort
header logic unchanged.
🧹 Nitpick comments (1)
packages/actions/src/prepare-request-body.ts (1)

725-725: Avoid the as unknown as ProviderRequestBody double-cast.

The return type Promise<ProviderRequestBody | FormData> doesn't account for OpenAIImageRequest. Either add OpenAIImageRequest to the ProviderRequestBody union in @llmgateway/models and export it from prepare-request-body.ts, or extend the return type to Promise<ProviderRequestBody | OpenAIImageRequest | FormData> so TypeScript can narrow the type properly without the cast.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/prepare-request-body.ts` at line 725, The code currently
forces a double-cast on openaiImageRequest in prepare-request-body.ts; update
the types instead of using as unknown as ProviderRequestBody: either add
OpenAIImageRequest to the ProviderRequestBody union in `@llmgateway/models` and
update any exports so prepare-request-body.ts can return it directly, or change
the function's return type to Promise<ProviderRequestBody | OpenAIImageRequest |
FormData> and export OpenAIImageRequest from prepare-request-body.ts so
TypeScript can narrow the value without the double-cast; adjust references to
openaiImageRequest and the function signature accordingly.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@packages/actions/src/prepare-request-body.ts`:
- Around line 33-68: fetchImageAsBlob is missing security and size protections
and doesn't use the available maxImageSizeMB; change its signature to accept
maxImageSizeMB (and isProd if needed) and enforce: validate URL scheme/host/IP
against SSRF-safe rules (reject localhost/169.254.169.254/private IPs when
isProd), apply a request timeout using AbortController or AbortSignal.timeout,
and stream/check response size to abort if it exceeds maxImageSizeMB; for data:
URLs, fix binary decoding by delegating to or reusing processImageUrl (which
returns {data,mimeType}) or correctly decode percent-encoded non-base64 payload
into bytes (don’t use Buffer.from(decodeURIComponent(...), "utf-8") which
corrupts binary), then construct the Blob and filename as before; update all
call sites to pass maxImageSizeMB/isProd.

---

Outside diff comments:
In `@apps/gateway/src/chat/tools/resolve-provider-context.ts`:
- Around line 396-411: The current block calls hasMaxTokens(requestBody) but
requestBody can be ProviderRequestBody | FormData; add a FormData type guard
before calling hasMaxTokens so multipart image-edit FormData bodies are skipped
(mirror the pattern used in chat.ts). Specifically, check that requestBody is
not an instance of FormData (or otherwise not FormData) before invoking
hasMaxTokens(requestBody), then proceed with the existing checks against
providerMappingForSelected.maxOutput and usedModel to throw the HTTPException
when requestBody.max_tokens exceeds providerMappingForSelected.maxOutput.
- Around line 413-425: The code unconditionally sets headers["Content-Type"] =
"application/json" which prevents fetch from auto-generating multipart
boundaries for FormData; instead, in resolve-provider-context.ts where headers
are built via getProviderHeaders and assigned to the headers variable, only set
the Content-Type when the outgoing requestBody is not multipart (e.g., guard
with a check like "requestBody is not FormData" or use a helper that detects
multipart bodies) so that FormData requests allow fetch to create the correct
"multipart/form-data; boundary=…" header; keep the existing anthropic effort
header logic unchanged.

---

Nitpick comments:
In `@packages/actions/src/prepare-request-body.ts`:
- Line 725: The code currently forces a double-cast on openaiImageRequest in
prepare-request-body.ts; update the types instead of using as unknown as
ProviderRequestBody: either add OpenAIImageRequest to the ProviderRequestBody
union in `@llmgateway/models` and update any exports so prepare-request-body.ts
can return it directly, or change the function's return type to
Promise<ProviderRequestBody | OpenAIImageRequest | FormData> and export
OpenAIImageRequest from prepare-request-body.ts so TypeScript can narrow the
value without the double-cast; adjust references to openaiImageRequest and the
function signature accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: a0b85c9c-6815-4055-a055-60df0a709436

📥 Commits

Reviewing files that changed from the base of the PR and between 88123ad and 1818a3a.

📒 Files selected for processing (3)
  • apps/gateway/src/chat/chat.ts
  • apps/gateway/src/chat/tools/resolve-provider-context.ts
  • packages/actions/src/prepare-request-body.ts

Comment on lines +33 to +68
async function fetchImageAsBlob(
url: string,
index: number,
): Promise<{ blob: Blob; filename: string }> {
const dataUrlMatch = url.match(/^data:([^;,]+)(?:;[^,]*)?,(.*)$/);
if (dataUrlMatch) {
const mimeType = dataUrlMatch[1] || "image/png";
const payload = dataUrlMatch[2] ?? "";
const isBase64 = /;base64,/i.test(url.slice(0, url.indexOf(",") + 1));
const raw = isBase64
? Buffer.from(payload, "base64")
: Buffer.from(decodeURIComponent(payload), "utf-8");
const buffer = new ArrayBuffer(raw.byteLength);
new Uint8Array(buffer).set(raw);
const ext = mimeType.split("/")[1]?.split("+")[0] ?? "png";
return {
blob: new Blob([buffer], { type: mimeType }),
filename: `image-${index}.${ext}`,
};
}

const response = await fetch(url);
if (!response.ok) {
throw new Error(
`Failed to fetch image ${url}: ${response.status} ${response.statusText}`,
);
}
const mimeType =
response.headers.get("content-type")?.split(";")[0]?.trim() || "image/png";
const buffer = await response.arrayBuffer();
const ext = mimeType.split("/")[1]?.split("+")[0] ?? "png";
return {
blob: new Blob([buffer], { type: mimeType }),
filename: `image-${index}.${ext}`,
};
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Confirm existing image-URL validation/SSRF/size utilities so we can reuse them here.
fd -e ts process-image-url
rg -nP -C2 '\b(maxImageSizeMB|process-image-url|processImageUrl|validateImageUrl|AbortSignal|AbortController)\b' packages/actions/src

Repository: theopenco/llmgateway

Length of output: 5365


🏁 Script executed:

#!/bin/bash
# First, examine the fetchImageAsBlob function in prepare-request-body.ts
head -80 packages/actions/src/prepare-request-body.ts | tail -50

echo "=== Checking process-image-url.ts for validation logic ==="
head -100 packages/actions/src/process-image-url.ts

Repository: theopenco/llmgateway

Length of output: 4761


🏁 Script executed:

#!/bin/bash
# Search for all calls to fetchImageAsBlob
echo "=== Looking for fetchImageAsBlob calls ==="
rg -n 'fetchImageAsBlob' packages/actions/src

echo "=== Context around those calls ==="
rg -B5 -A5 'fetchImageAsBlob' packages/actions/src

# Also check function signature and full implementation
echo "=== Full fetchImageAsBlob implementation context ==="
sed -n '30,80p' packages/actions/src/prepare-request-body.ts

Repository: theopenco/llmgateway

Length of output: 3762


🏁 Script executed:

#!/bin/bash
# Get more context around line 716 where fetchImageAsBlob is called
echo "=== Context around fetchImageAsBlob call at line 716 ==="
sed -n '700,730p' packages/actions/src/prepare-request-body.ts

echo "=== Check which provider/model this code path handles ==="
sed -n '650,720p' packages/actions/src/prepare-request-body.ts | head -50

echo "=== Check processImageUrl signature ==="
grep -A 10 'export async function processImageUrl' packages/actions/src/process-image-url.ts | head -15

Repository: theopenco/llmgateway

Length of output: 3154


🏁 Script executed:

#!/bin/bash
# Check the full function context around the fetchImageAsBlob call
echo "=== Full function containing fetchImageAsBlob call ==="
sed -n '610,750p' packages/actions/src/prepare-request-body.ts | grep -B 100 'fetchImageAsBlob' | head -80

echo "=== Check if maxImageSizeMB is available in this scope ==="
sed -n '610,720p' packages/actions/src/prepare-request-body.ts | grep -E '(maxImageSizeMB|function.*prepareRequestBody|isProd)'

Repository: theopenco/llmgateway

Length of output: 2786


🏁 Script executed:

#!/bin/bash
# Check what happens before and after fetchImageAsBlob call - look for any guards
echo "=== Full context of the fetchImageAsBlob call and surrounding code ==="
sed -n '680,750p' packages/actions/src/prepare-request-body.ts

echo "=== Check if there's any validation before this OpenAI image generation path ==="
sed -n '600,620p' packages/actions/src/prepare-request-body.ts

Repository: theopenco/llmgateway

Length of output: 3239


🏁 Script executed:

#!/bin/bash
# Verify the data-URL UTF-8 decoding issue more carefully
echo "=== Check how processImageUrl handles data URL decoding ==="
sed -n '50,120p' packages/actions/src/process-image-url.ts

# Also check if there's any upstream validation of URLs before they reach prepareRequestBody
echo "=== Check caller of prepareRequestBody to see if validation happens earlier ==="
rg -l 'prepareRequestBody' packages --type ts | head -3
rg -B3 'prepareRequestBody(' packages/actions/src/prepare-request-body.ts | head -20

Repository: theopenco/llmgateway

Length of output: 2506


🏁 Script executed:

#!/bin/bash
# Check the full processImageUrl fetch handling to see what else is validated
echo "=== Full fetch handling in processImageUrl ==="
sed -n '110,180p' packages/actions/src/process-image-url.ts

# Look for any validation that happens on imageUrls before they reach fetchImageAsBlob
echo "=== Search for any preprocessing of imageUrls in prepareRequestBody ==="
sed -n '640,720p' packages/actions/src/prepare-request-body.ts | grep -A 30 'imageUrls'

# Check if there's a timeout/AbortSignal utility available
echo "=== Check for timeout utilities in the codebase ==="
rg -n 'AbortSignal|timeout|AbortController' packages/actions/src --type ts | head -20

Repository: theopenco/llmgateway

Length of output: 3973


🏁 Script executed:

#!/bin/bash
# Verify there are no other calls to fetchImageAsBlob besides the one we found
echo "=== Comprehensive search for fetchImageAsBlob ==="
rg 'fetchImageAsBlob' packages

# Check if this code path handles edits specifically or if it's also for image generation
echo "=== Check OpenAI image generation vs edits flow ==="
sed -n '595,620p' packages/actions/src/prepare-request-body.ts

# Verify whether maxImageSizeMB parameter is in scope where fetchImageAsBlob is called
echo "=== Confirm maxImageSizeMB is in function scope ==="
sed -n '600,625p' packages/actions/src/prepare-request-body.ts | grep -E 'prepareRequestBody|maxImageSizeMB'

Repository: theopenco/llmgateway

Length of output: 1250


🏁 Script executed:

#!/bin/bash
# Final verification: confirm that processImageUrl is the proper replacement
echo "=== Verify processImageUrl handles both HTTP(S) URLs and data URLs correctly ==="
sed -n '39,110p' packages/actions/src/process-image-url.ts

# Check the return type compatibility
echo "=== Check return types ==="
echo "fetchImageAsBlob returns:"
sed -n '33,38p' packages/actions/src/prepare-request-body.ts | tail -1

echo ""
echo "processImageUrl returns:"
grep -A 1 'export async function processImageUrl' packages/actions/src/process-image-url.ts | tail -1

Repository: theopenco/llmgateway

Length of output: 2516


🏁 Script executed:

#!/bin/bash
# Check if AbortSignal.timeout is available (Node 17.3+) or if we need AbortController
echo "=== Check package.json Node version ==="
grep -E 'engines|node' packages/actions/package.json | head -5

# Check what version of Node is actually being used
echo "=== Check project Node version requirement ==="
cat package.json | grep -A 2 -B 2 'engines'

# Verify the exact issue with UTF-8 decoding for data URLs
echo "=== Inspect the data URL parsing issue more carefully ==="
sed -n '35,50p' packages/actions/src/prepare-request-body.ts

Repository: theopenco/llmgateway

Length of output: 885


Add size cap, timeout, SSRF validation, and enforce maxImageSizeMB parameter.

fetchImageAsBlob lacks all security safeguards on arbitrary user-supplied URLs:

  • No AbortSignal timeout (attacker can hang the request indefinitely)
  • No response size cap (can point at multi-GB files for DoS)
  • No scheme/host/IP validation (SSRF: http://169.254.169.254/, http://localhost/, internal endpoints)
  • maxImageSizeMB parameter available in scope but never passed or enforced

Additionally, the non-base64 data-URL branch (Buffer.from(decodeURIComponent(payload), "utf-8")) corrupts arbitrary binary bytes; use btoa() instead (as processImageUrl does).

The processImageUrl function in process-image-url.ts handles size and scheme validation, but has a different return type ({ data: string; mimeType: string } vs { blob: Blob; filename: string }). At minimum, pass maxImageSizeMB and isProd to fetchImageAsBlob, add size/scheme guards, and apply a timeout via AbortSignal.timeout() or AbortController. Fix the data-URL decoding path to use btoa() or delegate to processImageUrl with appropriate conversion logic.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@packages/actions/src/prepare-request-body.ts` around lines 33 - 68,
fetchImageAsBlob is missing security and size protections and doesn't use the
available maxImageSizeMB; change its signature to accept maxImageSizeMB (and
isProd if needed) and enforce: validate URL scheme/host/IP against SSRF-safe
rules (reject localhost/169.254.169.254/private IPs when isProd), apply a
request timeout using AbortController or AbortSignal.timeout, and stream/check
response size to abort if it exceeds maxImageSizeMB; for data: URLs, fix binary
decoding by delegating to or reusing processImageUrl (which returns
{data,mimeType}) or correctly decode percent-encoded non-base64 payload into
bytes (don’t use Buffer.from(decodeURIComponent(...), "utf-8") which corrupts
binary), then construct the Blob and filename as before; update all call sites
to pass maxImageSizeMB/isProd.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant