feat(models): add gpt-image-2 from openai - #2059
Conversation
Add GPT Image 2 with token-based pricing: text $5/$10 per 1M in/out, cached text $1.25/1M, image input $8/1M, image output $30/1M. Routes through the image generations pipeline. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughAdds OpenAI image-generation support: new Changes
Sequence Diagram(s)sequenceDiagram
participant Client
participant Gateway
participant Actions as "Actions\n(prepare-request-body / get-provider-endpoint)"
participant OpenAI as "OpenAI API"
participant Parser as "Gateway Parser\n(parse-provider-response)"
Client->>Gateway: POST image generation request
Gateway->>Actions: prepare-request-body (extract prompt, images, image_config)
Actions-->>Gateway: requestBody (JSON or FormData with images)
Gateway->>Actions: get-provider-endpoint (resolve endpoint)
Actions-->>Gateway: /v1/images/generations
Gateway->>Gateway: if multipart and edit flow -> rewrite to /v1/images/edits
Gateway->>OpenAI: forward request (JSON or multipart)
OpenAI-->>Gateway: response (json.data with b64_json or url)
Gateway->>Parser: parse-provider-response (map images, tokens, finishReason)
Parser-->>Gateway: normalized response with images
Gateway-->>Client: return normalized response
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@packages/models/src/models/openai.ts`:
- Around line 1825-1835: The model config currently enables imageGenerations:
true while leaving imageOutputTokensByResolution unset and cachedInputPrice
incorrect; update the model object (the block containing cachedInputPrice,
imageInputPrice, imageOutputPrice, imageGenerations) to (1) set cachedInputPrice
to OpenAI's $2/M equivalent (2 / 1e6) instead of 1.25 / 1e6, (2) add a complete
imageOutputTokensByResolution mapping keyed by quality/size (populate token
estimates for low/medium/high quality and small/medium/large sizes consistent
with OpenAI guidance to avoid falling back to LEGACY_DEFAULT_TOKENS_PER_IMAGE),
and (3) ensure any gateway cost logic that reads imageOutputTokensByResolution
(references: LEGACY_DEFAULT_TOKENS_PER_IMAGE and imageOutputTokensByResolution)
will find the new map before allowing imageGenerations: true to be exposed in
production.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: a1b32d2d-6a86-460f-9127-039f26935a38
📒 Files selected for processing (1)
packages/models/src/models/openai.ts
| cachedInputPrice: 1.25 / 1e6, | ||
| imageInputPrice: 8 / 1e6, | ||
| imageOutputPrice: 30 / 1e6, | ||
| requestPrice: 0, | ||
| contextSize: 128000, | ||
| maxOutput: 4096, | ||
| streaming: false, | ||
| vision: true, | ||
| tools: false, | ||
| jsonOutput: false, | ||
| imageGenerations: true, |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Description: Inspect image-generation cost paths for GPT Image 2-specific
# size/quality token handling or actual API usage-token accounting.
# Expected: cost code should not rely only on LEGACY_DEFAULT_TOKENS_PER_IMAGE
# for gpt-image-2, and should account for separate cached image input pricing.
rg -nP -C3 '\b(gpt-image-2|imageOutputTokensByResolution|LEGACY_DEFAULT_TOKENS_PER_IMAGE|imageOutputPrice|imageInputPrice|cachedInputPrice|maxOutput|quality|imageSize|usage)\b' \
--glob '!**/node_modules/**' \
--glob '!**/dist/**' \
--glob '!**/build/**'Repository: theopenco/llmgateway
Length of output: 50377
🏁 Script executed:
find . -name "costs.ts" -o -name "cost.ts" | grep -E '(apps|packages)' | head -20Repository: theopenco/llmgateway
Length of output: 95
🏁 Script executed:
fd -H "costs.ts|cost.ts" --type fRepository: theopenco/llmgateway
Length of output: 93
🏁 Script executed:
rg -l "LEGACY_DEFAULT_TOKENS_PER_IMAGE|imageOutputTokensByResolution" --type ts --type tsx --max-count=20Repository: theopenco/llmgateway
Length of output: 91
🏁 Script executed:
sed -n '1810,1838p' packages/models/src/models/openai.tsRepository: theopenco/llmgateway
Length of output: 758
🏁 Script executed:
cat -n apps/gateway/src/lib/costs.tsRepository: theopenco/llmgateway
Length of output: 17022
🏁 Script executed:
rg -n "imageGenerations|imageOutputTokensByResolution|imageOutputPrice" apps/gateway/src/lib/costs.ts -A 5 -B 5Repository: theopenco/llmgateway
Length of output: 1380
Add GPT Image 2 quality/size-dependent output token handling before enabling imageGenerations: true.
The cost calculation in apps/gateway/src/lib/costs.ts (lines 363–371) resolves imageOutputTokensByResolution for per-resolution token counts. When absent, it falls back to LEGACY_DEFAULT_TOKENS_PER_IMAGE = 1120 tokens per image. GPT Image 2 output tokens vary significantly by quality and size parameters—OpenAI's guide documents costs ranging from ~$0.006 (low quality, small size) to ~$0.211 (high quality, large size), which corresponds to roughly 200–7000 output tokens at the configured $30/M price. The static 1120-token fallback will systematically misbill these requests.
Additionally, the cachedInputPrice: 1.25 / 1e6 in the model definition does not align with OpenAI's separate $2/M cached image input pricing, and there is no corresponding cached image input cost handler in the cost path.
Populate imageOutputTokensByResolution with quality/size-specific token values (from OpenAI's documentation or API usage data) and ensure cached image input pricing is correctly represented before marking this model selectable for production traffic.
Sources: OpenAI pricing, OpenAI image generation guide.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@packages/models/src/models/openai.ts` around lines 1825 - 1835, The model
config currently enables imageGenerations: true while leaving
imageOutputTokensByResolution unset and cachedInputPrice incorrect; update the
model object (the block containing cachedInputPrice, imageInputPrice,
imageOutputPrice, imageGenerations) to (1) set cachedInputPrice to OpenAI's $2/M
equivalent (2 / 1e6) instead of 1.25 / 1e6, (2) add a complete
imageOutputTokensByResolution mapping keyed by quality/size (populate token
estimates for low/medium/high quality and small/medium/large sizes consistent
with OpenAI guidance to avoid falling back to LEGACY_DEFAULT_TOKENS_PER_IMAGE),
and (3) ensure any gateway cost logic that reads imageOutputTokensByResolution
(references: LEGACY_DEFAULT_TOKENS_PER_IMAGE and imageOutputTokensByResolution)
will find the new map before allowing imageGenerations: true to be exposed in
production.
- Route openai imageGenerations to /v1/images/generations
- Transform chat-shape payload to images API shape and normalize
"1K"/aspect-ratio into 1024x1024/1024x1536/1536x1024
- Parse { data:[{b64_json|url}] } responses back into images
- Swap to /v1/images/edits when an input image is present
Also fix infinite /orgs refetch loop: image/video page URL-sync
effects depended on searchParams, causing router.replace to
produce a new ref that retriggered the effect and refetched the
RSC forever. Match chat-page-client pattern (read
window.location.search, compare before replace, scroll:false).
There was a problem hiding this comment.
Actionable comments posted: 4
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@apps/gateway/src/chat/chat.ts`:
- Around line 3900-3908: In the OpenAI image-edit branch inside
apps/gateway/src/chat/chat.ts (the block that rewrites url from
"/v1/images/generations" to "/v1/images/edits"), change the payload presence
check to look for either "image" or "images" (e.g., ("image" in requestBody ||
"images" in requestBody)) so edits are selected for both singular and plural
image fields; also consider updating the prepareRequestBody logic that builds
OpenAI image payloads to emit an "images" array for multi-image requests to
match xAI's pattern and OpenAI's API expectations.
In `@apps/gateway/src/chat/tools/parse-provider-response.ts`:
- Around line 538-550: The mapping over imageData can produce undefined urls
because only json.data[0] was checked; update the images construction so each
item is validated: for each item in imageData check item.b64_json and item.url,
build url = `data:image/png;base64,...` when b64_json exists, otherwise use
item.url only if defined, and skip (or log and continue) any item lacking both
values instead of returning an ImageObject with an undefined url; adjust the
images assignment that produces ImageObject entries to filter out invalid items
and ensure returned objects always have a defined image_url.url.
In `@packages/actions/src/prepare-request-body.ts`:
- Around line 646-659: The code returns an object named openaiImageRequest typed
as any containing model/prompt/size/n and image fields, but when routing to the
OpenAI /v1/images/edits endpoint you must send multipart/form-data with binary
files rather than JSON; change the return type from any to an explicit interface
(e.g., ImageRequest { model: string; prompt?: string; size?: string; n?: number;
image?: string | string[] }) and detect the edits flow (the URL swap in chat.ts)
to build and return a FormData instead of a JSON object: decode data URLs in
imageUrls into binary Blobs/Uint8Arrays and append each as image[] with
filenames, plus append model/prompt/size/n as form fields, otherwise continue
returning the typed JSON object; ensure no use of any.
- Around line 610-644: The code in prepare-request-body.ts incorrectly restricts
allowed image sizes by mapping any non-whitelisted rawSize to "auto" (affecting
the openaiSize assigned from image_config?.image_size) and so breaks valid
gpt-image-2 dimensions; change the logic in the block that computes openaiSize
to accept numeric or arbitrary "WIDTHxHEIGHT" strings that meet OpenAI
constraints (multiples of 16, max edge ≤3840, aspect ratio ≤3:1, pixel count
between 655,360 and 8,294,400) instead of forcing "auto", and preserve the
existing aspectRatio-based fallback behavior for presets; also replace the
untyped openaiImageRequest (currently declared as any) with a proper
interface/type (e.g., OpenAIImageRequest) or a typed cast reflecting the OpenAI
image generation request shape so the request object is strongly typed rather
than any.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: bc5c6909-34e4-4867-862b-65412e52c0a4
📒 Files selected for processing (6)
apps/gateway/src/chat/chat.tsapps/gateway/src/chat/tools/parse-provider-response.tsapps/playground/src/components/playground/image-page-client.tsxapps/playground/src/components/playground/video-page-client.tsxpackages/actions/src/get-provider-endpoint.tspackages/actions/src/prepare-request-body.ts
| const imageData = json.data; | ||
| images = imageData.map((item: any): ImageObject => { | ||
| let url: string; | ||
| if (item.b64_json) { | ||
| url = `data:image/png;base64,${item.b64_json}`; | ||
| } else { | ||
| url = item.url; | ||
| } | ||
| return { | ||
| type: "image_url", | ||
| image_url: { url }, | ||
| }; | ||
| }); |
There was a problem hiding this comment.
Per-item fallback may produce an undefined URL.
The outer guard on Line 536 only checks json.data[0], but the map iterates every item. If a later element has neither b64_json nor url, the else branch assigns url = item.url (undefined), producing a malformed ImageObject. Consider guarding per item and skipping/erroring on invalid entries.
🛡️ Suggested hardening
- images = imageData.map((item: any): ImageObject => {
- let url: string;
- if (item.b64_json) {
- url = `data:image/png;base64,${item.b64_json}`;
- } else {
- url = item.url;
- }
- return {
- type: "image_url",
- image_url: { url },
- };
- });
+ images = imageData
+ .filter((item: any) => item?.b64_json || item?.url)
+ .map((item: any): ImageObject => ({
+ type: "image_url",
+ image_url: {
+ url: item.b64_json
+ ? `data:image/png;base64,${item.b64_json}`
+ : item.url,
+ },
+ }));📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| const imageData = json.data; | |
| images = imageData.map((item: any): ImageObject => { | |
| let url: string; | |
| if (item.b64_json) { | |
| url = `data:image/png;base64,${item.b64_json}`; | |
| } else { | |
| url = item.url; | |
| } | |
| return { | |
| type: "image_url", | |
| image_url: { url }, | |
| }; | |
| }); | |
| const imageData = json.data; | |
| images = imageData | |
| .filter((item: any) => item?.b64_json || item?.url) | |
| .map((item: any): ImageObject => ({ | |
| type: "image_url", | |
| image_url: { | |
| url: item.b64_json | |
| ? `data:image/png;base64,${item.b64_json}` | |
| : item.url, | |
| }, | |
| })); |
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@apps/gateway/src/chat/tools/parse-provider-response.ts` around lines 538 -
550, The mapping over imageData can produce undefined urls because only
json.data[0] was checked; update the images construction so each item is
validated: for each item in imageData check item.b64_json and item.url, build
url = `data:image/png;base64,...` when b64_json exists, otherwise use item.url
only if defined, and skip (or log and continue) any item lacking both values
instead of returning an ImageObject with an undefined url; adjust the images
assignment that produces ImageObject entries to filter out invalid items and
ensure returned objects always have a defined image_url.url.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
apps/gateway/src/chat/tools/resolve-provider-context.ts (2)
396-411:⚠️ Potential issue | 🔴 CriticalMissing
FormDatatype guard beforehasMaxTokens.
requestBodyis nowProviderRequestBody | FormData, buthasMaxTokensis typed to acceptProviderRequestBodyonly (packages/models/src/types.ts). Passing a potentialFormDatahere is both a TS type error and semantically wrong —"max_tokens" in formDatais alwaysfalse, but it shouldn't be calling this branch at all for image-edit multipart bodies.The sibling callsite in
apps/gateway/src/chat/chat.ts(3866-3870) already uses the correct pattern; mirror it here:🩹 Proposed fix
// Post-validation of max_tokens in request body if ( - hasMaxTokens(requestBody) && + !(requestBody instanceof FormData) && + hasMaxTokens(requestBody) && requestBody.max_tokens !== undefined && providerMappingForSelected ) {🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@apps/gateway/src/chat/tools/resolve-provider-context.ts` around lines 396 - 411, The current block calls hasMaxTokens(requestBody) but requestBody can be ProviderRequestBody | FormData; add a FormData type guard before calling hasMaxTokens so multipart image-edit FormData bodies are skipped (mirror the pattern used in chat.ts). Specifically, check that requestBody is not an instance of FormData (or otherwise not FormData) before invoking hasMaxTokens(requestBody), then proceed with the existing checks against providerMappingForSelected.maxOutput and usedModel to throw the HTTPException when requestBody.max_tokens exceeds providerMappingForSelected.maxOutput.
413-425:⚠️ Potential issue | 🔴 CriticalFix unconditional
Content-Type: application/jsonheader to allow FormData multipart boundaries.When
requestBodyisFormData, the unconditionalheaders["Content-Type"] = "application/json"at line 418 blocksfetchfrom auto-generating the requiredmultipart/form-data; boundary=…header, causing OpenAI's/v1/images/editsand similar endpoints to reject with 400. The same pattern is already correctly implemented elsewhere in the codebase (seechat.ts:7827–7829).Conditionally set the header only when the body is not multipart:
const headers = getProviderHeaders(usedProvider as Provider, usedToken, { requestId: options.requestId, webSearchEnabled: options.webSearchEnabled, }); -headers["Content-Type"] = "application/json"; +if (!(requestBody instanceof FormData)) { + headers["Content-Type"] = "application/json"; +}🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@apps/gateway/src/chat/tools/resolve-provider-context.ts` around lines 413 - 425, The code unconditionally sets headers["Content-Type"] = "application/json" which prevents fetch from auto-generating multipart boundaries for FormData; instead, in resolve-provider-context.ts where headers are built via getProviderHeaders and assigned to the headers variable, only set the Content-Type when the outgoing requestBody is not multipart (e.g., guard with a check like "requestBody is not FormData" or use a helper that detects multipart bodies) so that FormData requests allow fetch to create the correct "multipart/form-data; boundary=…" header; keep the existing anthropic effort header logic unchanged.
🧹 Nitpick comments (1)
packages/actions/src/prepare-request-body.ts (1)
725-725: Avoid theas unknown as ProviderRequestBodydouble-cast.The return type
Promise<ProviderRequestBody | FormData>doesn't account forOpenAIImageRequest. Either addOpenAIImageRequestto theProviderRequestBodyunion in@llmgateway/modelsand export it fromprepare-request-body.ts, or extend the return type toPromise<ProviderRequestBody | OpenAIImageRequest | FormData>so TypeScript can narrow the type properly without the cast.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@packages/actions/src/prepare-request-body.ts` at line 725, The code currently forces a double-cast on openaiImageRequest in prepare-request-body.ts; update the types instead of using as unknown as ProviderRequestBody: either add OpenAIImageRequest to the ProviderRequestBody union in `@llmgateway/models` and update any exports so prepare-request-body.ts can return it directly, or change the function's return type to Promise<ProviderRequestBody | OpenAIImageRequest | FormData> and export OpenAIImageRequest from prepare-request-body.ts so TypeScript can narrow the value without the double-cast; adjust references to openaiImageRequest and the function signature accordingly.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@packages/actions/src/prepare-request-body.ts`:
- Around line 33-68: fetchImageAsBlob is missing security and size protections
and doesn't use the available maxImageSizeMB; change its signature to accept
maxImageSizeMB (and isProd if needed) and enforce: validate URL scheme/host/IP
against SSRF-safe rules (reject localhost/169.254.169.254/private IPs when
isProd), apply a request timeout using AbortController or AbortSignal.timeout,
and stream/check response size to abort if it exceeds maxImageSizeMB; for data:
URLs, fix binary decoding by delegating to or reusing processImageUrl (which
returns {data,mimeType}) or correctly decode percent-encoded non-base64 payload
into bytes (don’t use Buffer.from(decodeURIComponent(...), "utf-8") which
corrupts binary), then construct the Blob and filename as before; update all
call sites to pass maxImageSizeMB/isProd.
---
Outside diff comments:
In `@apps/gateway/src/chat/tools/resolve-provider-context.ts`:
- Around line 396-411: The current block calls hasMaxTokens(requestBody) but
requestBody can be ProviderRequestBody | FormData; add a FormData type guard
before calling hasMaxTokens so multipart image-edit FormData bodies are skipped
(mirror the pattern used in chat.ts). Specifically, check that requestBody is
not an instance of FormData (or otherwise not FormData) before invoking
hasMaxTokens(requestBody), then proceed with the existing checks against
providerMappingForSelected.maxOutput and usedModel to throw the HTTPException
when requestBody.max_tokens exceeds providerMappingForSelected.maxOutput.
- Around line 413-425: The code unconditionally sets headers["Content-Type"] =
"application/json" which prevents fetch from auto-generating multipart
boundaries for FormData; instead, in resolve-provider-context.ts where headers
are built via getProviderHeaders and assigned to the headers variable, only set
the Content-Type when the outgoing requestBody is not multipart (e.g., guard
with a check like "requestBody is not FormData" or use a helper that detects
multipart bodies) so that FormData requests allow fetch to create the correct
"multipart/form-data; boundary=…" header; keep the existing anthropic effort
header logic unchanged.
---
Nitpick comments:
In `@packages/actions/src/prepare-request-body.ts`:
- Line 725: The code currently forces a double-cast on openaiImageRequest in
prepare-request-body.ts; update the types instead of using as unknown as
ProviderRequestBody: either add OpenAIImageRequest to the ProviderRequestBody
union in `@llmgateway/models` and update any exports so prepare-request-body.ts
can return it directly, or change the function's return type to
Promise<ProviderRequestBody | OpenAIImageRequest | FormData> and export
OpenAIImageRequest from prepare-request-body.ts so TypeScript can narrow the
value without the double-cast; adjust references to openaiImageRequest and the
function signature accordingly.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: a0b85c9c-6815-4055-a055-60df0a709436
📒 Files selected for processing (3)
apps/gateway/src/chat/chat.tsapps/gateway/src/chat/tools/resolve-provider-context.tspackages/actions/src/prepare-request-body.ts
| async function fetchImageAsBlob( | ||
| url: string, | ||
| index: number, | ||
| ): Promise<{ blob: Blob; filename: string }> { | ||
| const dataUrlMatch = url.match(/^data:([^;,]+)(?:;[^,]*)?,(.*)$/); | ||
| if (dataUrlMatch) { | ||
| const mimeType = dataUrlMatch[1] || "image/png"; | ||
| const payload = dataUrlMatch[2] ?? ""; | ||
| const isBase64 = /;base64,/i.test(url.slice(0, url.indexOf(",") + 1)); | ||
| const raw = isBase64 | ||
| ? Buffer.from(payload, "base64") | ||
| : Buffer.from(decodeURIComponent(payload), "utf-8"); | ||
| const buffer = new ArrayBuffer(raw.byteLength); | ||
| new Uint8Array(buffer).set(raw); | ||
| const ext = mimeType.split("/")[1]?.split("+")[0] ?? "png"; | ||
| return { | ||
| blob: new Blob([buffer], { type: mimeType }), | ||
| filename: `image-${index}.${ext}`, | ||
| }; | ||
| } | ||
|
|
||
| const response = await fetch(url); | ||
| if (!response.ok) { | ||
| throw new Error( | ||
| `Failed to fetch image ${url}: ${response.status} ${response.statusText}`, | ||
| ); | ||
| } | ||
| const mimeType = | ||
| response.headers.get("content-type")?.split(";")[0]?.trim() || "image/png"; | ||
| const buffer = await response.arrayBuffer(); | ||
| const ext = mimeType.split("/")[1]?.split("+")[0] ?? "png"; | ||
| return { | ||
| blob: new Blob([buffer], { type: mimeType }), | ||
| filename: `image-${index}.${ext}`, | ||
| }; | ||
| } |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Confirm existing image-URL validation/SSRF/size utilities so we can reuse them here.
fd -e ts process-image-url
rg -nP -C2 '\b(maxImageSizeMB|process-image-url|processImageUrl|validateImageUrl|AbortSignal|AbortController)\b' packages/actions/srcRepository: theopenco/llmgateway
Length of output: 5365
🏁 Script executed:
#!/bin/bash
# First, examine the fetchImageAsBlob function in prepare-request-body.ts
head -80 packages/actions/src/prepare-request-body.ts | tail -50
echo "=== Checking process-image-url.ts for validation logic ==="
head -100 packages/actions/src/process-image-url.tsRepository: theopenco/llmgateway
Length of output: 4761
🏁 Script executed:
#!/bin/bash
# Search for all calls to fetchImageAsBlob
echo "=== Looking for fetchImageAsBlob calls ==="
rg -n 'fetchImageAsBlob' packages/actions/src
echo "=== Context around those calls ==="
rg -B5 -A5 'fetchImageAsBlob' packages/actions/src
# Also check function signature and full implementation
echo "=== Full fetchImageAsBlob implementation context ==="
sed -n '30,80p' packages/actions/src/prepare-request-body.tsRepository: theopenco/llmgateway
Length of output: 3762
🏁 Script executed:
#!/bin/bash
# Get more context around line 716 where fetchImageAsBlob is called
echo "=== Context around fetchImageAsBlob call at line 716 ==="
sed -n '700,730p' packages/actions/src/prepare-request-body.ts
echo "=== Check which provider/model this code path handles ==="
sed -n '650,720p' packages/actions/src/prepare-request-body.ts | head -50
echo "=== Check processImageUrl signature ==="
grep -A 10 'export async function processImageUrl' packages/actions/src/process-image-url.ts | head -15Repository: theopenco/llmgateway
Length of output: 3154
🏁 Script executed:
#!/bin/bash
# Check the full function context around the fetchImageAsBlob call
echo "=== Full function containing fetchImageAsBlob call ==="
sed -n '610,750p' packages/actions/src/prepare-request-body.ts | grep -B 100 'fetchImageAsBlob' | head -80
echo "=== Check if maxImageSizeMB is available in this scope ==="
sed -n '610,720p' packages/actions/src/prepare-request-body.ts | grep -E '(maxImageSizeMB|function.*prepareRequestBody|isProd)'Repository: theopenco/llmgateway
Length of output: 2786
🏁 Script executed:
#!/bin/bash
# Check what happens before and after fetchImageAsBlob call - look for any guards
echo "=== Full context of the fetchImageAsBlob call and surrounding code ==="
sed -n '680,750p' packages/actions/src/prepare-request-body.ts
echo "=== Check if there's any validation before this OpenAI image generation path ==="
sed -n '600,620p' packages/actions/src/prepare-request-body.tsRepository: theopenco/llmgateway
Length of output: 3239
🏁 Script executed:
#!/bin/bash
# Verify the data-URL UTF-8 decoding issue more carefully
echo "=== Check how processImageUrl handles data URL decoding ==="
sed -n '50,120p' packages/actions/src/process-image-url.ts
# Also check if there's any upstream validation of URLs before they reach prepareRequestBody
echo "=== Check caller of prepareRequestBody to see if validation happens earlier ==="
rg -l 'prepareRequestBody' packages --type ts | head -3
rg -B3 'prepareRequestBody(' packages/actions/src/prepare-request-body.ts | head -20Repository: theopenco/llmgateway
Length of output: 2506
🏁 Script executed:
#!/bin/bash
# Check the full processImageUrl fetch handling to see what else is validated
echo "=== Full fetch handling in processImageUrl ==="
sed -n '110,180p' packages/actions/src/process-image-url.ts
# Look for any validation that happens on imageUrls before they reach fetchImageAsBlob
echo "=== Search for any preprocessing of imageUrls in prepareRequestBody ==="
sed -n '640,720p' packages/actions/src/prepare-request-body.ts | grep -A 30 'imageUrls'
# Check if there's a timeout/AbortSignal utility available
echo "=== Check for timeout utilities in the codebase ==="
rg -n 'AbortSignal|timeout|AbortController' packages/actions/src --type ts | head -20Repository: theopenco/llmgateway
Length of output: 3973
🏁 Script executed:
#!/bin/bash
# Verify there are no other calls to fetchImageAsBlob besides the one we found
echo "=== Comprehensive search for fetchImageAsBlob ==="
rg 'fetchImageAsBlob' packages
# Check if this code path handles edits specifically or if it's also for image generation
echo "=== Check OpenAI image generation vs edits flow ==="
sed -n '595,620p' packages/actions/src/prepare-request-body.ts
# Verify whether maxImageSizeMB parameter is in scope where fetchImageAsBlob is called
echo "=== Confirm maxImageSizeMB is in function scope ==="
sed -n '600,625p' packages/actions/src/prepare-request-body.ts | grep -E 'prepareRequestBody|maxImageSizeMB'Repository: theopenco/llmgateway
Length of output: 1250
🏁 Script executed:
#!/bin/bash
# Final verification: confirm that processImageUrl is the proper replacement
echo "=== Verify processImageUrl handles both HTTP(S) URLs and data URLs correctly ==="
sed -n '39,110p' packages/actions/src/process-image-url.ts
# Check the return type compatibility
echo "=== Check return types ==="
echo "fetchImageAsBlob returns:"
sed -n '33,38p' packages/actions/src/prepare-request-body.ts | tail -1
echo ""
echo "processImageUrl returns:"
grep -A 1 'export async function processImageUrl' packages/actions/src/process-image-url.ts | tail -1Repository: theopenco/llmgateway
Length of output: 2516
🏁 Script executed:
#!/bin/bash
# Check if AbortSignal.timeout is available (Node 17.3+) or if we need AbortController
echo "=== Check package.json Node version ==="
grep -E 'engines|node' packages/actions/package.json | head -5
# Check what version of Node is actually being used
echo "=== Check project Node version requirement ==="
cat package.json | grep -A 2 -B 2 'engines'
# Verify the exact issue with UTF-8 decoding for data URLs
echo "=== Inspect the data URL parsing issue more carefully ==="
sed -n '35,50p' packages/actions/src/prepare-request-body.tsRepository: theopenco/llmgateway
Length of output: 885
Add size cap, timeout, SSRF validation, and enforce maxImageSizeMB parameter.
fetchImageAsBlob lacks all security safeguards on arbitrary user-supplied URLs:
- No
AbortSignaltimeout (attacker can hang the request indefinitely) - No response size cap (can point at multi-GB files for DoS)
- No scheme/host/IP validation (SSRF:
http://169.254.169.254/,http://localhost/, internal endpoints) maxImageSizeMBparameter available in scope but never passed or enforced
Additionally, the non-base64 data-URL branch (Buffer.from(decodeURIComponent(payload), "utf-8")) corrupts arbitrary binary bytes; use btoa() instead (as processImageUrl does).
The processImageUrl function in process-image-url.ts handles size and scheme validation, but has a different return type ({ data: string; mimeType: string } vs { blob: Blob; filename: string }). At minimum, pass maxImageSizeMB and isProd to fetchImageAsBlob, add size/scheme guards, and apply a timeout via AbortSignal.timeout() or AbortController. Fix the data-URL decoding path to use btoa() or delegate to processImageUrl with appropriate conversion logic.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@packages/actions/src/prepare-request-body.ts` around lines 33 - 68,
fetchImageAsBlob is missing security and size protections and doesn't use the
available maxImageSizeMB; change its signature to accept maxImageSizeMB (and
isProd if needed) and enforce: validate URL scheme/host/IP against SSRF-safe
rules (reject localhost/169.254.169.254/private IPs when isProd), apply a
request timeout using AbortController or AbortSignal.timeout, and stream/check
response size to abort if it exceeds maxImageSizeMB; for data: URLs, fix binary
decoding by delegating to or reusing processImageUrl (which returns
{data,mimeType}) or correctly decode percent-encoded non-base64 payload into
bytes (don’t use Buffer.from(decodeURIComponent(...), "utf-8") which corrupts
binary), then construct the Blob and filename as before; update all call sites
to pass maxImageSizeMB/isProd.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Summary
Test plan
🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Bug Fixes