Skip to content

Guard image decode against oversized (decompression-bomb) images - #28588

Closed
hhhhhhhhhhhhhhhhho wants to merge 1 commit into
sgl-project:mainfrom
hhhhhhhhhhhhhhhhho:guard-oversized-image-decode
Closed

hhhhhhhhhhhhhhhhho wants to merge 1 commit into
sgl-project:mainfrom
hhhhhhhhhhhhhhhhho:guard-oversized-image-decode

Conversation

@hhhhhhhhhhhhhhhhho

@hhhhhhhhhhhhhhhhho hhhhhhhhhhhhhhhhho commented Jun 18, 2026 •

Copy link
Copy Markdown

Motivation

Closes #28587.

_load_image (python/sglang/srt/utils/common.py) decodes images via Image.open(BytesIO(...)) with no pixel-count guard. The per-model cap (SGLANG_IMAGE_MAX_PIXELS, used by smart_resize in qwen_vl.py) is only applied after PIL has fully decoded the image. So an oversized image — e.g. a 300-DPI rasterized PDF page, or a 12000×12000 PNG (144 M px) — is fully decoded into memory (≈ 432 MB RGB buffer) and only then downscaled.

Effect on a single Qwen3-VL request with such an image:

  • a CPU core pinned at ~100% for a long time (the reported case took ~1m50s),
  • RSS grows by ~1 GB per request.

With --max-running-requests 256, concurrent oversized-image requests scale memory linearly and can OOM the server before any GPU inference — a pre-inference DoS vector. vLLM short-circuits the same input via PIL's DecompressionBombWarning.

Note: #27451 truncated the exception message for malformed inputs to 100 chars, which fixed the raw-base64-in-logs / log-flood symptom for the malformed path. It does not help here: a valid oversized image never raises, so the decode cost and memory growth remain. See the issue for the full breakdown.

Modification

  • Add _check_image_pixels(width, height) in utils/common.py. Image.open only reads the header, so width/height are known before the expensive decode (.convert() / .load()). The check rejects images whose pixel count exceeds the limit with a clear ValueError, before any pixels are decoded.
  • Add SGLANG_IMAGE_MAX_DECODE_PIXELS (environ.py), default 89_478_485 (PIL's default decompression-bomb threshold); set 0 to disable.
  • Add a CPU unit test (test/registered/unit/utils/test_load_image_guard.py) covering within-limit, over-limit, and disabled cases.

This intentionally only guards the PIL decode path (the bomb vector). Malformed inputs that fail Image.open outright are already handled by #27451.

Checklist


CI States

Latest PR Test (Base): ❌ Run #27730377137
Latest PR Test (Extra): ❌ Run #27730376966

`_load_image` decoded images via `Image.open(BytesIO(...))` with no pixel-count
guard. The per-model cap (`SGLANG_IMAGE_MAX_PIXELS` in qwen_vl `smart_resize`) is
only applied *after* a full decode, so an oversized image (e.g. a high-DPI
rasterized PDF page, or a 12000x12000 PNG) is fully decoded into memory before
being downscaled. This pins a CPU core at ~100% for minutes and grows RSS by
~1GB per request; with many concurrent requests it is a pre-inference DoS vector.

`Image.open` only reads the header, so width/height are known before the
expensive decode. Add `_check_image_pixels` to reject images whose pixel count
exceeds `SGLANG_IMAGE_MAX_DECODE_PIXELS` (default = PIL's bomb threshold,
89,478,485; set 0 to disable) right after `Image.open`, before `.convert()`
forces the decode.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Missing pre-decode pixel-count validation enables pre-inference DoS and request log amplification

1 participant