Skip to content

perf(multimodal): optimize preprocessing serialization, tensor conversion, and pad fusion - #1012

Merged
CatherineSue merged 5 commits into
mainfrom
perf/multimodal-profiling
Apr 1, 2026
Merged

CatherineSue merged 5 commits into
mainfrom
perf/multimodal-profiling

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Apr 1, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

Perf profiling of the multimodal preprocessing pipeline during real MMMU benchmarks (900 images, Qwen3-VL-8B + Llama4-Scout-17B) revealed three CPU hotspots in SMG that together consumed 66% of all SMG CPU time:

  1. serialize_pixel_values (31% CPU) — FlattenCompat::fold iterating per-element f32::to_le_bytes() to serialize pixel tensors for gRPC
  2. build_planar_tensor (22% CPU) — scalar RGB deinterleave loop that the compiler couldn't auto-vectorize
  3. pad_image (13% CPU, Llama4 only) — image::overlay + get_pixel creating an intermediate padded RgbImage

Solution

Three targeted optimizations based on perf flamegraph analysis:

  1. serialize_pixel_values → bytemuck::cast_slice: Zero-copy reinterpret &[f32] as &[u8] on little-endian, replacing per-element flat_map(to_le_bytes).collect(). Eliminates the the top CPU hotspot entirely.

  2. build_planar_tensor block processing: Restructure the RGB deinterleave loop into fixed 8-pixel blocks so the compiler can unroll and auto-vectorize the stride-3 gather pattern.

  3. pad_and_normalize_to_tensor fusion (Llama4): Build the padded f32 tensor directly from resized RGB bytes, skipping the intermediate padded RgbImage allocation. Pre-fills padding regions with the normalized black value, then overwrites the image region row-by-row.

Changes

  • model_gateway/src/routers/grpc/multimodal.rs — Replace flat_map(to_le_bytes) with bytemuck::cast_slice in serialize_pixel_values
  • crates/multimodal/src/vision/transforms.rs — Restructure build_planar_tensor into 8-pixel blocks; add pad_and_normalize_to_tensor fused function
  • crates/multimodal/src/vision/processors/llama4_vision.rs — Use fused pad_and_normalize_to_tensor, remove dead pad_image method
  • Cargo.toml — Add [profile.profiling] for future perf profiling with debug symbols (Removed in later commits)

Test Plan

Profiling methodology

Profiled with perf record -F 4999 -g --call-graph fp during real MMMU benchmark runs. To get meaningful CPU samples (SMG is >99% idle during inference), sent 500 concurrent requests with max_tokens=1.

Built with a new [profile.profiling] that inherits release but keeps debug symbols:

[profile.profiling]
inherits = "release"
debug = 1
strip = false

Commands used

# vLLM gRPC (Qwen3-VL)
CUDA_VISIBLE_DEVICES=0 vllm serve /raid/models/Qwen/Qwen3-VL-8B-Instruct \
  --tensor-parallel-size 1 --port 8080 --grpc --generation-config vllm

# vLLM gRPC (Llama4-Scout, 4x H100)
CUDA_VISIBLE_DEVICES=0,1,2,3 vllm serve /raid/models/meta-llama/Llama-4-Scout-17B-16E-Instruct \
  --tensor-parallel-size 4 --port 8080 --grpc --generation-config vllm --max-model-len 131072

# SMG gateway
~/.cargo/target/profiling/smg \
  --host 0.0.0.0 --port 3002 --prometheus-port 9321 \
  --worker-urls grpc://127.0.0.1:8080 \
  --model-path /raid/models/<model> --log-level debug

# MMMU benchmark (900 images)
OPENAI_API_KEY=EMPTY python3 -m lmms_eval \
  --model openai \
  --model_args "model_version=<model>,base_url=http://localhost:3002/v1" \
  --tasks mmmu_val --batch_size 1 --output_path ~/mmmu_results/<run>

Perf profile: before vs after (Llama4-Scout, 66K samples)

Function Baseline % CPU Optimized % CPU Change
FlattenCompat::fold (serialize) 23.8% gone eliminated by opt-1
serialize_pixel_values (outer) 7.5% gone eliminated by opt-1
build_planar_tensor 21.9% 5.0% -77% by opt-2
image::get_pixel (pad) 3.2% gone eliminated by opt-3
image::overlay (pad) 2.4% gone eliminated by opt-3
ImageBuffer::from_pixel (pad) 1.7% gone eliminated by opt-3
ImageBuffer::get_pixel_mut (pad) 1.9% gone eliminated by opt-3
FIR resize (AVX2) 2.6% 4.5% same absolute, larger share
ndarray::assign (tile) 3.4% 6.8% same absolute, larger share

Preprocessing wall-clock timing (Llama4-Scout, 900 MMMU images)

Step Baseline Optimized Change
resize 0.8ms 1.1ms ~same
pad 5.2ms 0.0ms eliminated
tensor 8.5ms 6.4ms -25%
tile 11.0ms 8.6ms -22%
total avg 25.5ms 16.1ms -37%
total median 21.6ms 9.8ms -55%
total P95 60.9ms 50.9ms -16%

Perf profile confirms same hotspots on Qwen3-VL (209 samples)

% CPU Function Category
13.3% FlattenCompat::fold serialize_pixel_values
9.2% build_planar_tensor tensor conversion
5.3% fdeflate::Decompressor PNG decode
4.5% Vec::resize patchify alloc
4.1% serialize_pixel_values serialization

MMMU accuracy (Llama4-Scout, 900 images)

Score
Baseline 0.41667
Optimized 0.42778

No regression — scores are within noise (optimizations don't change output values).

Correctness

  • bytemuck::cast_slice produces bit-identical output to to_le_bytes() on little-endian (guarded by #[cfg(target_endian = "little")])
  • build_planar_tensor block processing: same arithmetic, just restructured for auto-vectorization
  • pad_and_normalize_to_tensor: same normalized values, verified by pre-fill with bias (normalized black) + row-by-row overwrite
End-to-end MMMU wall-clock time (Llama4-Scout, 900 images)

Preprocessing is <0.01% of total request time — model inference dominates. The ~8s saved across 900 images matches the per-image improvement (900 × 9.4ms ≈ 8.5s). The real benefit is reduced SMG CPU utilization under concurrent load.

Baseline

Throughput Summary

|      Metric      |  Value   |  Unit  |
|------------------|---------:|--------|
|total_gen_tokens  |78352.0000|tokens  |
|total_elapsed_time| 2329.9494|seconds |
|avg_speed         |   33.6282|tokens/s|

Optimized

Throughput Summary

|      Metric      |  Value   |  Unit  |
|------------------|---------:|--------|
|total_gen_tokens  |78952.0000|tokens  |
|total_elapsed_time| 2321.6685|seconds |
|avg_speed         |   34.0066|tokens/s|
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Refactor
    • Fused image padding and normalization into a single, more efficient processing path for faster and lower‑overhead image handling.
    • Improved pixel serialization for significantly better performance on common platforms.
  • Chores
    • Enhanced runtime diagnostics to include image size info and updated tests for image handling.

…rumentation

Signed-off-by: Chang Su <chang.s.su@oracle.com>
…r, and pad+tensor fusion

Three optimizations based on perf profiling of MMMU benchmark:

1. serialize_pixel_values: replace per-element flat_map(to_le_bytes) with
   bytemuck::cast_slice zero-copy reinterpret. Eliminates FlattenCompat
   which was #1 CPU hotspot (24-31% of SMG CPU).

2. build_planar_tensor: restructure RGB deinterleave loop into 8-pixel
   blocks for better auto-vectorization. Reduces tensor conversion from
   22% to 5% of CPU.

3. Llama4 pad_and_normalize_to_tensor: fuse pad_image + to_tensor into
   one pass, eliminating intermediate padded RgbImage allocation.
   Removes 13% CPU overhead from image::overlay + get_pixel.

Combined effect on Llama4 preprocessing: avg 25.5ms -> 16.1ms (-37%),
median 21.6ms -> 9.8ms (-55%). MMMU accuracy unchanged (0.417 -> 0.428).

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added dependencies Dependency updates grpc gRPC client and router changes model-gateway Model gateway crate changes labels Apr 1, 2026
@coderabbitai

coderabbitai Bot commented Apr 1, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Fuses image padding and normalization into a single tensor construction path, exposes and centralizes RGB deinterleaving, and optimizes pixel byte serialization on little-endian targets.

Changes

Cohort / File(s) Summary
Vision pad+normalize & processor
crates/multimodal/src/vision/processors/llama4_vision.rs
Added pad_and_normalize_to_tensor and updated process_single_image to use fused pad+normalize when resized != target canvas; removed intermediate padded RgbImage path and adjusted comments/step numbering.
Vision transforms & helpers
crates/multimodal/src/vision/transforms.rs
Made rgb_bytes public and added pub fn deinterleave_rgb_to_planes(...) to deinterleave RGB into planar f32 slices with an 8-pixel block loop and debug assertions; replaced per-pixel planar build with this helper.
gRPC serialization
model_gateway/src/routers/grpc/multimodal.rs
Logged image sizes in process_multimodal_parts; in serialize_pixel_values used bytemuck::cast_slice to produce contiguous u8 from &[f32] on little-endian targets, preserving per-element conversion on other endianness.
Tests (imports)
crates/multimodal/src/vision/... (test files)
Added imports for Rgb and RgbImage to support constructed test images.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested labels

multimodal

Suggested reviewers

  • slin1237

Poem

🐰 I padded the canvas, neat and bright,
Wrote pixels in planes by morning light,
Block by block the colors slide,
Zero-copy bytes now take the ride—
Hop, hop, tensors hum tonight. 🥕

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the three main optimizations: serialize_pixel_values, tensor conversion (build_planar_tensor), and pad fusion (pad_and_normalize_to_tensor), all of which are the primary focus of this performance-focused PR.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/multimodal-profiling

Comment @coderabbitai help to get the list of available commands and usage tips.

Comment thread crates/multimodal/Cargo.toml Outdated

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces several performance optimizations for multimodal image processing, including the fusion of padding and normalization steps to eliminate intermediate allocations and the use of zero-copy reinterpretation for pixel serialization. The RGB deinterleaving process was also optimized for auto-vectorization using block-based processing. Feedback suggests further optimizing the hot path by moving eager allocations for debug logging into the logging macros and extending the block-based vectorization to the newly added fused padding function to ensure consistent performance gains.

Comment thread model_gateway/src/routers/grpc/multimodal.rs Outdated
Comment thread crates/multimodal/src/vision/transforms.rs Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@crates/multimodal/src/vision/transforms.rs`:
- Around line 377-426: Add a focused unit test for pad_and_normalize_to_tensor
that constructs a small known DynamicImage (e.g., 2x2 RGB with explicit bytes),
calls pad_and_normalize_to_tensor with a larger canvas (e.g., 4x4) and known
mean/std, then asserts: (1) the padded region values equal the computed pad_val
(bias) per channel, and (2) the image region values equal the per-pixel
normalized values computed as (pixel/255 - mean) / std (use the same fused
scale/bias math used in the function). Place the test as a #[test] (e.g.,
test_pad_and_normalize_to_tensor) near other vision tests, and compare
elementwise values from the returned Array3::from_shape_vec output for channels
0/1/2 to the expected floats with a small epsilon for float equality.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: aaf9e62e-cd3b-40ed-8083-38d9618e517d

📥 Commits

Reviewing files that changed from the base of the PR and between 041017f and 1c91aca.

📒 Files selected for processing (5)
  • Cargo.toml
  • crates/multimodal/Cargo.toml
  • crates/multimodal/src/vision/processors/llama4_vision.rs
  • crates/multimodal/src/vision/transforms.rs
  • model_gateway/src/routers/grpc/multimodal.rs

Comment thread crates/multimodal/src/vision/transforms.rs Outdated
Signed-off-by: Chang Su <chang.s.su@oracle.com>
@github-actions github-actions Bot removed the dependencies Dependency updates label Apr 1, 2026
… llama4, remove timing

- Extract `deinterleave_rgb_to_planes` as shared public helper with
  8-pixel block optimization, used by both `build_planar_tensor` and
  Llama4's `pad_and_normalize_to_tensor`
- Move `pad_and_normalize_to_tensor` from transforms.rs to
  llama4_vision.rs since it's only used there
- Remove eager `image_sizes` allocation from hot path (review feedback)
- Remove timing instrumentation from gateway multimodal handler

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
Comment thread crates/multimodal/src/vision/processors/llama4_vision.rs
Comment thread model_gateway/src/routers/grpc/multimodal.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 688-699: Add a byte-for-byte regression unit test that verifies
the fast little-endian path (the branch using bytemuck::cast_slice on
pixel_slice) produces identical output to the reference path (the
pixel_slice.iter().flat_map(|v| v.to_le_bytes()).collect()) for at least one
standard-layout tensor and one non-standard-layout case; the test should
construct a known &[f32] input (and a non-contiguous or differently laid-out
variant), call the code path that yields the Vec<u8> from bytemuck::cast_slice
and compare it byte-for-byte to the Vec<u8> produced by mapping each f32 via
to_le_bytes, failing the test if any byte differs.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: a71dbf87-a565-42be-844d-7f1ff6bb2665

📥 Commits

Reviewing files that changed from the base of the PR and between 56f7727 and dc5b579.

📒 Files selected for processing (1)
  • model_gateway/src/routers/grpc/multimodal.rs

Comment thread model_gateway/src/routers/grpc/multimodal.rs
@CatherineSue
CatherineSue merged commit 687319f into main Apr 1, 2026
45 checks passed
@CatherineSue
CatherineSue deleted the perf/multimodal-profiling branch April 1, 2026 18:22
smfirmin pushed a commit to smfirmin/smg that referenced this pull request Apr 2, 2026
…sion, and pad fusion (smg-project#1012)

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant