perf(multimodal): optimize preprocessing serialization, tensor conversion, and pad fusion - #1012
Conversation
…rumentation Signed-off-by: Chang Su <chang.s.su@oracle.com>
…r, and pad+tensor fusion Three optimizations based on perf profiling of MMMU benchmark: 1. serialize_pixel_values: replace per-element flat_map(to_le_bytes) with bytemuck::cast_slice zero-copy reinterpret. Eliminates FlattenCompat which was #1 CPU hotspot (24-31% of SMG CPU). 2. build_planar_tensor: restructure RGB deinterleave loop into 8-pixel blocks for better auto-vectorization. Reduces tensor conversion from 22% to 5% of CPU. 3. Llama4 pad_and_normalize_to_tensor: fuse pad_image + to_tensor into one pass, eliminating intermediate padded RgbImage allocation. Removes 13% CPU overhead from image::overlay + get_pixel. Combined effect on Llama4 preprocessing: avg 25.5ms -> 16.1ms (-37%), median 21.6ms -> 9.8ms (-55%). MMMU accuracy unchanged (0.417 -> 0.428). Signed-off-by: Chang Su <chang.s.su@oracle.com>
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
📝 WalkthroughWalkthroughFuses image padding and normalization into a single tensor construction path, exposes and centralizes RGB deinterleaving, and optimizes pixel byte serialization on little-endian targets. Changes
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces several performance optimizations for multimodal image processing, including the fusion of padding and normalization steps to eliminate intermediate allocations and the use of zero-copy reinterpretation for pixel serialization. The RGB deinterleaving process was also optimized for auto-vectorization using block-based processing. Feedback suggests further optimizing the hot path by moving eager allocations for debug logging into the logging macros and extending the block-based vectorization to the newly added fused padding function to ensure consistent performance gains.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@crates/multimodal/src/vision/transforms.rs`:
- Around line 377-426: Add a focused unit test for pad_and_normalize_to_tensor
that constructs a small known DynamicImage (e.g., 2x2 RGB with explicit bytes),
calls pad_and_normalize_to_tensor with a larger canvas (e.g., 4x4) and known
mean/std, then asserts: (1) the padded region values equal the computed pad_val
(bias) per channel, and (2) the image region values equal the per-pixel
normalized values computed as (pixel/255 - mean) / std (use the same fused
scale/bias math used in the function). Place the test as a #[test] (e.g.,
test_pad_and_normalize_to_tensor) near other vision tests, and compare
elementwise values from the returned Array3::from_shape_vec output for channels
0/1/2 to the expected floats with a small epsilon for float equality.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: aaf9e62e-cd3b-40ed-8083-38d9618e517d
📒 Files selected for processing (5)
Cargo.tomlcrates/multimodal/Cargo.tomlcrates/multimodal/src/vision/processors/llama4_vision.rscrates/multimodal/src/vision/transforms.rsmodel_gateway/src/routers/grpc/multimodal.rs
Signed-off-by: Chang Su <chang.s.su@oracle.com>
… llama4, remove timing - Extract `deinterleave_rgb_to_planes` as shared public helper with 8-pixel block optimization, used by both `build_planar_tensor` and Llama4's `pad_and_normalize_to_tensor` - Move `pad_and_normalize_to_tensor` from transforms.rs to llama4_vision.rs since it's only used there - Remove eager `image_sizes` allocation from hot path (review feedback) - Remove timing instrumentation from gateway multimodal handler Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@model_gateway/src/routers/grpc/multimodal.rs`:
- Around line 688-699: Add a byte-for-byte regression unit test that verifies
the fast little-endian path (the branch using bytemuck::cast_slice on
pixel_slice) produces identical output to the reference path (the
pixel_slice.iter().flat_map(|v| v.to_le_bytes()).collect()) for at least one
standard-layout tensor and one non-standard-layout case; the test should
construct a known &[f32] input (and a non-contiguous or differently laid-out
variant), call the code path that yields the Vec<u8> from bytemuck::cast_slice
and compare it byte-for-byte to the Vec<u8> produced by mapping each f32 via
to_le_bytes, failing the test if any byte differs.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: a71dbf87-a565-42be-844d-7f1ff6bb2665
📒 Files selected for processing (1)
model_gateway/src/routers/grpc/multimodal.rs
…sion, and pad fusion (smg-project#1012) Signed-off-by: Chang Su <chang.s.su@oracle.com>
Description
Problem
Perf profiling of the multimodal preprocessing pipeline during real MMMU benchmarks (900 images, Qwen3-VL-8B + Llama4-Scout-17B) revealed three CPU hotspots in SMG that together consumed 66% of all SMG CPU time:
serialize_pixel_values(31% CPU) —FlattenCompat::folditerating per-elementf32::to_le_bytes()to serialize pixel tensors for gRPCbuild_planar_tensor(22% CPU) — scalar RGB deinterleave loop that the compiler couldn't auto-vectorizepad_image(13% CPU, Llama4 only) —image::overlay+get_pixelcreating an intermediate paddedRgbImageSolution
Three targeted optimizations based on perf flamegraph analysis:
serialize_pixel_values→bytemuck::cast_slice: Zero-copy reinterpret&[f32]as&[u8]on little-endian, replacing per-elementflat_map(to_le_bytes).collect(). Eliminates the the top CPU hotspot entirely.build_planar_tensorblock processing: Restructure the RGB deinterleave loop into fixed 8-pixel blocks so the compiler can unroll and auto-vectorize the stride-3 gather pattern.pad_and_normalize_to_tensorfusion (Llama4): Build the padded f32 tensor directly from resized RGB bytes, skipping the intermediate paddedRgbImageallocation. Pre-fills padding regions with the normalized black value, then overwrites the image region row-by-row.Changes
model_gateway/src/routers/grpc/multimodal.rs— Replaceflat_map(to_le_bytes)withbytemuck::cast_sliceinserialize_pixel_valuescrates/multimodal/src/vision/transforms.rs— Restructurebuild_planar_tensorinto 8-pixel blocks; addpad_and_normalize_to_tensorfused functioncrates/multimodal/src/vision/processors/llama4_vision.rs— Use fusedpad_and_normalize_to_tensor, remove deadpad_imagemethod(Removed in later commits)Cargo.toml— Add[profile.profiling]for future perf profiling with debug symbolsTest Plan
Profiling methodology
Profiled with
perf record -F 4999 -g --call-graph fpduring real MMMU benchmark runs. To get meaningful CPU samples (SMG is >99% idle during inference), sent 500 concurrent requests withmax_tokens=1.Built with a new
[profile.profiling]that inherits release but keeps debug symbols:Commands used
Perf profile: before vs after (Llama4-Scout, 66K samples)
FlattenCompat::fold(serialize)serialize_pixel_values(outer)build_planar_tensorimage::get_pixel(pad)image::overlay(pad)ImageBuffer::from_pixel(pad)ImageBuffer::get_pixel_mut(pad)ndarray::assign(tile)Preprocessing wall-clock timing (Llama4-Scout, 900 MMMU images)
Perf profile confirms same hotspots on Qwen3-VL (209 samples)
FlattenCompat::foldbuild_planar_tensorfdeflate::DecompressorVec::resizeserialize_pixel_valuesMMMU accuracy (Llama4-Scout, 900 images)
No regression — scores are within noise (optimizations don't change output values).
Correctness
bytemuck::cast_sliceproduces bit-identical output toto_le_bytes()on little-endian (guarded by#[cfg(target_endian = "little")])build_planar_tensorblock processing: same arithmetic, just restructured for auto-vectorizationpad_and_normalize_to_tensor: same normalized values, verified by pre-fill withbias(normalized black) + row-by-row overwriteEnd-to-end MMMU wall-clock time (Llama4-Scout, 900 images)
Preprocessing is <0.01% of total request time — model inference dominates. The ~8s saved across 900 images matches the per-image improvement (900 × 9.4ms ≈ 8.5s). The real benefit is reduced SMG CPU utilization under concurrent load.
Baseline
Optimized
Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit