Conversation
|
Warning Rate limit exceeded
⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughRefactored Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
Hi @slin1237, the DCO sign-off check has failed. All commits must include a To fix existing commits: # Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-leaseTo sign off future commits automatically:
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3a04edd14c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| .collect::<Result<Vec<_>, TransformError>>()?; | ||
|
|
||
| // Merge results sequentially to preserve image order | ||
| let total_patch_floats: usize = per_image_results.iter().map(|(p, _, _, _)| p.len()).sum(); | ||
| let mut all_patches: Vec<f32> = Vec::with_capacity(total_patch_floats); |
There was a problem hiding this comment.
Avoid duplicating full patch buffers in preprocess
Collecting per_image_results stores every image's local_patches in memory at once, and then Vec::with_capacity(total_patch_floats) allocates a second full buffer before merge. For large batches/high-resolution images this roughly doubles peak memory versus the previous streaming append path and can trigger OOMs in production inference workers. This regression is introduced by the parallel refactor; consider a flattening strategy that consumes per-image vectors without preallocating a second full-sized buffer (or writes into final storage by offset).
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Code Review
This pull request introduces parallel image preprocessing in the QwenVLProcessorBase by integrating the rayon library. The implementation refactors the resizing, normalization, and patchification steps to execute in parallel, followed by a sequential aggregation phase to maintain the original image order. A recommendation was made to refactor the final result collection into a more functional fold operation to enhance code clarity and follow idiomatic Rust patterns.
| let mut all_patches: Vec<f32> = Vec::with_capacity(total_patch_floats); | ||
| let mut patches_per_image: Vec<i64> = Vec::with_capacity(images.len()); | ||
| let mut grid_thw_data = Vec::with_capacity(images.len() * 3); | ||
| let mut num_img_tokens = Vec::with_capacity(images.len()); | ||
|
|
||
| for image in images { | ||
| let (w, h) = image.dimensions(); | ||
| let (target_h, target_w) = self.smart_resize(h as usize, w as usize)?; | ||
|
|
||
| // Resize to the image's own target size (skip if dimensions match) | ||
| let (tw32, th32) = (target_w as u32, target_h as u32); | ||
| let needs_resize = config.do_resize.unwrap_or(true) && (w != tw32 || h != th32); | ||
| let resized; | ||
| let img_ref = if needs_resize { | ||
| resized = resize(image, tw32, th32, filter); | ||
| &resized | ||
| } else { | ||
| image | ||
| }; | ||
|
|
||
| // Grid dimensions based on the target size | ||
| let (grid_t, grid_h, grid_w) = self.calculate_grid_thw(target_h, target_w, 1); | ||
| grid_thw_data.push(grid_t as i64); | ||
| grid_thw_data.push(grid_h as i64); | ||
| grid_thw_data.push(grid_w as i64); | ||
|
|
||
| let num_patches = grid_t * grid_h * grid_w; | ||
| let tokens = self.calculate_tokens_from_grid(grid_t, grid_h, grid_w); | ||
| num_img_tokens.push(tokens); | ||
|
|
||
| // Convert to tensor [C, H, W] and normalize in one fused pass | ||
| let tensor = if config.do_normalize.unwrap_or(true) { | ||
| to_tensor_and_normalize(img_ref, &mean, &std) | ||
| } else { | ||
| to_tensor(img_ref) | ||
| }; | ||
|
|
||
| // Patchify directly into all_patches to avoid intermediate Vec + copy | ||
| self.patchify_into(&tensor, grid_t, grid_h, grid_w, &mut all_patches)?; | ||
| for (local_patches, num_patches, grid_thw, tokens) in per_image_results { | ||
| all_patches.extend(local_patches); | ||
| patches_per_image.push(num_patches as i64); | ||
| grid_thw_data.extend_from_slice(&grid_thw); | ||
| num_img_tokens.push(tokens); | ||
| } |
There was a problem hiding this comment.
For improved code clarity and to follow a more functional style, you can replace the multiple mutable Vec declarations and the subsequent for loop with a single Iterator::fold operation. This consolidates the result aggregation logic into one expression, which can make the data flow easier to follow.
let (all_patches, patches_per_image, grid_thw_data, num_img_tokens) =
per_image_results.into_iter().fold(
(
Vec::with_capacity(total_patch_floats),
Vec::with_capacity(images.len()),
Vec::with_capacity(images.len() * 3),
Vec::with_capacity(images.len()),
),
|(mut all_patches, mut patches_per_image, mut grid_thw_data, mut num_img_tokens),
(local_patches, num_patches, grid_thw, tokens)| {
all_patches.extend(local_patches);
patches_per_image.push(num_patches as i64);
grid_thw_data.extend_from_slice(&grid_thw);
num_img_tokens.push(tokens);
(all_patches, patches_per_image, grid_thw_data, num_img_tokens)
},
);There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@crates/multimodal/src/vision/processors/qwen_vl_base.rs`:
- Around line 372-381: The parallel path currently creates and retains
per_image_results (holding local_patches) and then allocates all_patches and
copies every float again, doubling peak memory and incurring a full copy;
instead compute per-image patch counts/offsets first, preallocate a single
all_patches Vec<f32> of the total size, and have each parallel task call a
variant of patchify_into that writes directly into its disjoint slice of that
preallocated buffer (use per-image index/offset from the enumerated parallel
iterator); update or overload patchify_into (or add patchify_into_buffer) to
accept &mut [f32] plus the write offset so threads write in-place into
all_patches (preserving image order by using the same indexing used for
calculating offsets) and eliminate keeping local_patches/per_image_results.
- Around line 345-373: The code currently keeps img_ref at the original image
when do_resize is false even if smart_resize computed a different target size,
causing patchify_into/calculate_grid_thw to see inconsistent dimensions; update
the branch around needs_resize (and where img_ref is chosen) to detect the case
where !do_resize but (w != tw32 || h != th32) and immediately return a
recoverable error (use an existing TransformError variant or add one, e.g.,
TransformError::InvalidImageDimensions) instead of proceeding, and ensure any
places using expect/unreachable switch to ok_or(...) where appropriate to avoid
panics (refer to smart_resize, calculate_grid_thw, patchify_into,
to_tensor/to_tensor_and_normalize, and the do_resize flag).
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: bb5a19b7-e16f-4971-b172-16240cc0ddbe
📒 Files selected for processing (2)
crates/multimodal/Cargo.tomlcrates/multimodal/src/vision/processors/qwen_vl_base.rs
| let (w, h) = image.dimensions(); | ||
| let (target_h, target_w) = self.smart_resize(h as usize, w as usize)?; | ||
|
|
||
| // Resize to the image's own target size (skip if dimensions match) | ||
| let (tw32, th32) = (target_w as u32, target_h as u32); | ||
| let needs_resize = do_resize && (w != tw32 || h != th32); | ||
| let resized; | ||
| let img_ref = if needs_resize { | ||
| resized = resize(image, tw32, th32, filter); | ||
| &resized | ||
| } else { | ||
| image | ||
| }; | ||
|
|
||
| // Grid dimensions based on the target size | ||
| let (grid_t, grid_h, grid_w) = self.calculate_grid_thw(target_h, target_w, 1); | ||
| let num_patches = grid_t * grid_h * grid_w; | ||
| let tokens = self.calculate_tokens_from_grid(grid_t, grid_h, grid_w); | ||
|
|
||
| // Convert to tensor [C, H, W] and normalize in one fused pass | ||
| let tensor = if do_normalize { | ||
| to_tensor_and_normalize(img_ref, &mean, &std) | ||
| } else { | ||
| to_tensor(img_ref) | ||
| }; | ||
|
|
||
| // Patchify into a local buffer | ||
| let mut local_patches = Vec::new(); | ||
| self.patchify_into(&tensor, grid_t, grid_h, grid_w, &mut local_patches)?; |
There was a problem hiding this comment.
Return an error when do_resize is false but the image still needs resizing.
smart_resize() always computes target_h/target_w, but this branch keeps img_ref at the original size when resizing is disabled. The later calculate_grid_thw(target_*, ...) / patchify_into(...) path then sees inconsistent tensor and grid dimensions, which can silently crop oversized inputs and panic on undersized ones.
🛡️ Minimal fix
let (tw32, th32) = (target_w as u32, target_h as u32);
- let needs_resize = do_resize && (w != tw32 || h != th32);
+ let needs_resize = w != tw32 || h != th32;
+ if !do_resize && needs_resize {
+ return Err(TransformError::InvalidShape {
+ expected: format!(
+ "dimensions already match smart_resize output ({target_h}x{target_w}) when do_resize=false"
+ ),
+ actual: vec![h as usize, w as usize],
+ });
+ }
let resized;
- let img_ref = if needs_resize {
+ let img_ref = if do_resize && needs_resize {
resized = resize(image, tw32, th32, filter);
&resized
} else {
image
};🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@crates/multimodal/src/vision/processors/qwen_vl_base.rs` around lines 345 -
373, The code currently keeps img_ref at the original image when do_resize is
false even if smart_resize computed a different target size, causing
patchify_into/calculate_grid_thw to see inconsistent dimensions; update the
branch around needs_resize (and where img_ref is chosen) to detect the case
where !do_resize but (w != tw32 || h != th32) and immediately return a
recoverable error (use an existing TransformError variant or add one, e.g.,
TransformError::InvalidImageDimensions) instead of proceeding, and ensure any
places using expect/unreachable switch to ok_or(...) where appropriate to avoid
panics (refer to smart_resize, calculate_grid_thw, patchify_into,
to_tensor/to_tensor_and_normalize, and the do_resize flag).
| let mut local_patches = Vec::new(); | ||
| self.patchify_into(&tensor, grid_t, grid_h, grid_w, &mut local_patches)?; | ||
|
|
||
| Ok((local_patches, num_patches, [grid_t as i64, grid_h as i64, grid_w as i64], tokens)) | ||
| }) | ||
| .sum(); | ||
| let mut all_patches: Vec<f32> = Vec::with_capacity(estimated_total); | ||
| .collect::<Result<Vec<_>, TransformError>>()?; | ||
|
|
||
| // Merge results sequentially to preserve image order | ||
| let total_patch_floats: usize = per_image_results.iter().map(|(p, _, _, _)| p.len()).sum(); | ||
| let mut all_patches: Vec<f32> = Vec::with_capacity(total_patch_floats); |
There was a problem hiding this comment.
This parallel path now materializes the patch buffer twice.
Collecting per_image_results keeps every local_patches alive until the merge, and the merge then allocates all_patches and copies every float again. For max-sized inputs that is tens of MiB per image, so large batches now pay roughly 2× peak patch-buffer memory plus an extra full copy. Please consider precomputing per-image lengths/offsets and patchifying directly into disjoint slices of one preallocated output buffer instead.
Also applies to: 386-387
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@crates/multimodal/src/vision/processors/qwen_vl_base.rs` around lines 372 -
381, The parallel path currently creates and retains per_image_results (holding
local_patches) and then allocates all_patches and copies every float again,
doubling peak memory and incurring a full copy; instead compute per-image patch
counts/offsets first, preallocate a single all_patches Vec<f32> of the total
size, and have each parallel task call a variant of patchify_into that writes
directly into its disjoint slice of that preallocated buffer (use per-image
index/offset from the enumerated parallel iterator); update or overload
patchify_into (or add patchify_into_buffer) to accept &mut [f32] plus the write
offset so threads write in-place into all_patches (preserving image order by
using the same indexing used for calculating offsets) and eliminate keeping
local_patches/per_image_results.
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@crates/multimodal/src/vision/processors/qwen_vl_base.rs`:
- Around line 413-420: The conditional in the sequential path that sets img_ref
(using do_resize, needs_resize, resized, resize()) has the same bug as the
parallel branch: when do_resize is false but the incoming image size differs
from smart_resize() output, downstream patchify_into() gets mismatched
tensor/grid sizes; update the sequential branch to compute the target size (as
smart_resize does), and always ensure the image is resized to that target when
dimensions differ (i.e., replace the current do_resize-only check with a
size-comparison check like in the parallel fix), so that img_ref passed to
patchify_into() is guaranteed to match the expected grid dimensions.
- Around line 341-443: The image-processing logic is duplicated between the
parallel branch and the sequential branch; extract it into a helper (e.g., fn
process_single_image(&self, image: &DynamicImage, do_resize: bool, do_normalize:
bool, filter: FilterType, mean: &[f64;3], std: &[f64;3]) -> Result<(Vec<f32>,
usize, [i64;3], usize), TransformError>) that calls self.smart_resize, applies
the do_resize check + resize, calls self.calculate_grid_thw,
self.calculate_tokens_from_grid, chooses to_tensor_or to_tensor_and_normalize,
and (optionally) invokes self.patchify_into or returns the patch buffer so
callers can decide zero-copy; then replace the duplicated blocks in the parallel
path (where per_image_results are built) and the sequential loop (where you
currently call self.patchify_into directly) to call this helper so behavior
(including the do_resize logic) is consistent.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: f4c2d80f-a1fe-4e07-bfb4-5c5649be03fa
📒 Files selected for processing (1)
crates/multimodal/src/vision/processors/qwen_vl_base.rs
| let (all_patches, patches_per_image, grid_thw_data, num_img_tokens) = if images.len() > 1 { | ||
| // Parallel path: process each image independently, then merge | ||
| let per_image_results: Vec<_> = images | ||
| .par_iter() | ||
| .map(|image| { | ||
| let (w, h) = image.dimensions(); | ||
| let (target_h, target_w) = self.smart_resize(h as usize, w as usize)?; | ||
|
|
||
| let (tw32, th32) = (target_w as u32, target_h as u32); | ||
| let needs_resize = do_resize && (w != tw32 || h != th32); | ||
| let resized; | ||
| let img_ref = if needs_resize { | ||
| resized = resize(image, tw32, th32, filter); | ||
| &resized | ||
| } else { | ||
| image | ||
| }; | ||
|
|
||
| let (grid_t, grid_h, grid_w) = self.calculate_grid_thw(target_h, target_w, 1); | ||
| let num_patches = grid_t * grid_h * grid_w; | ||
| let tokens = self.calculate_tokens_from_grid(grid_t, grid_h, grid_w); | ||
|
|
||
| let tensor = if do_normalize { | ||
| to_tensor_and_normalize(img_ref, &mean, &std) | ||
| } else { | ||
| to_tensor(img_ref) | ||
| }; | ||
|
|
||
| let mut local_patches = Vec::new(); | ||
| self.patchify_into(&tensor, grid_t, grid_h, grid_w, &mut local_patches)?; | ||
|
|
||
| Ok((local_patches, num_patches, [grid_t as i64, grid_h as i64, grid_w as i64], tokens)) | ||
| }) | ||
| .collect::<Result<Vec<_>, TransformError>>()?; | ||
|
|
||
| let total_patch_floats: usize = | ||
| per_image_results.iter().map(|(p, _, _, _)| p.len()).sum(); | ||
| let mut all_patches: Vec<f32> = Vec::with_capacity(total_patch_floats); | ||
| let mut patches_per_image: Vec<i64> = Vec::with_capacity(images.len()); | ||
| let mut grid_thw_data = Vec::with_capacity(images.len() * 3); | ||
| let mut num_img_tokens = Vec::with_capacity(images.len()); | ||
|
|
||
| for (local_patches, num_patches, grid_thw, tokens) in per_image_results { | ||
| all_patches.extend(local_patches); | ||
| patches_per_image.push(num_patches as i64); | ||
| grid_thw_data.extend_from_slice(&grid_thw); | ||
| num_img_tokens.push(tokens); | ||
| } | ||
|
|
||
| (all_patches, patches_per_image, grid_thw_data, num_img_tokens) | ||
| } else { | ||
| // Sequential path: patchify directly into output buffer (zero-copy) | ||
| let estimated_total: usize = images | ||
| .iter() | ||
| .map(|img| { | ||
| let (w, h) = img.dimensions(); | ||
| (w as usize * h as usize) | ||
| / (self.config.merge_size * self.config.merge_size) | ||
| * patch_features | ||
| / (patch_size * patch_size) | ||
| }) | ||
| .sum(); | ||
| let mut all_patches: Vec<f32> = Vec::with_capacity(estimated_total); | ||
| let mut patches_per_image: Vec<i64> = Vec::with_capacity(images.len()); | ||
| let mut grid_thw_data = Vec::with_capacity(images.len() * 3); | ||
| let mut num_img_tokens = Vec::with_capacity(images.len()); | ||
|
|
||
| for image in images { | ||
| let (w, h) = image.dimensions(); | ||
| let (target_h, target_w) = self.smart_resize(h as usize, w as usize)?; | ||
|
|
||
| let (tw32, th32) = (target_w as u32, target_h as u32); | ||
| let needs_resize = do_resize && (w != tw32 || h != th32); | ||
| let resized; | ||
| let img_ref = if needs_resize { | ||
| resized = resize(image, tw32, th32, filter); | ||
| &resized | ||
| } else { | ||
| image | ||
| }; | ||
|
|
||
| let (grid_t, grid_h, grid_w) = self.calculate_grid_thw(target_h, target_w, 1); | ||
| grid_thw_data.push(grid_t as i64); | ||
| grid_thw_data.push(grid_h as i64); | ||
| grid_thw_data.push(grid_w as i64); | ||
|
|
||
| let num_patches = grid_t * grid_h * grid_w; | ||
| let tokens = self.calculate_tokens_from_grid(grid_t, grid_h, grid_w); | ||
| num_img_tokens.push(tokens); | ||
|
|
||
| let tensor = if do_normalize { | ||
| to_tensor_and_normalize(img_ref, &mean, &std) | ||
| } else { | ||
| to_tensor(img_ref) | ||
| }; | ||
|
|
||
| // Patchify directly into all_patches to avoid intermediate Vec + copy | ||
| self.patchify_into(&tensor, grid_t, grid_h, grid_w, &mut all_patches)?; | ||
| patches_per_image.push(num_patches as i64); | ||
| } | ||
|
|
||
| (all_patches, patches_per_image, grid_thw_data, num_img_tokens) | ||
| }; |
There was a problem hiding this comment.
🧹 Nitpick | 🔵 Trivial
Consider extracting shared image-processing logic to reduce duplication.
The per-image processing logic (smart_resize → conditional resize → grid calculation → tensor conversion) is nearly identical between the parallel path (lines 346-372) and sequential path (lines 408-439). A helper method could reduce duplication and ensure fixes (like the do_resize issue) are applied consistently.
♻️ Sketch of a possible helper
fn process_single_image(
&self,
image: &DynamicImage,
do_resize: bool,
do_normalize: bool,
filter: FilterType,
mean: &[f64; 3],
std: &[f64; 3],
) -> Result<(Array3<f32>, usize, usize, usize, [i64; 3], usize), TransformError> {
let (w, h) = image.dimensions();
let (target_h, target_w) = self.smart_resize(h as usize, w as usize)?;
// ... resize/normalize/grid logic ...
Ok((tensor, grid_t, grid_h, grid_w, grid_thw, tokens))
}Then both paths call this helper—the parallel path collects and merges, while the sequential path patchifies directly.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@crates/multimodal/src/vision/processors/qwen_vl_base.rs` around lines 341 -
443, The image-processing logic is duplicated between the parallel branch and
the sequential branch; extract it into a helper (e.g., fn
process_single_image(&self, image: &DynamicImage, do_resize: bool, do_normalize:
bool, filter: FilterType, mean: &[f64;3], std: &[f64;3]) -> Result<(Vec<f32>,
usize, [i64;3], usize), TransformError>) that calls self.smart_resize, applies
the do_resize check + resize, calls self.calculate_grid_thw,
self.calculate_tokens_from_grid, chooses to_tensor_or to_tensor_and_normalize,
and (optionally) invokes self.patchify_into or returns the patch buffer so
callers can decide zero-copy; then replace the duplicated blocks in the parallel
path (where per_image_results are built) and the sequential loop (where you
currently call self.patchify_into directly) to call this helper so behavior
(including the do_resize logic) is consistent.
| let needs_resize = do_resize && (w != tw32 || h != th32); | ||
| let resized; | ||
| let img_ref = if needs_resize { | ||
| resized = resize(image, tw32, th32, filter); | ||
| &resized | ||
| } else { | ||
| image | ||
| }; |
There was a problem hiding this comment.
Same do_resize=false dimension mismatch applies here.
The sequential path has the identical logic issue flagged for the parallel path: when do_resize=false but the image dimensions don't match smart_resize() output, the tensor dimensions will be inconsistent with the grid passed to patchify_into().
When addressing the fix proposed in the earlier review comment, apply it to both paths.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@crates/multimodal/src/vision/processors/qwen_vl_base.rs` around lines 413 -
420, The conditional in the sequential path that sets img_ref (using do_resize,
needs_resize, resized, resize()) has the same bug as the parallel branch: when
do_resize is false but the incoming image size differs from smart_resize()
output, downstream patchify_into() gets mismatched tensor/grid sizes; update the
sequential branch to compute the target size (as smart_resize does), and always
ensure the image is resized to that target when dimensions differ (i.e., replace
the current do_resize-only check with a size-comparison check like in the
parallel fix), so that img_ref passed to patchify_into() is guaranteed to match
the expected grid dimensions.
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
…d thread pool overhead Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
ea9de7c to
f46ebff
Compare
Benchmark ResultsBaseline (main) vs this PR (with batch-size threshold fix):
Single images use the original zero-copy sequential path (no rayon overhead). Batches >1 use Note: Initial implementation (without threshold) caused 5.3x regression on single images. Fixed by gating |
MMMU Validation — Qwen3-VL-8B
No accuracy degradation. The slight latency increase is expected since MMMU sends single images sequentially (rayon parallelism only activates for batch > 1). The real-world benefit is in multi-image batch scenarios where batch3/5 see 74-76% speedup. |
Summary
rayondependency tollm-multimodalcrateQwenVLProcessorBase::preprocessto userayon::par_iter(), processing each image's resize/normalize/patchify pipeline in parallelTest plan
cargo test -p llm-multimodal-- all 135 unit tests passcargo test -p llm-multimodal -- vision_golden-- all 81 golden tests pass (bit-exact output preserved)Summary by CodeRabbit