Skip to content

refactor(multimodal): split registry.rs into per-model spec modules - #593

Merged
CatherineSue merged 1 commit into
mainfrom
chang/split-registry
Mar 3, 2026
Merged

CatherineSue merged 1 commit into
mainfrom
chang/split-registry

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 3, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

multimodal/src/registry.rs was a 975-line monolith containing the ModelProcessorSpec trait, all per-model spec implementations, ModelRegistry, error types, and test helpers. This made it difficult to navigate and maintain as new model specs are added.

Solution

Extract the monolithic file into a registry/ module directory with one file per model spec, keeping co-located tests alongside each implementation.

Changes

  • traits.rs: ModelProcessorSpec trait, ModelMetadata, ModelRegistryError, image_sizes_hw helper
  • mod.rs: re-exports, ModelRegistry, LazySpec, test_helpers
  • llama4.rs: Llama 4 spec + tests
  • llava.rs: LLaVA spec + tests
  • phi3_v.rs: Phi-3 Vision spec + tests
  • qwen3_vl.rs: Qwen3-VL spec + tests
  • qwen_vl.rs: Qwen-VL (Qwen2-VL) spec + tests
  • Uses crate::registry:: imports instead of super::super:: chains

Test Plan

cargo check -p llm-multimodal && cargo test -p llm-multimodal  # all tests pass
pre-commit run --all-files  # passes
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Refactor

    • Reorganized the multimodal model registry into a modular architecture with trait-based processor specifications for improved maintainability.
  • New Features

    • Added support for LLaVA, Phi3-V, Qwen3-VL, Qwen-VL, and Llama4 vision models with dedicated processor implementations.
    • Introduced lazy initialization for efficient model spec loading.

Extract the monolithic registry.rs (975 lines) into a registry/ module:
- traits.rs: ModelProcessorSpec trait, ModelMetadata, error types
- mod.rs: re-exports, ModelRegistry, LazySpec, test_helpers
- llama4.rs, llava.rs, phi3_v.rs, qwen3_vl.rs, qwen_vl.rs: per-model specs with co-located tests

Uses crate::registry:: imports instead of super::super:: chains.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@CatherineSue
CatherineSue requested a review from slin1237 as a code owner March 3, 2026 19:46
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added the multimodal Multimodal crate changes label Mar 3, 2026
@coderabbitai

coderabbitai Bot commented Mar 3, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

The multimodal registry module is being refactored from a single monolithic file into a modular architecture. The existing registry.rs file is split into traits.rs (core types and traits), mod.rs (registry container), and separate spec files for each model implementation (Llama4, Llava, Phi3V, Qwen3VL, QwenVL).

Changes

Cohort / File(s) Summary
Registry Restructuring
multimodal/src/registry.rs
Entire file removed; monolithic registry module (~975 lines including error types, traits, structs, concrete specs, and tests) being decomposed into modular structure.
Core Traits & Types
multimodal/src/registry/traits.rs
New file introducing foundational types: ModelRegistryError, RegistryResult, ModelMetadata (with token_id and config_u32 methods), ModelProcessorSpec trait, and image_sizes_hw utility function.
Registry Container & Test Utilities
multimodal/src/registry/mod.rs
New file defining ModelRegistry with lazy initialization via LazySpec, lookup method, Default implementation, and pub(super) test_helpers module with TestTokenizer and test data constructors.
Model Specs - Vision-focused
multimodal/src/registry/llava.rs, multimodal/src/registry/phi3_v.rs, multimodal/src/registry/qwen_vl.rs, multimodal/src/registry/qwen3_vl.rs
Four new spec files implementing ModelProcessorSpec for individual models with model-specific token handling, configuration helpers, prompt replacement logic, field layouts, and unit tests.
Model Spec - Advanced Tiling
multimodal/src/registry/llama4.rs
New spec file for Llama4 with advanced image-tiling support: patch/tile size configuration, aspect ratio extraction, multi-tile token sequences with grid separators, and corresponding tests.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

multimodal, tests

Suggested reviewers

  • slin1237
  • key4ng

Poem

🐰 One large file split into pieces so neat,
Registry traits now dance down their own street,
Specs for each model in homes of their own,
Lazy initialization ensures they're well-grown!
From mono to modular, clean code takes flight! ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 27.71% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'refactor(multimodal): split registry.rs into per-model spec modules' accurately describes the main change in the PR, which is a structural reorganization of a monolithic registry file into modular per-model spec files.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/split-registry

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.


Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly refactors the multimodal model registry by deconstructing a single, large file into a more modular and organized directory structure. The change aims to enhance the navigability and long-term maintainability of the codebase, particularly as new model specifications are integrated. This reorganization improves code clarity without altering the existing functional behavior of the model processing logic.

Highlights

  • Code Organization: The monolithic multimodal/src/registry.rs file has been split into a new registry/ module directory for improved structure.
  • Modularization of Model Specifications: Individual files have been created for each model specification (Llama4, LLaVA, Phi-3 Vision, Qwen3-VL, Qwen-VL), enhancing clarity and maintainability.
  • Core Registry Components: The ModelProcessorSpec trait, ModelMetadata, ModelRegistryError, and image_sizes_hw helper have been moved to a new traits.rs module, centralizing common definitions.
  • Registry and Test Helpers: The ModelRegistry, LazySpec, and common test helpers are now consolidated within the mod.rs file of the new registry/ module.
  • Import Path Refinement: Internal imports have been updated to use crate::registry:: for better readability and consistency, replacing verbose super::super:: chains.
Changelog
  • multimodal/src/registry.rs
    • Removed the monolithic registry implementation file.
  • multimodal/src/registry/llama4.rs
    • Added the dedicated Llama4 model specification, including its logic and tests.
  • multimodal/src/registry/llava.rs
    • Added the dedicated LLaVA model specification, including its logic and tests.
  • multimodal/src/registry/mod.rs
    • Introduced the new registry module, managing re-exports, the ModelRegistry structure, and common test helpers.
  • multimodal/src/registry/phi3_v.rs
    • Added the dedicated Phi-3 Vision model specification, including its logic and tests.
  • multimodal/src/registry/qwen3_vl.rs
    • Added the dedicated Qwen3-VL model specification, including its logic and tests.
  • multimodal/src/registry/qwen_vl.rs
    • Added the dedicated Qwen-VL model specification, including its logic and tests.
  • multimodal/src/registry/traits.rs
    • Introduced a new file to define the ModelProcessorSpec trait, ModelMetadata, ModelRegistryError, and the image_sizes_hw helper.
Activity
  • No specific activity (comments, reviews, or progress updates) has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the monolithic registry.rs file into a more maintainable, modular structure, with each model specification now residing in its own file with co-located tests, significantly improving code organization and clarity. However, it preserves several resource exhaustion vulnerabilities. As these are cross-cutting security concerns requiring design decisions on resource limits, they should be addressed in a dedicated pull request. Additionally, my review includes suggestions for minor improvements, such as using more idiomatic Rust constructs and ensuring consistency.

Comment on lines +134 to +166
Ok(aspect_ratios
.iter()
.map(|&(h_tiles, w_tiles)| {
let num_tiles = h_tiles * w_tiles;

let mut tokens = Vec::new();

// <|image_start|>
tokens.push(image_start_id);

// Grid tiles with separators (only for multi-tile images)
if num_tiles > 1 {
for _row in 0..h_tiles {
for col in 0..w_tiles {
tokens.extend(std::iter::repeat_n(patch_token_id, tokens_per_tile));
if col < w_tiles - 1 {
tokens.push(tile_x_sep_id);
}
}
tokens.push(tile_y_sep_id);
}
}

// Global/cover tile: <|image|> + <|patch|> * tokens_per_tile
tokens.push(image_id);
tokens.extend(std::iter::repeat_n(patch_token_id, tokens_per_tile));

// <|image_end|>
tokens.push(image_end_id);

PromptReplacement::sequence(Modality::Image, &placeholder, tokens)
})
.collect())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

This presents a resource exhaustion vulnerability. Addressing this and similar issues across the system, which require design decisions on resource limits, should be done in a dedicated pull request.

                let capacity = if num_tiles > 1 {
                    num_tiles * (tokens_per_tile + 1)
                } else {
                    0
                } + tokens_per_tile + 3;
                let mut tokens = Vec::with_capacity(capacity);
References
  1. Cross-cutting concerns, especially security-related ones that require design decisions, should be addressed in a dedicated pull request rather than being patched within a feature-specific PR.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

Comment on lines +67 to +73
Ok(image_sizes
.iter()
.map(|size| {
let count = Self::tokens_per_image(metadata, *size);
PromptReplacement::repeated(Modality::Image, &token, token_id, count)
})
.collect())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

This presents a resource exhaustion vulnerability. Addressing this and similar issues across the system, which require design decisions on resource limits, should be done in a dedicated pull request.

References
  1. Cross-cutting concerns, especially security-related ones that require design decisions, should be addressed in a dedicated pull request rather than being patched within a feature-specific PR.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

Comment on lines +58 to +62
Ok(preprocessed
.image_sizes
.iter()
.map(|_| PromptReplacement::repeated(Modality::Image, &token, token_id, count))
.collect())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

This presents a resource exhaustion vulnerability. Addressing this and similar issues across the system, which require design decisions on resource limits, should be done in a dedicated pull request.

References
  1. Cross-cutting concerns, especially security-related ones that require design decisions, should be addressed in a dedicated pull request rather than being patched within a feature-specific PR.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

Comment on lines +37 to +45
pub fn lookup<'a>(&'a self, metadata: &ModelMetadata) -> Option<&'a dyn ModelProcessorSpec> {
for spec in &self.specs {
let spec_ref = spec.get();
if spec_ref.matches(metadata) {
return Some(spec_ref);
}
}
None
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This for loop with a manual return can be expressed more concisely and idiomatically using iterator methods like find.

    pub fn lookup<'a>(&'a self, metadata: &ModelMetadata) -> Option<&'a dyn ModelProcessorSpec> {
        self.specs
            .iter()
            .map(|spec| spec.get())
            .find(|spec_ref| spec_ref.matches(metadata))
    }

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

}

#[cfg(test)]
pub(super) mod test_helpers {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using pub(super) makes the test_helpers module visible to the parent module of registry. Since these helpers are only used for tests within the llm-multimodal crate, pub(crate) is more idiomatic and correctly scopes the visibility to the crate level.

pub(crate) mod test_helpers {

Comment on lines +45 to +51
fn find_value<'v>(value: &'v Value, path: &[&str]) -> Option<&'v Value> {
let mut current = value;
for key in path {
current = current.get(*key)?;
}
Some(current)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This function can be implemented more concisely using try_fold, which is well-suited for iterating through a sequence and potentially failing early.

    fn find_value<'v>(value: &'v Value, path: &[&str]) -> Option<&'v Value> {
        path.iter().try_fold(value, |current, key| current.get(key))
    }

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@multimodal/src/registry/llava.rs`:
- Around line 20-25: The Llava processor is saving image_sizes from
img.dimensions() (which returns (width, height)) but image_sizes_hw expects
(height, width), causing a semantic swap; update the Llava processor code that
builds image_sizes / image_sizes_hw (the place calling img.dimensions()) to
store tuples as (height, width) instead of (width, height) so ImageSize
constructions match other processors; ensure any use sites (e.g., where
ImageSize is created for tokens_per_image and the patch_size/patch_size methods)
continue to use ImageSize { height, width } with the corrected ordering.

In `@multimodal/src/registry/mod.rs`:
- Around line 58-63: The _id parameter on LazySpec::new is unused; either remove
it or document its intent—update the signature of LazySpec::new to remove the
_id parameter (and all call sites, e.g. ModelRegistry::new) if not needed, or
add a short comment explaining why _id is kept (e.g., reserved for future
logging/debugging) and prefix it with underscore to suppress warnings; adjust
all invocations that construct LazySpec (notably in ModelRegistry::new) to match
the new signature or retain passing the id if you choose to keep the parameter
with documentation.

In `@multimodal/src/registry/qwen_vl.rs`:
- Around line 38-41: In the matches method of qwen_vl.rs, avoid calling
metadata.model_id.to_ascii_lowercase() twice by assigning the lowercased model
id to a local variable (e.g., let id = metadata.model_id.to_ascii_lowercase())
and then use id.contains("qwen") && id.contains("vl"); update the matches(&self,
metadata: &ModelMetadata) implementation accordingly to reference that local
variable.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 770237d and ab80c64.

📒 Files selected for processing (8)
  • multimodal/src/registry.rs
  • multimodal/src/registry/llama4.rs
  • multimodal/src/registry/llava.rs
  • multimodal/src/registry/mod.rs
  • multimodal/src/registry/phi3_v.rs
  • multimodal/src/registry/qwen3_vl.rs
  • multimodal/src/registry/qwen_vl.rs
  • multimodal/src/registry/traits.rs
💤 Files with no reviewable changes (1)
  • multimodal/src/registry.rs

Comment on lines +20 to +25
fn tokens_per_image(metadata: &ModelMetadata, size: ImageSize) -> usize {
let patch = Self::patch_size(metadata);
let cols = size.width.div_ceil(patch) as usize;
let rows = size.height.div_ceil(patch) as usize;
cols * rows
}

@coderabbitai coderabbitai Bot Mar 3, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# Search for llava processor implementation and image_sizes handling
find . -type f -name "*.rs" | xargs grep -l "llava" | grep -i processor

Repository: lightseekorg/smg

Length of output: 188


🏁 Script executed:

# Look at the llava processor file to see how image_sizes are set
rg -n "image_sizes" multimodal/src/vision/processors/ --type rust -A 3 -B 3

Repository: lightseekorg/smg

Length of output: 20904


🏁 Script executed:

# Check the ImageSize struct definition
rg -n "struct ImageSize" multimodal/src/ --type rust -A 5

Repository: lightseekorg/smg

Length of output: 311


🏁 Script executed:

# Search for any conversion or helper functions related to image size ordering
rg -n "image_sizes_hw|width.*height|height.*width" multimodal/src/registry.rs --type rust -B 2 -A 2

Repository: lightseekorg/smg

Length of output: 126


🏁 Script executed:

# Search for registry.rs file more carefully
find . -name "registry.rs" -o -name "registry" -type d | head -20

Repository: lightseekorg/smg

Length of output: 154


🏁 Script executed:

# Look at the registry llava.rs file mentioned in the review
cat -n multimodal/src/registry/llava.rs | head -50

Repository: lightseekorg/smg

Length of output: 1853


🏁 Script executed:

# Search for image_sizes_hw function
rg -n "image_sizes_hw" multimodal/src/ --type rust -B 2 -A 5

Repository: lightseekorg/smg

Length of output: 2410


🏁 Script executed:

# Check PreprocessedImages struct definition
rg -n "struct PreprocessedImages" multimodal/src/ --type rust -A 10

Repository: lightseekorg/smg

Length of output: 935


Fix image size tuple ordering in Llava processor to match registry expectations.

The Llava processor stores image_sizes from img.dimensions() which returns (width, height) tuples, but image_sizes_hw interprets them as (height, width). This causes dimension swapping when creating ImageSize objects used by tokens_per_image. While the token count calculation remains correct for square images due to multiplication commutativity, the semantically swapped dimensions violate the expected convention used by other processors (Phi4, Pixtral, Llama4) and could break code that relies on the ImageSize fields having the correct semantic meaning.

Update the Llava processor to explicitly store (height, width) tuples like other processors:

Fix in multimodal/src/vision/processors/llava.rs
- let image_sizes: Vec<(u32, u32)> = images.iter().map(|img| img.dimensions()).collect();
+ let image_sizes: Vec<(u32, u32)> = images.iter().map(|img| {
+     let (w, h) = img.dimensions();
+     (h, w)
+ }).collect();
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@multimodal/src/registry/llava.rs` around lines 20 - 25, The Llava processor
is saving image_sizes from img.dimensions() (which returns (width, height)) but
image_sizes_hw expects (height, width), causing a semantic swap; update the
Llava processor code that builds image_sizes / image_sizes_hw (the place calling
img.dimensions()) to store tuples as (height, width) instead of (width, height)
so ImageSize constructions match other processors; ensure any use sites (e.g.,
where ImageSize is created for tokens_per_image and the patch_size/patch_size
methods) continue to use ImageSize { height, width } with the corrected
ordering.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CatherineSue, understood! You're right—this is a refactoring PR focused on code organization, and the image size ordering issue exists in the processor implementation, not in the code you're introducing here.

If you'd like to track this for a future fix, I can open a follow-up issue for it. Just let me know!


🧠 Learnings used
Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 570
File: multimodal/src/registry.rs:133-143
Timestamp: 2026-03-01T06:00:39.427Z
Learning: In repo lightseekorg/smg, the image_sizes field in PreprocessedImages will be standardized to use `(height, width)` tuple ordering across all processors in multimodal/src/vision/processors/. Currently llama4, phi4, and pixtral use `(height, width)` while llava and phi3 use `(width, height)`. The image_sizes_hw helper in multimodal/src/registry.rs correctly interprets tuples as (h, w) for the majority convention.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/src/vision/processors/phi4_vision.rs:486-490
Timestamp: 2026-02-21T02:39:17.064Z
Learning: In Rust files under multimodal/src (e.g., multimodal/src/vision/processors/phi4_vision.rs), prefer returning recoverable errors via ok_or(TransformError::EmptyBatch) (or similar error) in Phi4VisionProcessor::preprocess when a batch is empty, instead of panicking with expect/unreachable. This aligns with the repo guideline to avoid panics in production code (PR `#489`) and lint-only efforts. Apply this pattern to similar preprocessing paths in this module and related code paths where a non-fatal error conveys meaningful failure to callers.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/tests/vision_golden_tests.rs:576-578
Timestamp: 2026-02-21T02:39:51.670Z
Learning: Repo lightseekorg/smg — For PR `#489` (clippy/lint-only), do not replace unwrap() with expect(...) in test files when a file-/crate-level #![expect(clippy::unwrap_used)] is present (e.g., multimodal/tests/vision_golden_tests.rs). Such stylistic swaps are out of scope.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/benches/wasm_middleware_latency.rs:88-91
Timestamp: 2026-02-21T02:37:02.009Z
Learning: Repo: lightseekorg/smg — For clippy-only/enforcement PRs (e.g., PR `#489`), even micro-optimizations (like replacing an async closure with std::future::ready in benches such as model_gateway/benches/wasm_middleware_latency.rs) should be deferred to a follow-up PR rather than included inline.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/tests/vision_golden_tests.rs:402-404
Timestamp: 2026-02-21T02:39:23.481Z
Learning: Repo lightseekorg/smg — For clippy-only/lint-enforcement PRs (e.g., PR `#489`), do not replace unwrap() with expect() across tests/benches when a crate-level `#![expect(clippy::unwrap_used)]` is present. Such per-call swaps are treated as out-of-scope stylistic changes. Example: multimodal/tests/vision_golden_tests.rs.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: mesh/src/crdt.rs:296-299
Timestamp: 2026-02-21T02:36:31.543Z
Learning: Repo lightseekorg/smg — For clippy/lint-only PRs (e.g., PR `#489`), avoid requesting stylistic doc comments when an item is already annotated with #[expect(...)] (e.g., #[expect(dead_code)] on SyncCRDTMap::contains_key in mesh/src/crdt.rs); such style changes are considered out of scope.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/benches/wasm_middleware_latency.rs:1-1
Timestamp: 2026-02-21T02:37:04.633Z
Learning: Repo: lightseekorg/smg — For benchmark files (e.g., model_gateway/benches/*.rs), using a crate-level `#![expect(clippy::unwrap_used, clippy::disallowed_methods)]` is preferred when unwrap/spawn are used throughout. Do not push to per-function scoping; a single crate-level `reason` is acceptable when justification is required.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 588
File: model_gateway/src/routers/grpc/multimodal.rs:453-514
Timestamp: 2026-03-03T18:03:37.820Z
Learning: In repo lightseekorg/smg, backend assembly functions in model_gateway/src/routers/grpc/multimodal.rs (e.g., assemble_sglang, assemble_vllm, assemble_trtllm) are tested via E2E tests rather than unit tests, as unit tests for these functions are not considered worthwhile.

Learnt from: XinyueZhang369
Repo: lightseekorg/smg PR: 399
File: protocols/src/interactions.rs:505-509
Timestamp: 2026-02-19T03:08:50.192Z
Learning: In code reviews for Rust projects using the validator crate (v0.20.0), ensure that custom validation functions for numeric primitive types (e.g., f32, i32, u32, i16, etc.) accept the value by value, not by reference. Example: fn validate(value: f32) { ... }. The validator derive macro has a hardcoded list of numeric types that are passed by value, while all other types are passed by reference. Apply this guideline whenever validating numeric fields to align with the derive macro behavior.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/src/core/token_bucket.rs:58-63
Timestamp: 2026-02-21T02:30:51.443Z
Learning: For lint-only/Clippy enforcement PRs in this repository, avoid introducing behavioral changes (e.g., new input validation or logic changes). Treat such PRs as non-functional changes and plan a separate follow-up issue/PR for hardening or behavior changes. This applies broadly to Rust files across the repo; during review, focus on lint/style corrections and clearly note any intentional exceptions. 

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: protocols/src/responses.rs:928-931
Timestamp: 2026-02-21T02:36:00.882Z
Learning: In Rust code across the repository, use the marker INVARIANT: to document assumptions in safe code. Reserve SAFETY: for explaining why unsafe blocks are sound. This improves clarity of invariants and safety reasoning. Example reference: protocols/src/responses.rs near validate_tool_choice_with_tools().

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: mesh/src/sync.rs:83-83
Timestamp: 2026-02-21T02:37:01.416Z
Learning: General Rust formatting rule: format! with implicit captures only supports simple identifiers, not full expressions like {state.model_id}. For cases where you want to interpolate a field or expression, bind the value first and interpolate the binding, e.g., let model_id = &state.model_id; and then use format!("policy:{}", model_id). In the specific file mesh/src/sync.rs, prefer format!("policy:{}", state.model_id) or bind to a local variable if you need named interpolation, to keep clarity and avoid unintended captures.

Comment on lines +58 to +63
impl LazySpec {
fn new(_id: &'static str, factory: fn() -> Box<dyn ModelProcessorSpec>) -> Self {
Self {
inner: Lazy::new(factory),
}
}

@coderabbitai coderabbitai Bot Mar 3, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Consider removing unused _id parameter or documenting its purpose.

The _id parameter in LazySpec::new is unused. If it's intended for future debugging/logging, consider adding a comment. Otherwise, it could be removed.

♻️ Option 1: Remove unused parameter
-    fn new(_id: &'static str, factory: fn() -> Box<dyn ModelProcessorSpec>) -> Self {
+    fn new(factory: fn() -> Box<dyn ModelProcessorSpec>) -> Self {
         Self {
             inner: Lazy::new(factory),
         }
     }

This would also require updating the call sites in ModelRegistry::new().

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@multimodal/src/registry/mod.rs` around lines 58 - 63, The _id parameter on
LazySpec::new is unused; either remove it or document its intent—update the
signature of LazySpec::new to remove the _id parameter (and all call sites, e.g.
ModelRegistry::new) if not needed, or add a short comment explaining why _id is
kept (e.g., reserved for future logging/debugging) and prefix it with underscore
to suppress warnings; adjust all invocations that construct LazySpec (notably in
ModelRegistry::new) to match the new signature or retain passing the id if you
choose to keep the parameter with documentation.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CatherineSue, understood—thank you for clarifying. Since this is a refactoring PR focused on splitting the module, keeping the existing signature as-is makes sense. The unused parameter can be addressed in a separate cleanup if needed.


🧠 Learnings used
Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 495
File: model_gateway/Cargo.toml:115-116
Timestamp: 2026-02-21T11:57:48.335Z
Learning: In repo lightseekorg/smg, the `image` crate dependency in model_gateway/Cargo.toml is intentionally declared without explicit codec features (`image = { version = "0.25.4", default-features = false }`) because codec features are provided transitively through the `llm-multimodal` workspace dependency, and Cargo's feature unification handles this correctly. The direct dependency is needed for type usage (e.g., `DynamicImage`) in model_gateway code.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 570
File: multimodal/src/registry.rs:133-143
Timestamp: 2026-03-01T06:00:39.427Z
Learning: In repo lightseekorg/smg, the image_sizes field in PreprocessedImages will be standardized to use `(height, width)` tuple ordering across all processors in multimodal/src/vision/processors/. Currently llama4, phi4, and pixtral use `(height, width)` while llava and phi3 use `(width, height)`. The image_sizes_hw helper in multimodal/src/registry.rs correctly interprets tuples as (h, w) for the majority convention.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 497
File: model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs:108-113
Timestamp: 2026-02-21T23:56:04.191Z
Learning: In repo lightseekorg/smg, file model_gateway/src/routers/grpc/regular/stages/chat/request_building.rs: When fetching tokenizer_source via ctx.components.tokenizer_registry.get_by_name(model_id).map(|e| e.source).unwrap_or_default(), an empty string is an intentional valid fallback when the model isn't in the registry. This allows config loading to proceed with the default path. Comments documenting this fallback are not necessary to keep noise down.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/benches/wasm_middleware_latency.rs:88-91
Timestamp: 2026-02-21T02:37:02.009Z
Learning: Repo: lightseekorg/smg — For clippy-only/enforcement PRs (e.g., PR `#489`), even micro-optimizations (like replacing an async closure with std::future::ready in benches such as model_gateway/benches/wasm_middleware_latency.rs) should be deferred to a follow-up PR rather than included inline.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/tests/vision_golden_tests.rs:576-578
Timestamp: 2026-02-21T02:39:51.670Z
Learning: Repo lightseekorg/smg — For PR `#489` (clippy/lint-only), do not replace unwrap() with expect(...) in test files when a file-/crate-level #![expect(clippy::unwrap_used)] is present (e.g., multimodal/tests/vision_golden_tests.rs). Such stylistic swaps are out of scope.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/tests/vision_golden_tests.rs:402-404
Timestamp: 2026-02-21T02:39:23.481Z
Learning: Repo lightseekorg/smg — For clippy-only/lint-enforcement PRs (e.g., PR `#489`), do not replace unwrap() with expect() across tests/benches when a crate-level `#![expect(clippy::unwrap_used)]` is present. Such per-call swaps are treated as out-of-scope stylistic changes. Example: multimodal/tests/vision_golden_tests.rs.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: mesh/src/crdt.rs:296-299
Timestamp: 2026-02-21T02:36:31.543Z
Learning: Repo lightseekorg/smg — For clippy/lint-only PRs (e.g., PR `#489`), avoid requesting stylistic doc comments when an item is already annotated with #[expect(...)] (e.g., #[expect(dead_code)] on SyncCRDTMap::contains_key in mesh/src/crdt.rs); such style changes are considered out of scope.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/src/vision/processors/phi4_vision.rs:486-490
Timestamp: 2026-02-21T02:39:17.064Z
Learning: In Rust files under multimodal/src (e.g., multimodal/src/vision/processors/phi4_vision.rs), prefer returning recoverable errors via ok_or(TransformError::EmptyBatch) (or similar error) in Phi4VisionProcessor::preprocess when a batch is empty, instead of panicking with expect/unreachable. This aligns with the repo guideline to avoid panics in production code (PR `#489`) and lint-only efforts. Apply this pattern to similar preprocessing paths in this module and related code paths where a non-fatal error conveys meaningful failure to callers.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/benches/wasm_middleware_latency.rs:1-1
Timestamp: 2026-02-21T02:37:04.633Z
Learning: Repo: lightseekorg/smg — For benchmark files (e.g., model_gateway/benches/*.rs), using a crate-level `#![expect(clippy::unwrap_used, clippy::disallowed_methods)]` is preferred when unwrap/spawn are used throughout. Do not push to per-function scoping; a single crate-level `reason` is acceptable when justification is required.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 588
File: model_gateway/src/routers/grpc/multimodal.rs:453-514
Timestamp: 2026-03-03T18:03:37.820Z
Learning: In repo lightseekorg/smg, backend assembly functions in model_gateway/src/routers/grpc/multimodal.rs (e.g., assemble_sglang, assemble_vllm, assemble_trtllm) are tested via E2E tests rather than unit tests, as unit tests for these functions are not considered worthwhile.

Learnt from: XinyueZhang369
Repo: lightseekorg/smg PR: 399
File: protocols/src/interactions.rs:505-509
Timestamp: 2026-02-19T03:08:50.192Z
Learning: In code reviews for Rust projects using the validator crate (v0.20.0), ensure that custom validation functions for numeric primitive types (e.g., f32, i32, u32, i16, etc.) accept the value by value, not by reference. Example: fn validate(value: f32) { ... }. The validator derive macro has a hardcoded list of numeric types that are passed by value, while all other types are passed by reference. Apply this guideline whenever validating numeric fields to align with the derive macro behavior.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/src/core/token_bucket.rs:58-63
Timestamp: 2026-02-21T02:30:51.443Z
Learning: For lint-only/Clippy enforcement PRs in this repository, avoid introducing behavioral changes (e.g., new input validation or logic changes). Treat such PRs as non-functional changes and plan a separate follow-up issue/PR for hardening or behavior changes. This applies broadly to Rust files across the repo; during review, focus on lint/style corrections and clearly note any intentional exceptions. 

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: protocols/src/responses.rs:928-931
Timestamp: 2026-02-21T02:36:00.882Z
Learning: In Rust code across the repository, use the marker INVARIANT: to document assumptions in safe code. Reserve SAFETY: for explaining why unsafe blocks are sound. This improves clarity of invariants and safety reasoning. Example reference: protocols/src/responses.rs near validate_tool_choice_with_tools().

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: mesh/src/sync.rs:83-83
Timestamp: 2026-02-21T02:37:01.416Z
Learning: General Rust formatting rule: format! with implicit captures only supports simple identifiers, not full expressions like {state.model_id}. For cases where you want to interpolate a field or expression, bind the value first and interpolate the binding, e.g., let model_id = &state.model_id; and then use format!("policy:{}", model_id). In the specific file mesh/src/sync.rs, prefer format!("policy:{}", state.model_id) or bind to a local variable if you need named interpolation, to keep clarity and avoid unintended captures.

Comment on lines +38 to +41
fn matches(&self, metadata: &ModelMetadata) -> bool {
metadata.model_id.to_ascii_lowercase().contains("qwen")
&& metadata.model_id.to_ascii_lowercase().contains("vl")
}

@coderabbitai coderabbitai Bot Mar 3, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Minor: Avoid duplicate to_ascii_lowercase() calls.

The model ID is lowercased twice in the same check. Consider storing the result.

♻️ Proposed refactor
     fn matches(&self, metadata: &ModelMetadata) -> bool {
-        metadata.model_id.to_ascii_lowercase().contains("qwen")
-            && metadata.model_id.to_ascii_lowercase().contains("vl")
+        let id = metadata.model_id.to_ascii_lowercase();
+        id.contains("qwen") && id.contains("vl")
     }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@multimodal/src/registry/qwen_vl.rs` around lines 38 - 41, In the matches
method of qwen_vl.rs, avoid calling metadata.model_id.to_ascii_lowercase() twice
by assigning the lowercased model id to a local variable (e.g., let id =
metadata.model_id.to_ascii_lowercase()) and then use id.contains("qwen") &&
id.contains("vl"); update the matches(&self, metadata: &ModelMetadata)
implementation accordingly to reference that local variable.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-existing code, not introduced by this PR.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CatherineSue, understood! Since this is a refactoring PR focused on splitting files, it makes sense to keep the code as-is and address this optimization separately if needed. Thanks for the clarification.


🧠 Learnings used
Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/benches/wasm_middleware_latency.rs:88-91
Timestamp: 2026-02-21T02:37:02.009Z
Learning: Repo: lightseekorg/smg — For clippy-only/enforcement PRs (e.g., PR `#489`), even micro-optimizations (like replacing an async closure with std::future::ready in benches such as model_gateway/benches/wasm_middleware_latency.rs) should be deferred to a follow-up PR rather than included inline.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/tests/vision_golden_tests.rs:576-578
Timestamp: 2026-02-21T02:39:51.670Z
Learning: Repo lightseekorg/smg — For PR `#489` (clippy/lint-only), do not replace unwrap() with expect(...) in test files when a file-/crate-level #![expect(clippy::unwrap_used)] is present (e.g., multimodal/tests/vision_golden_tests.rs). Such stylistic swaps are out of scope.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/tests/vision_golden_tests.rs:402-404
Timestamp: 2026-02-21T02:39:23.481Z
Learning: Repo lightseekorg/smg — For clippy-only/lint-enforcement PRs (e.g., PR `#489`), do not replace unwrap() with expect() across tests/benches when a crate-level `#![expect(clippy::unwrap_used)]` is present. Such per-call swaps are treated as out-of-scope stylistic changes. Example: multimodal/tests/vision_golden_tests.rs.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: mesh/src/crdt.rs:296-299
Timestamp: 2026-02-21T02:36:31.543Z
Learning: Repo lightseekorg/smg — For clippy/lint-only PRs (e.g., PR `#489`), avoid requesting stylistic doc comments when an item is already annotated with #[expect(...)] (e.g., #[expect(dead_code)] on SyncCRDTMap::contains_key in mesh/src/crdt.rs); such style changes are considered out of scope.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/benches/wasm_middleware_latency.rs:1-1
Timestamp: 2026-02-21T02:37:04.633Z
Learning: Repo: lightseekorg/smg — For benchmark files (e.g., model_gateway/benches/*.rs), using a crate-level `#![expect(clippy::unwrap_used, clippy::disallowed_methods)]` is preferred when unwrap/spawn are used throughout. Do not push to per-function scoping; a single crate-level `reason` is acceptable when justification is required.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 570
File: multimodal/src/registry.rs:133-143
Timestamp: 2026-03-01T06:00:39.427Z
Learning: In repo lightseekorg/smg, the image_sizes field in PreprocessedImages will be standardized to use `(height, width)` tuple ordering across all processors in multimodal/src/vision/processors/. Currently llama4, phi4, and pixtral use `(height, width)` while llava and phi3 use `(width, height)`. The image_sizes_hw helper in multimodal/src/registry.rs correctly interprets tuples as (h, w) for the majority convention.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: multimodal/src/vision/processors/phi4_vision.rs:486-490
Timestamp: 2026-02-21T02:39:17.064Z
Learning: In Rust files under multimodal/src (e.g., multimodal/src/vision/processors/phi4_vision.rs), prefer returning recoverable errors via ok_or(TransformError::EmptyBatch) (or similar error) in Phi4VisionProcessor::preprocess when a batch is empty, instead of panicking with expect/unreachable. This aligns with the repo guideline to avoid panics in production code (PR `#489`) and lint-only efforts. Apply this pattern to similar preprocessing paths in this module and related code paths where a non-fatal error conveys meaningful failure to callers.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 588
File: model_gateway/src/routers/grpc/multimodal.rs:453-514
Timestamp: 2026-03-03T18:03:37.820Z
Learning: In repo lightseekorg/smg, backend assembly functions in model_gateway/src/routers/grpc/multimodal.rs (e.g., assemble_sglang, assemble_vllm, assemble_trtllm) are tested via E2E tests rather than unit tests, as unit tests for these functions are not considered worthwhile.

Learnt from: XinyueZhang369
Repo: lightseekorg/smg PR: 399
File: protocols/src/interactions.rs:505-509
Timestamp: 2026-02-19T03:08:50.192Z
Learning: In code reviews for Rust projects using the validator crate (v0.20.0), ensure that custom validation functions for numeric primitive types (e.g., f32, i32, u32, i16, etc.) accept the value by value, not by reference. Example: fn validate(value: f32) { ... }. The validator derive macro has a hardcoded list of numeric types that are passed by value, while all other types are passed by reference. Apply this guideline whenever validating numeric fields to align with the derive macro behavior.

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: model_gateway/src/core/token_bucket.rs:58-63
Timestamp: 2026-02-21T02:30:51.443Z
Learning: For lint-only/Clippy enforcement PRs in this repository, avoid introducing behavioral changes (e.g., new input validation or logic changes). Treat such PRs as non-functional changes and plan a separate follow-up issue/PR for hardening or behavior changes. This applies broadly to Rust files across the repo; during review, focus on lint/style corrections and clearly note any intentional exceptions. 

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: protocols/src/responses.rs:928-931
Timestamp: 2026-02-21T02:36:00.882Z
Learning: In Rust code across the repository, use the marker INVARIANT: to document assumptions in safe code. Reserve SAFETY: for explaining why unsafe blocks are sound. This improves clarity of invariants and safety reasoning. Example reference: protocols/src/responses.rs near validate_tool_choice_with_tools().

Learnt from: slin1237
Repo: lightseekorg/smg PR: 489
File: mesh/src/sync.rs:83-83
Timestamp: 2026-02-21T02:37:01.416Z
Learning: General Rust formatting rule: format! with implicit captures only supports simple identifiers, not full expressions like {state.model_id}. For cases where you want to interpolate a field or expression, bind the value first and interpolate the binding, e.g., let model_id = &state.model_id; and then use format!("policy:{}", model_id). In the specific file mesh/src/sync.rs, prefer format!("policy:{}", state.model_id) or bind to a local variable if you need named interpolation, to keep clarity and avoid unintended captures.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

multimodal Multimodal crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant