Skip to content

fix(multimodal): fall back to config.model_type for aliased model IDs - #898

Merged
CatherineSue merged 1 commit into
mainfrom
fix/add-fallback-mm
Mar 25, 2026
Merged

CatherineSue merged 1 commit into
mainfrom
fix/add-fallback-mm

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 24, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

Multimodal family detection fails when a vision model is served under a custom/aliased name (e.g. custom-model) because both the model spec registry and image processor registry match exclusively on the request-side model_id string. Even though config.json and preprocessor_config.json are loaded successfully from the real model path, the family detection never consults config.model_type.

Solution

Add a config.model_type fallback to each spec's matches() method and to ImageProcessorRegistry::find(). The existing model_id substring check remains the fast path; config.model_type is only consulted when model_id does not match any known family. Unlike the model_id path which uses substring matching, model_type uses exact equality since it is a known fixed value from config.json.

Changes

  • Add config_model_type() helper to ModelMetadata for reading config.json's model_type field
  • Update each multimodal spec's matches() to fall back to exact config.model_type matching when model_id substring check fails (qwen3_vl, qwen_vl, llava, llama4, phi3_v)
  • Extend ImageProcessorRegistry::find() to accept an optional model_type parameter and fall back to it when model_id doesn't match
  • Register phi3_v pattern in image processor defaults so model_type fallback can match it
  • Thread model_type from loaded config through the caller in multimodal.rs
  • Remove unused has_processor() from ImageProcessorRegistry
  • Add regression tests for aliased model IDs across all multimodal specs and image processor lookup

Supersedes #755
Fixes #754

Test Plan

  1. cargo test -p llm-multimodal — all 220 tests pass
  2. cargo check -p smg — gateway compiles cleanly
  3. Configure a gRPC worker to serve a supported multimodal model under an aliased name such as custom-model
  4. Send a chat completion request with image content using the aliased model name
  5. Verify SMG resolves the correct multimodal spec and image processor
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Refactor
    • Improved model detection to recognize several vision and multimodal models via an optional config "model_type", and enhanced image processor selection to fall back to this model type when needed.
  • Tests
    • Added unit tests covering model-type-based identification and the new processor selection fallback behavior.

Fixes #754

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Mar 24, 2026
@coderabbitai

coderabbitai Bot commented Mar 24, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: d5a08730-8072-42f8-830b-00b4ece9e4a8

📥 Commits

Reviewing files that changed from the base of the PR and between 84b7593 and 5ce2178.

📒 Files selected for processing (8)
  • crates/multimodal/src/registry/llama4.rs
  • crates/multimodal/src/registry/llava.rs
  • crates/multimodal/src/registry/phi3_v.rs
  • crates/multimodal/src/registry/qwen3_vl.rs
  • crates/multimodal/src/registry/qwen_vl.rs
  • crates/multimodal/src/registry/traits.rs
  • crates/multimodal/src/vision/image_processor.rs
  • model_gateway/src/routers/grpc/multimodal.rs

📝 Walkthrough

Walkthrough

Added a config-based fallback for multimodal family detection: ModelMetadata gains config_model_type(), vision specs and image processor registry can match by config.model_type when model_id substring checks fail, and the multimodal gRPC handler passes the extracted model_type into the image-processor lookup. Tests added.

Changes

Cohort / File(s) Summary
Registry Foundation
crates/multimodal/src/registry/traits.rs
Added pub fn config_model_type(&self) -> Option<&str> to read "model_type" from model config.
Vision Spec Matchers
crates/multimodal/src/registry/llama4.rs, crates/multimodal/src/registry/llava.rs, crates/multimodal/src/registry/phi3_v.rs, crates/multimodal/src/registry/qwen3_vl.rs, crates/multimodal/src/registry/qwen_vl.rs
Extended matches() implementations to accept config_model_type() as an alias/match path; added unit tests exercising alias-via-config for each spec.
Image Processor Registry
crates/multimodal/src/vision/image_processor.rs
Changed find(&self, model_id, model_type: Option<&str>) with fallback search by model_type, removed has_processor(), refactored matching via find_in_candidate(), registered "phi3_v" default, and updated tests.
gRPC Multimodal Handler
model_gateway/src/routers/grpc/multimodal.rs
Extracts optional model_type from loaded config and passes it into ImageProcessorRegistry::find(model_id, model_type) during multimodal processing.

Sequence Diagram

sequenceDiagram
    participant Client as gRPC Client
    participant Handler as Multimodal Handler
    participant Config as Config Loader
    participant Registry as Spec Registry
    participant Processor as Image Processor Registry

    Client->>Handler: Request with custom model_id
    Handler->>Config: Load tokenizer/config (config.json)
    Config-->>Handler: config (may include model_type)
    Handler->>Registry: lookup(metadata with model_id + config_model_type)
    
    rect rgba(100, 150, 255, 0.5)
    Note over Registry: Spec Matching
    Registry->>Registry: Try match by model_id
    alt model_id matches
        Registry-->>Handler: Return spec
    else no model_id match
        Registry->>Registry: Try match by config_model_type
        Registry-->>Handler: Return spec (if any)
    end
    end

    Handler->>Processor: find(model_id, model_type)
    
    rect rgba(150, 200, 150, 0.5)
    Note over Processor: Processor Selection
    Processor->>Processor: Search by model_id
    alt model_id matches
        Processor-->>Handler: Return processor
    else no model_id match
        Processor->>Processor: Fallback to model_type search
        Processor-->>Handler: Return processor (if any)
    end
    end

    Handler-->>Client: Proceed with multimodal processing / error if none found
Loading

Estimated Code Review Effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Possibly Related PRs

Suggested Labels

multimodal, tests

Suggested Reviewers

  • slin1237
  • key4ng

Poem

🐰 I hop through configs, sniffing model_type,
When custom names hide the family’s type,
I nudge the registry to look inside,
So vision processors find where they should hide,
Hooray — aliased models now take flight! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'fix(multimodal): fall back to config.model_type for aliased model IDs' accurately and concisely describes the main change: enabling multimodal family detection to use config.model_type as a fallback when model IDs don't match known patterns.
Linked Issues check ✅ Passed The PR fully addresses issue #754's requirements: added config_model_type() helper [traits.rs], updated all spec matches() implementations to fall back to exact config.model_type equality [llama4, llava, phi3_v, qwen3_vl, qwen_vl], extended ImageProcessorRegistry::find() with model_type parameter [image_processor.rs], and threaded model_type through multimodal.rs.
Out of Scope Changes check ✅ Passed All changes are directly scoped to fixing the multimodal family detection problem: config_model_type() extraction, spec matching updates, image processor registry changes, and removing the unused has_processor() method. No unrelated refactoring or feature additions were introduced.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/add-fallback-mm

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request resolves an issue where multimodal vision models could not be correctly identified when served under custom or aliased names. The core solution involves implementing a fallback mechanism that consults the model_type field within the config.json if the primary model_id matching fails. This ensures that both model specifications and their corresponding image processors are accurately resolved, significantly improving the flexibility and reliability of multimodal model deployment.

Highlights

  • Enhanced Multimodal Model Detection: Multimodal model detection now robustly handles custom or aliased model IDs by falling back to the config.model_type field when the initial model_id substring matching fails.
  • Improved Image Processor Lookup: The image processor registry's lookup mechanism has been updated to also utilize config.model_type as a fallback, ensuring correct processor identification for aliased models.
  • New Configuration Helper: A config_model_type() helper function was introduced in ModelMetadata to facilitate reading the model_type from config.json.
  • Comprehensive Regression Tests: Extensive regression tests have been added to verify the new aliased model ID detection logic across all supported multimodal specifications and image processor lookups.
  • Code Cleanup: An unused has_processor() method was removed from the ImageProcessorRegistry.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request enhances model identification by introducing a fallback mechanism to use the model_type from model configurations when the model_id doesn't directly match. This new logic is applied to various model specifications (Llama4, Llava, Phi3-V, Qwen3-VL, Qwen-VL) and the ImageProcessorRegistry, with corresponding updates in the model_gateway and new tests to validate the fallback. However, there are two high-severity issues identified: the LlavaSpec's prompt_replacements incorrectly calculates image tokens for LLaVA-NeXT models, and the Phi3VisionSpec's prompt_replacements might inaccurately recalculate image tokens, both of which should be updated to use the preprocessed.num_img_tokens for correctness and consistency.

Comment thread crates/multimodal/src/registry/llava.rs
Comment thread crates/multimodal/src/registry/phi3_v.rs
Multimodal family detection failed when a vision model was served under
a custom name (e.g. "custom-model") because both the model spec registry
and image processor registry matched exclusively on the request-side
model_id string.

Add a config.model_type fallback to each spec's matches() method and to
ImageProcessorRegistry::find(). The existing model_id substring check
remains the fast path; config.model_type is only consulted when model_id
does not match any known family. Unlike substring matching, model_type
uses exact equality since it is a known fixed value from config.json.

Also removes unused has_processor() from ImageProcessorRegistry.

Fixes #754

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Multimodal family detection fails for aliased/custom model IDs even when tokenizer/config point to the correct vision model

1 participant