Skip to content

feat(mm): share multimodal config registry keyed by tokenizer UUID - #1247

Closed
key4ng wants to merge 7 commits into
smg-project:mainfrom
key4ng:keyang/mm-config-registry
Closed

key4ng wants to merge 7 commits into
smg-project:mainfrom
key4ng:keyang/mm-config-registry

Conversation

@key4ng

@key4ng key4ng commented Apr 20, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

In IGW/K8s mode the gateway loads tokenizers via GetTokenizer gRPC streaming. The bundle contains config.json and preprocessor_config.json, but those files are extracted to a tempdir that is deleted immediately after the tokenizer is loaded into memory.

When a multimodal request then arrives, multimodal.rs tries to read these files from disk using the tokenizer source — which in IGW is a worker-only path the gateway can't reach. The HF fallback also fails in air-gapped environments. Multimodal therefore breaks for any model whose tokenizer was loaded via GetTokenizer.

On top of that, the cache lived inside MultimodalComponents (per-router, unreachable from tokenizer registration) and was keyed by the request-time model_id — which is ambiguous because callers may pass either a tokenizer name or a tokenizer UUID.

Solution

  • New gateway-local MultimodalConfigRegistry owned by AppContext, keyed by TokenizerEntry.id (UUID).
  • MultimodalComponents now holds an Arc reference to the shared registry instead of its own cache.
  • GetTokenizer bundle extraction now reads both JSON files out of the tempdir before cleanup and inserts the parsed MultimodalModelConfig into the registry under the tokenizer UUID. Failure is non-fatal — text-only tokenizers legitimately have no preprocessor_config.json.
  • Request-time code resolves the TokenizerEntry once at the preparation stage boundary and threads both tokenizer_id and tokenizer_source into the multimodal helpers.

Changes

  • model_gateway/src/routers/grpc/multimodal.rs — add MultimodalConfigRegistry; drop model_configs from MultimodalComponents; route multimodal lookups through the shared registry keyed by tokenizer UUID.
  • model_gateway/src/app_context.rs — new multimodal_config_registry: Arc<MultimodalConfigRegistry> field initialized in AppContextBuilder::build().
  • model_gateway/src/workflow/tokenizer_registration.rs — read and parse mm config files inside the bundle-extraction closure (before tempdir cleanup); insert into the registry under the tokenizer UUID.
  • model_gateway/src/routers/grpc/router.rs + pd_router.rs — pass the shared registry into MultimodalComponents::new.
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs + messages/preparation.rs — resolve TokenizerEntry once; pass tokenizer_id + tokenizer_source downstream.
  • model_gateway/src/service_discovery.rs — update test fixture for the new field.

Test Plan

  • cargo test -p smg --lib routers::grpc::multimodal::tests — 14/14 passing.
  • cargo test -p smg --lib — full gateway library suite green.
  • cargo check -p smg --all-targets — clean.
  • rg 'get_or_load_config|model_configs' model_gateway/src — zero orphaned references.

New tests (all in `model_gateway/src/routers/grpc/multimodal.rs`):

  • `registry_new_is_empty`, `registry_insert_and_get_roundtrip`, `registry_get_or_load_reads_from_local_dir_and_caches` — core registry behavior with tempdir fixtures.
  • `registry_get_or_load_hits_preloaded_entry_without_touching_source` — unit-level regression fence: a preloaded entry is returned without consulting a deliberately unreachable `tokenizer_source`.
  • `multimodal_components_serves_preloaded_config_without_touching_source` — integration-level regression fence: same invariant through `MultimodalComponents::new(registry.clone())`, confirming the `AppContext` wiring threads the same `Arc` end-to-end.
Checklist
  • `cargo +nightly fmt` passes
  • `cargo clippy --all-targets --all-features -- -D warnings` passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Refactor
    • Introduced centralized multimodal configuration registry for improved caching and resource sharing across the application
    • Changed configuration resolution from model-based to tokenizer-based lookup strategy
    • Enabled automatic preloading and caching of multimodal configurations when tokenizers are registered
    • Optimized configuration access patterns for multimodal processing components

key4ng added 7 commits April 19, 2026 23:14
Shared gateway-local cache for config.json + preprocessor_config.json,
used as the preload target for GetTokenizer bundles and the lookup surface
for multimodal request handlers. Not yet wired into AppContext or
MultimodalComponents; that happens in the following commits.

Signed-off-by: key4ng <rukeyang@gmail.com>
Expose MultimodalConfigRegistry at AppContext so tokenizer registration
(preload path) and request handlers (lookup path) can share one cache.

Signed-off-by: key4ng <rukeyang@gmail.com>
Drop per-router model_configs DashMap from MultimodalComponents; hold an
Arc<MultimodalConfigRegistry> reference instead. Multimodal helpers now
take tokenizer_id + tokenizer_source (resolved once at the stage
boundary) so the canonical cache key is the TokenizerEntry UUID, not the
ambiguous request-time model_id.

Signed-off-by: key4ng <rukeyang@gmail.com>
Read config.json + preprocessor_config.json from the extracted bundle
tempdir before cleanup, and insert the parsed MultimodalModelConfig into
AppContext.multimodal_config_registry keyed by tokenizer UUID. Fixes
multimodal requests in IGW/K8s mode where the worker-reported tokenizer
source is unreachable from the gateway.

Signed-off-by: key4ng <rukeyang@gmail.com>
Asserts that an entry preloaded into MultimodalConfigRegistry under the
tokenizer UUID is served to MultimodalComponents.config_registry lookups
without falling through to tokenizer_source. Stops the original IGW bug
from returning silently.

Signed-off-by: key4ng <rukeyang@gmail.com>
Silence private_interfaces warnings: the registry is held in Arc on a
pub AppContext field, but its methods don't need to be callable outside
the model_gateway crate. Struct stays pub; methods become pub(crate).

Signed-off-by: key4ng <rukeyang@gmail.com>
Nightly rustfmt required minor layout fixes to the tokenizer_registration.rs
preload helpers and the multimodal.rs tests module. service_discovery.rs
gained a `use` import for MultimodalConfigRegistry to silence the
clippy::absolute_paths lint triggered by the long crate::-prefixed path.

No behavior change.

Signed-off-by: key4ng <rukeyang@gmail.com>
@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Apr 20, 2026
@coderabbitai

coderabbitai Bot commented Apr 20, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 5b0f5782-bc2e-44e1-a7c4-b0e4706f2e77

📥 Commits

Reviewing files that changed from the base of the PR and between 5c9bd90 and e860b56.

📒 Files selected for processing (8)
  • model_gateway/src/app_context.rs
  • model_gateway/src/routers/grpc/multimodal.rs
  • model_gateway/src/routers/grpc/pd_router.rs
  • model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs
  • model_gateway/src/routers/grpc/regular/stages/messages/preparation.rs
  • model_gateway/src/routers/grpc/router.rs
  • model_gateway/src/service_discovery.rs
  • model_gateway/src/workflow/tokenizer_registration.rs

📝 Walkthrough

Walkthrough

This pull request introduces a shared MultimodalConfigRegistry in the application context to centralize multimodal model configuration management. The registry replaces per-router local caches, is keyed by tokenizer_id, and supports lazy loading from tokenizer sources. The multimodal pipeline is updated to thread tokenizer_id throughout, and configurations are preloaded during tokenizer registration.

Changes

Cohort / File(s) Summary
Core Registry Implementation
model_gateway/src/app_context.rs, model_gateway/src/routers/grpc/multimodal.rs
Added MultimodalConfigRegistry struct with get, insert, and get_or_load methods for caching configs by tokenizer_id. Integrated registry into AppContext as a shared field initialized during startup.
Router Integration
model_gateway/src/routers/grpc/pd_router.rs, model_gateway/src/routers/grpc/router.rs
Updated MultimodalComponents::new() calls to accept and use the shared multimodal_config_registry from context instead of creating isolated caches.
Multimodal Pipeline Updates
model_gateway/src/routers/grpc/regular/stages/chat/preparation.rs, model_gateway/src/routers/grpc/regular/stages/messages/preparation.rs
Extended multimodal processing functions to thread tokenizer_id through the call chain alongside tokenizer_source. Updated tokenizer resolution to extract and validate both fields before proceeding.
Tokenizer Registration & Preloading
model_gateway/src/workflow/tokenizer_registration.rs
Added extraction and preloading of multimodal model config from tokenizer bundles into the shared registry during worker tokenizer registration to populate the cache proactively.
Test Infrastructure
model_gateway/src/service_discovery.rs
Updated test AppContext setup to initialize MultimodalConfigRegistry dependency.

Sequence Diagram(s)

sequenceDiagram
    participant Router as Multimodal<br/>Router
    participant Prep as Preparation<br/>Stage
    participant Registry as Multimodal<br/>Config Registry
    participant TokenSrc as Tokenizer<br/>Source
    participant Worker as Tokenizer<br/>Registration

    Worker->>Registry: preload_config(tokenizer_id, config)
    Note over Registry: Cache entry created
    
    Router->>Prep: process_message(tokenizer_id, tokenizer_source)
    Prep->>Registry: get_or_load(tokenizer_id, tokenizer_source)
    
    alt Config in cache
        Registry-->>Prep: return Arc<MultimodalModelConfig>
    else Cache miss
        Registry->>TokenSrc: resolve_model_config_dir(tokenizer_source)
        TokenSrc-->>Registry: path
        Registry->>TokenSrc: read config.json + preprocessor_config.json
        TokenSrc-->>Registry: parsed config
        Registry->>Registry: insert(tokenizer_id, config)
        Registry-->>Prep: return Arc<MultimodalModelConfig>
    end
    
    Prep->>Prep: process_multimodal(tokenizer_id, config)
    Prep-->>Router: processed_messages
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~55 minutes

Possibly related PRs

Suggested labels

model-gateway, multimodal, grpc, tests

Suggested reviewers

  • slin1237

Poem

🐰 A registry shared, no longer scattered,
Each tokenizer cached, configuration mattered.
Through preparation stages, the ID threads true,
Preloaded and ready—the multimodal's debut! ✨

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands and usage tips.

@key4ng key4ng closed this Apr 20, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a centralized MultimodalConfigRegistry within the AppContext to manage and cache multimodal model configurations across the gateway. The changes update the tokenizer registration workflow to preload these configurations from bundles and refactor the multimodal processing pipeline to utilize the shared registry via tokenizer IDs. Review feedback highlights a potential cache stampede vulnerability in the registry's loading logic and identifies several instances of synchronous file I/O that should be converted to asynchronous or offloaded to prevent blocking the async runtime.

Comment on lines +72 to 76
pub(crate) async fn get_or_load(
&self,
model_id: &str,
tokenizer_id: &str,
tokenizer_source: &str,
) -> Result<Arc<MultimodalModelConfig>> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The get_or_load method is susceptible to a cache stampede (thundering herd) because multiple concurrent requests for the same tokenizer_id can simultaneously pass the cache check and trigger redundant loading operations. Consider using a synchronization mechanism (e.g., tokio::sync::OnceCell or an async-aware cache) to ensure that only one loading operation is performed per key.

Comment on lines 95 to 96
let config: serde_json::Value = std::fs::read_to_string(&config_path)
.with_context(|| format!("Failed to read config.json at {}", config_path.display()))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The use of std::fs::read_to_string inside an async function blocks the current thread, which can degrade the performance of the asynchronous runtime. Consider using tokio::fs::read_to_string instead.

Comment on lines 104 to 105
let preprocessor_config = std::fs::read_to_string(&pp_config_path)
.with_context(|| {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Similar to the config.json read above, this synchronous file read blocks the async executor. It should be replaced with an asynchronous alternative like tokio::fs::read_to_string.

@@ -305,7 +371,19 @@ async fn fetch_tokenizer_from_worker(
);

match load_tokenizer_from_bundle(&bundle) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

load_tokenizer_from_bundle performs synchronous file IO and tokenizer creation, which blocks the async executor when called directly from fetch_tokenizer_from_worker. This should be wrapped in tokio::task::spawn_blocking to maintain runtime responsiveness. Note that data passed to spawned background tasks must have a 'static lifetime; use owned types or reference-counted pointers like Arc instead of passing references to ensure the data outlives the task.

References
  1. Data passed to spawned background tasks must have a 'static lifetime. Use owned types or reference-counted pointers like Arc instead of passing references to ensure the data outlives the task.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant