Skip to content

fix(grpc_servicer): return served_model_name in vLLM GetModelInfo response - #727

Merged
slin1237 merged 1 commit into
mainfrom
fix/grpc-served-model-name
Mar 11, 2026
Merged

slin1237 merged 1 commit into
mainfrom
fix/grpc-served-model-name

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 11, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

When using smg serve --backend vllm --served-model-name vllm-model, the gateway registers the worker with the raw filesystem model path (e.g. /raid/models/...) instead of the user-specified --served-model-name.

This happens because the vLLM gRPC servicer returns model_config.model (always the filesystem path) in the GetModelInfo response, ignoring the served_model_name that vLLM sets when --served-model-name is provided.

Solution

Use model_config.served_model_name (which vLLM sets via get_served_model_name() — falls back to model_config.model when not specified) instead of model_config.model in the GetModelInfo response.

Changes

  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py: Return model_config.served_model_name or model_config.model in model_path field of GetModelInfoResponse

Test Plan

  1. Launch vLLM gRPC server with --served-model-name vllm-model
  2. Connect smg gateway to the worker
  3. Verify smg registers the worker with model name vllm-model instead of the filesystem path

Before
Screenshot 2026-03-11 at 12 05 08 AM

After
Screenshot 2026-03-11 at 12 54 52 PM
Screenshot 2026-03-11 at 12 54 59 PM

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • New Features
    • Model info responses now include a "served model name" field so interfaces can display the configured served name when available.
  • Bug Fixes
    • Improved model-name reporting to prefer the served model name with a fallback to the original model identifier, reducing ambiguity in displays and logs.

@CatherineSue
CatherineSue requested a review from njhill as a code owner March 11, 2026 07:07
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@coderabbitai

coderabbitai Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

GetModelInfo now includes a new response field served_model_name and the servicer populates it from model_config.served_model_name when present, otherwise from model_config.model. No function signatures were changed.

Changes

Cohort / File(s) Summary
gRPC Servicer
grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Populate GetModelInfoResponse.served_model_name using model_config.served_model_name if available, else fallback to model_config.model.
Proto
crates/grpc_client/proto/vllm_engine.proto
Added string served_model_name = 6; to GetModelInfoResponse message (new response field).

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Suggested labels

grpc, model-gateway

Suggested reviewers

  • njhill
  • slin1237

Poem

🐰 I nibble on bytes and hop through names so spry,
Served name first, then fallback — watch me fly!
A tiny tweak, a careful little tweak,
Models feel snug and responses speak. 🥕✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding served_model_name to the vLLM GetModelInfo response, addressing the core problem described in the PR.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch fix/grpc-served-model-name

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses an issue where the vLLM backend, when integrated with the smg gateway, incorrectly reported its model identifier as the filesystem path rather than the intended served_model_name. By adjusting the GetModelInfo response, the gateway will now correctly register the worker using the user-defined model name, improving clarity and consistency in model identification.

Highlights

  • vLLM GetModelInfo Response: Modified the vLLM gRPC servicer to return the user-specified served model name (served_model_name) instead of the raw filesystem path in the GetModelInfo response.
Changelog
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py
    • Updated the GetModelInfo method to prioritize 'served_model_name' over 'model' for the 'model_path' field.
Activity
  • The author has completed the code formatting and linting checks.
  • No documentation updates were provided in this pull request.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request addresses a bug where the vLLM gRPC servicer would incorrectly return the model's filesystem path instead of the user-provided served-model-name. The fix correctly prioritizes served_model_name in the GetModelInfo response, falling back to the model path, which ensures that the gateway registers the worker with the intended model identifier. The change is correct and aligns with the problem description.

…ponse

The gRPC servicer was returning only model_config.model in the model_path
field, with no served_model_name. This caused smg to register workers with
the raw filesystem path instead of the user-specified --served-model-name.

Add a served_model_name field to the vLLM GetModelInfoResponse proto and
populate it from model_config.served_model_name. This keeps model_path as
the filesystem path (needed for tokenizer loading) while providing the
served name separately for model_id resolution.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@CatherineSue
CatherineSue force-pushed the fix/grpc-served-model-name branch from 03a36d7 to a90f492 Compare March 11, 2026 07:32
@CatherineSue
CatherineSue requested a review from slin1237 as a code owner March 11, 2026 07:32

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
grpc_servicer/smg_grpc_servicer/vllm/servicer.py (1)

268-274: ⚠️ Potential issue | 🔴 Critical

Update model_path too, or the gateway bug remains.

Line 269 still returns model_config.model for model_path, so existing GetModelInfo consumers will keep seeing the filesystem path. Adding served_model_name alongside it does not fix the registration flow described in this PR.

Suggested fix
+        served_model_name = model_config.served_model_name or model_config.model
         return vllm_engine_pb2.GetModelInfoResponse(
-            model_path=model_config.model,
+            model_path=served_model_name,
             is_generation=model_config.runner_type == "generate",
             max_context_length=model_config.max_model_len,
             vocab_size=model_config.get_vocab_size(),
             supports_vision=model_config.is_multimodal_model,
-            served_model_name=model_config.served_model_name or model_config.model,
+            served_model_name=served_model_name,
         )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py` around lines 268 - 274, The
GetModelInfoResponse is still returning the filesystem path via
model_config.model for the model_path field, so update the model_path assignment
to prefer model_config.served_model_name (falling back to model_config.model) so
consumers receive the served model name; change the model_path in the
GetModelInfoResponse construction (in the GetModelInfoResponse return block) to
something like served_model_name or model as a fallback using
model_config.served_model_name or model_config.model, keeping served_model_name
logic for served_model_name unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 268-274: The GetModelInfoResponse is still returning the
filesystem path via model_config.model for the model_path field, so update the
model_path assignment to prefer model_config.served_model_name (falling back to
model_config.model) so consumers receive the served model name; change the
model_path in the GetModelInfoResponse construction (in the GetModelInfoResponse
return block) to something like served_model_name or model as a fallback using
model_config.served_model_name or model_config.model, keeping served_model_name
logic for served_model_name unchanged.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 3144910a-0867-4fa1-91bc-1e886a6a5a73

📥 Commits

Reviewing files that changed from the base of the PR and between 03a36d7 and a90f492.

📒 Files selected for processing (2)
  • crates/grpc_client/proto/vllm_engine.proto
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

@slin1237
slin1237 merged commit de0f695 into main Mar 11, 2026
25 checks passed
@slin1237
slin1237 deleted the fix/grpc-served-model-name branch March 11, 2026 07:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants