Skip to content

refactor(grpc): optimize vLLM servicer and remove duplicate server.py - #662

Merged
CatherineSue merged 8 commits into
mainfrom
chang/grpc
Mar 6, 2026
Merged

CatherineSue merged 8 commits into
mainfrom
chang/grpc

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 6, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

The vLLM gRPC servicer had several efficiency issues (unnecessary allocations in logprobs building, unvectorized placeholder mask ops, numpy intermediate in tensor deserialization) and maintained a duplicate server.py that is now upstream in vLLM.

Solution

Optimize hot paths in the servicer, remove the duplicate server entrypoint, and consolidate kv_transfer_params validation.

Changes

  • Optimize tensor ops and logprobs building (servicer.py):

    • Guard _build_top_logprobs in _build_input_logprobs to avoid allocating empty proto objects for every prompt token when top logprobs aren't requested
    • Vectorize placeholder mask building with torch tensor ops instead of Python list comprehensions
    • Eliminate numpy intermediate in _tensor_from_proto — use torch.frombuffer(bytearray(...)) directly (1 copy instead of 2, removed numpy dependency)
    • Cache flat sizes tensor conversion to avoid redundant .flatten().to(torch.int64) calls
  • Remove duplicate server.py: The canonical entrypoint is now vllm.entrypoints.grpc_server upstream. Updated __init__.py and README.md references.

  • Move kv_transfer_params validation into _sampling_params_from_proto: Consolidates validation and logging into one place instead of inline in Generate(). Raises ValueError on invalid params, caught by existing error handler.

Test Plan

  • Existing gRPC e2e tests cover the servicer behavior
  • Changes are pure refactors — no functional behavior changes
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Documentation

    • README updated to recommend invoking the public vLLM gRPC entrypoint.
  • Refactor

    • Removed an internal gRPC server entry point and simplified public exports.
    • Optimized multimodal tensor handling and added caching to reduce overhead.
  • Bug Fixes

    • Strengthened validation for transfer-related parameters and guarded top-logprob construction.
  • Chores

    • CI workflow updated to include grpc_servicer paths for PR tests.

@CatherineSue
CatherineSue requested a review from njhill as a code owner March 6, 2026 21:05
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Mar 6, 2026
@coderabbitai

coderabbitai Bot commented Mar 6, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Removed the vLLM gRPC server module and its export, updated README to use the public gRPC entrypoint, and refactored servicer internals to use torch dtypes, add caching and vectorized masking, and tighten kv_transfer_params validation.

Changes

Cohort / File(s) Summary
Docs & exports
grpc_servicer/README.md, grpc_servicer/smg_grpc_servicer/vllm/__init__.py
README example now invokes vllm.entrypoints.grpc_server; removed serve_grpc import and its removal from __all__.
Server module removal
grpc_servicer/smg_grpc_servicer/vllm/server.py
Deleted entire module that provided serve_grpc() and main() (CLI parsing, AsyncLLM setup, gRPC aio server, reflection, signal handling, graceful shutdown).
Servicer internals
grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Switched tensor deserialization to use torch dtypes, added flat_sizes_cache, vectorized prompt/mask handling and pre-converted prompt tensors, validated/propagated kv_transfer_params, and guarded top-logprobs construction.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested labels

grpc, ci

Suggested reviewers

  • key4ng
  • slin1237
  • XinyueZhang369

Poem

🐰 I hopped through lines of code tonight,
The server tucked; the README writes,
Torch paws pat caches warm and bright,
Tokens march in neat, swift light,
A rabbit nods — all systems right.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main changes: refactoring the gRPC servicer for optimization and removing the duplicate server.py module, which aligns with the primary objectives of the PR.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/grpc

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request refactors the vLLM gRPC servicer to significantly improve its efficiency and streamline its architecture. It addresses several performance bottlenecks by optimizing tensor deserialization, logprobs construction, and placeholder mask generation, leading to a more performant system. Furthermore, it eliminates a redundant server entrypoint by adopting the upstream vLLM implementation and centralizes validation logic for kv_transfer_params, enhancing code clarity and reducing duplication.

Highlights

  • Performance Optimizations: Optimized tensor operations and logprobs building within the vLLM gRPC servicer, reducing unnecessary allocations and intermediate steps.
  • Code Consolidation: Removed the duplicate server.py entrypoint, now relying on the canonical upstream vLLM implementation, and updated related references.
  • Validation Centralization: Consolidated kv_transfer_params validation logic into a dedicated helper function for better maintainability and error handling.
Changelog
  • grpc_servicer/README.md
    • Updated the example command to reflect the new upstream server entrypoint.
  • grpc_servicer/smg_grpc_servicer/vllm/init.py
    • Removed the import and __all__ entry for the serve_grpc function.
  • grpc_servicer/smg_grpc_servicer/vllm/server.py
    • Deleted the entire file, as its functionality is now provided by the upstream vLLM.
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py
    • Removed the numpy import and replaced numpy.dtype with torch.dtype in _PROTO_DTYPE_MAP.
    • Modified _tensor_from_proto to directly use torch.frombuffer with bytearray, eliminating an intermediate numpy array.
    • Removed kv_transfer_params validation logic from the Generate method.
    • Introduced a flat_sizes_cache in _build_preprocessed_mm_inputs to prevent redundant tensor conversions.
    • Vectorized the placeholder mask building process in _build_preprocessed_mm_inputs using torch operations.
    • Moved kv_transfer_params validation and logging into the _sampling_params_from_proto function.
    • Added a conditional check in _build_input_logprobs to only call _build_top_logprobs when num_top_logprobs is requested.
Activity
  • No human activity has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a solid set of optimizations and refactorings for the vLLM gRPC servicer, including removing the numpy dependency for tensor deserialization, vectorizing mask building, caching tensor conversions, removing the duplicate server.py, and consolidating validation logic. However, a potential Server-Side Request Forgery (SSRF) vulnerability was identified in the handling of kv_transfer_params where the remote_host is not properly validated against an allow-list or restricted to authorized internal addresses. Additionally, there is one minor suggestion for a micro-optimization to further improve the code.

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/__init__.py`:
- Line 5: The change removed serve_grpc from the public exports causing a
compatibility break; restore a deprecated shim named serve_grpc in
grpc_servicer/smg_grpc_servicer/vllm/__init__.py that imports and calls the
current implementation (or forwards to VllmEngineServicer where appropriate),
add "serve_grpc" back into __all__ alongside "VllmEngineServicer", and mark the
shim with a deprecation warning (e.g., using warnings.warn) so consumers get
notified while existing imports continue to work for one release.

In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 455-459: The INFO-level logging of KV transfer parameters in the
request hot path (the logger.info call that prints
"kv_transfer_params={remote_host=%s, remote_port=%d}") should be removed or
demoted to debug and redact sensitive topology; update the logging inside the
Generate request path (in servicer.py where the kv_transfer_params logger.info
is used) to either delete the log statement or change it to logger.debug and
avoid emitting raw remote_host/remote_port (e.g., redact or replace with masked
values or a boolean flag like "kv_transfer_params_present"). Ensure the change
touches the exact logger.info invocation so the hot path no longer emits
INFO-level host/port details.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6c5f1c15-6b48-47e7-90bc-d8f26f904ebd

📥 Commits

Reviewing files that changed from the base of the PR and between 24b2424 and ca2282c.

📒 Files selected for processing (4)
  • grpc_servicer/README.md
  • grpc_servicer/smg_grpc_servicer/vllm/__init__.py
  • grpc_servicer/smg_grpc_servicer/vllm/server.py
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py
💤 Files with no reviewable changes (1)
  • grpc_servicer/smg_grpc_servicer/vllm/server.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/__init__.py
Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
grpc_servicer/smg_grpc_servicer/vllm/servicer.py (1)

449-464: ⚠️ Potential issue | 🟠 Major

Keep the validation, but drop the hot-path host/port INFO log.

The validation move is fine, but Line 455 still emits internal topology for every KV-transfer request on the Generate path. This should be debug-level at most and ideally redacted.

Suggested change
-            logger.info(
-                "kv_transfer_params={remote_host=%s, remote_port=%d}",
-                remote_host,
-                remote_port,
-            )
+            logger.debug("kv_transfer_params received")
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py` around lines 449 - 464, The
validation for kv_transfer_params (checking kv_transfer_params.remote_host and 1
<= kv_transfer_params.remote_port <= 65535) should remain, but replace the
hot-path logger.info call in servicer.py with a lower-verbosity, redacted debug
log: change logger.info(...) to logger.debug(...) and avoid printing raw
internal topology values (use placeholders like "<redacted>" or log only that
kv_transfer_params was provided without including remote_host/remote_port),
leaving extra_args construction (kv_transfer_params mapping) intact so behavior
is unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 449-464: The validation for kv_transfer_params (checking
kv_transfer_params.remote_host and 1 <= kv_transfer_params.remote_port <= 65535)
should remain, but replace the hot-path logger.info call in servicer.py with a
lower-verbosity, redacted debug log: change logger.info(...) to
logger.debug(...) and avoid printing raw internal topology values (use
placeholders like "<redacted>" or log only that kv_transfer_params was provided
without including remote_host/remote_port), leaving extra_args construction
(kv_transfer_params mapping) intact so behavior is unchanged.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 392b95c9-6452-4ef3-a9a6-595335705921

📥 Commits

Reviewing files that changed from the base of the PR and between ca2282c and 205ffb0.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 582-585: The condition checking whether to append top_logprobs is
functionally correct but stylistically inconsistent with the check in
_build_output_logprobs; change the conditional around proto.top_logprobs (the
block appending VllmEngineServicer._build_top_logprobs) to match the simpler
style used in _build_output_logprobs (e.g. use "if num_top_logprobs:"), so both
places use the same falsy-check pattern for num_top_logprobs.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 133b7d33-bd23-46c9-94eb-fe538cb86a73

📥 Commits

Reviewing files that changed from the base of the PR and between 205ffb0 and 968ac2a.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 42-47: The _tensor_from_proto function currently lets
torch.frombuffer()/reshape raise RuntimeError for malformed payloads, causing
server-side INTERNAL errors; wrap the deserialization/reshape logic in a
try/except that catches RuntimeError (and any torch-specific buffer/shape
errors) and re-raise them as ValueError with a clear message so clients get
INVALID_ARGUMENT; reference _tensor_from_proto, vllm_engine_pb2.TensorData and
_PROTO_DTYPE_MAP when locating the change.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 35fca93e-feb0-41b2-9393-b32c341b1625

📥 Commits

Reviewing files that changed from the base of the PR and between 968ac2a and e0fdfc8.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Guard _build_top_logprobs in _build_input_logprobs to avoid allocating
empty proto objects for every prompt token. Vectorize placeholder mask
building with torch ops. Eliminate numpy intermediate in tensor
deserialization. Cache flat sizes tensor conversion. Extract
_build_top_logprobs to module level.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
…points.grpc_server

Signed-off-by: Chang Su <chang.s.su@oracle.com>
…ams_from_proto

Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
Signed-off-by: Chang Su <chang.s.su@oracle.com>
@github-actions github-actions Bot added the ci CI/CD configuration changes label Mar 6, 2026
torch.topk() already returns entries sorted descending by logprob
value, and Python dict insertion order preserves this (3.7+).

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 54c10a81d7

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant