Skip to content

refactor(grpc): use EngineClient interface instead of AsyncLLM in vLLM servicer - #949

Merged
CatherineSue merged 3 commits into
mainfrom
refactor/vllm-servicer-engine-client
Mar 27, 2026
Merged

CatherineSue merged 3 commits into
mainfrom
refactor/vllm-servicer-engine-client

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 27, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

The vLLM gRPC servicer directly depends on AsyncLLM, a concrete implementation, rather than the EngineClient protocol interface. This was flagged in vllm-project/vllm#36169 by @njhill.

The servicer also accessed output_processor.get_num_unfinished_requests() which is not on EngineClient, and reached through input_processor.input_preprocessor.renderer instead of using EngineClient.renderer directly.

Solution

  • Replace AsyncLLM type with EngineClient protocol
  • Use self.engine.renderer.process_for_engine() directly (available on EngineClient since vLLM v0.18.0)
  • Remove output_processor.get_num_unfinished_requests() from GetServerInfo — SMG strips active_requests, is_paused, last_receive_timestamp, uptime_seconds, and server_type during worker metadata discovery (discover_metadata.rs:268-276), so these fields were never consumed. Only kv_connector and kv_role are returned (used for PD routing).
  • Keep async_llm as the constructor param name to avoid breaking the caller in vLLM's grpc_server.py

Changes

  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py:
    • Type annotation: AsyncLLM → EngineClient
    • Remove from vllm.v1.engine.async_llm import AsyncLLM
    • Renderer access: self.engine.input_processor.input_preprocessor.renderer → self.engine.renderer
    • Text prompt: use renderer.process_for_engine() instead of input_preprocessor.preprocess()
    • GetServerInfo: return only kv_connector and kv_role

Test Plan

  • vllm serve --grpc with tokenized input (normal SMG flow) → works as before
  • vllm serve --grpc with text input → renderer handles tokenization
  • GetServerInfo → returns kv_connector/kv_role, other fields default to zero
  • Verify SMG worker discovery still works (unused fields are stripped anyway)
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Refactor
    • Request preprocessing, streaming, and runtime operations now use a new backend engine; health, abort, and model-info operations use the same engine.
  • Behavior Change
    • Tokenized inputs use the engine's renderer; non-tokenized text prompts are passed through without prior preprocessing.
    • Server-info payload reduced (fewer runtime/telemetry fields returned).
  • Chore
    • Bumped optional vllm dependency to >=0.17.0.

@CatherineSue
CatherineSue requested a review from slin1237 as a code owner March 27, 2026 18:10
@coderabbitai

coderabbitai Bot commented Mar 27, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Vllm servicer replaced stored AsyncLLM with a vLLM EngineClient (self.engine); Generate, preprocessing, streaming, health/abort/model-info paths now call self.engine. GetServerInfo removed several server-state fields. vLLM optional dependency bumped to >=0.17.0. (≤50 words)

Changes

Cohort / File(s) Summary
Servicer: EngineClient migration
grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Replaced AsyncLLM with EngineClient (self.engine) across constructor and call sites. Consolidated arrival_time calculation, moved renderer calls to self.engine.renderer.process_for_engine(...), removed non-tokenized preprocessing path, switched generate loop to self.engine.generate(...), and updated HealthCheck/Abort/GetModelInfo/GetServerInfo to use self.engine and its configs. Model dtype lookup now uses self.engine.model_config.dtype.
Dependency bump
grpc_servicer/pyproject.toml
Updated optional dependency vllm constraint from vllm>=0.16.0 to vllm>=0.17.0.

Sequence Diagram(s)

sequenceDiagram
    autonumber
    participant Client as Client
    participant Servicer as VllmEngineServicer
    participant Renderer as Renderer
    participant Engine as EngineClient
    participant Model as vLLM_Model

    Client->>Servicer: Send Generate request
    Servicer->>Servicer: capture arrival_time
    Servicer->>Renderer: process_for_engine(inputs, arrival_time)
    Renderer-->>Servicer: preprocessed inputs
    Servicer->>Engine: generate(preprocessed inputs)
    Engine->>Model: request generation
    Model-->>Engine: tokens/chunks
    Engine-->>Servicer: streamed responses
    Servicer-->>Client: stream responses
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~40 minutes

Possibly related PRs

Suggested labels

grpc

Suggested reviewers

  • njhill
  • slin1237

Poem

🐰 A hop, the engine takes the lead,
Arrival stamped for every feed,
Renderer trims and hands the light,
Tokens stream through day and night,
Rabbit cheers — the servicer's speed!

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main refactoring: replacing AsyncLLM with EngineClient in the vLLM servicer, which is the primary change across all modifications.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch refactor/vllm-servicer-engine-client

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request migrates the VllmEngineServicer to use EngineClient instead of AsyncLLM, involving the renaming of the internal engine instance and updating the Generate method to utilize the engine's renderer for prompt processing. Additionally, several fields were removed from the GetServerInfo response. The review feedback recommends renaming the init parameter for consistency with the new type and refactoring the Generate method to capture a single timestamp for request arrival to ensure consistency and reduce redundant calls.

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py
Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f0d6d90b5f

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
grpc_servicer/smg_grpc_servicer/vllm/servicer.py (1)

295-320: 🧹 Nitpick | 🔵 Trivial

Consider populating uptime_seconds and server_type for consistency.

Per PR objectives, SMG strips certain fields during worker discovery, so omitting them is acceptable. However, uptime_seconds is trivially computable from self.start_time, and server_type is a static string. Populating these would maintain parity with the SGLang servicer (see sglang/servicer.py:490-498) and support any non-SMG consumers of this RPC.

♻️ Optional: populate additional fields
         return vllm_engine_pb2.GetServerInfoResponse(
             kv_connector=kv_connector,
             kv_role=kv_role,
+            uptime_seconds=time.time() - self.start_time,
+            server_type="vllm-grpc",
         )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py` around lines 295 - 320,
GetServerInfo currently returns only kv_connector/kv_role; compute
uptime_seconds as int(time.time() - self.start_time) and set server_type to the
static string used elsewhere (e.g., "vllm") before returning the
vllm_engine_pb2.GetServerInfoResponse. Update the GetServerInfo method to
import/use time, reference self.start_time to calculate uptime_seconds, and
include server_type and uptime_seconds in the returned GetServerInfoResponse
alongside kv_connector and kv_role.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 295-320: GetServerInfo currently returns only
kv_connector/kv_role; compute uptime_seconds as int(time.time() -
self.start_time) and set server_type to the static string used elsewhere (e.g.,
"vllm") before returning the vllm_engine_pb2.GetServerInfoResponse. Update the
GetServerInfo method to import/use time, reference self.start_time to calculate
uptime_seconds, and include server_type and uptime_seconds in the returned
GetServerInfoResponse alongside kv_connector and kv_role.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 962eafd2-77ab-48fc-a787-35e2844c5e73

📥 Commits

Reviewing files that changed from the base of the PR and between 5a593cf and f0d6d90.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

Replace the concrete AsyncLLM dependency with the EngineClient protocol
interface, as suggested in vllm-project/vllm#36169.

Changes:
- Constructor takes EngineClient instead of AsyncLLM
- Use self.engine.renderer directly instead of reaching through
  input_processor.input_preprocessor.renderer
- Text prompt path now uses renderer.process_for_engine() which
  handles tokenization internally
- Remove output_processor.get_num_unfinished_requests() from
  GetServerInfo — SMG strips active_requests, is_paused,
  last_receive_timestamp, uptime_seconds, and server_type during
  worker metadata discovery (discover_metadata.rs:268-276), so
  these fields were never consumed
- Only kv_connector and kv_role are returned (used for PD routing)

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@CatherineSue
CatherineSue force-pushed the refactor/vllm-servicer-engine-client branch from f0d6d90 to 9bd34be Compare March 27, 2026 18:21
@github-actions github-actions Bot added the dependencies Dependency updates label Mar 27, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9bd34be3a5

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 121-124: The code incorrectly passes raw text via
self.engine.renderer.process_for_engine({"prompt": request.text}, ...) which
expects a tokenized TokensPrompt; instead, call the engine's generate flow with
raw text and tokenization parameters: use self.engine.generate(...) passing
request.text (or a PromptType containing the raw text) plus the prepared
tokenization_kwargs (the variable at line ~134) so the engine handles
tokenization; keep the tokenized path unchanged (which builds TokensPrompt and
calls process_for_engine), but for the text path remove the process_for_engine
call and invoke engine.generate with request.text and tokenization_kwargs to
match EngineClient's expected usage.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6bcb7f9f-7239-4749-8b31-c06bdf5644f2

📥 Commits

Reviewing files that changed from the base of the PR and between f0d6d90 and 9bd34be.

📒 Files selected for processing (2)
  • grpc_servicer/pyproject.toml
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated
Address review feedback from #949:

- Reject text prompts with a clear ValueError instead of passing
  them to process_for_engine which expects tokenized input. SMG
  always sends tokenized input via gRPC, so this path was dead
  code. The ValueError is caught and mapped to INVALID_ARGUMENT.

- Consolidate time.time() into a single arrival_time variable
  reused across all branches.

- Bump vllm dependency to >= 0.17.0 (process_for_engine added
  in v0.17.0, not v0.15.0 as previously stated).

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@CatherineSue
CatherineSue requested a review from njhill as a code owner March 27, 2026 18:36
Signed-off-by: Chang Su <chang.s.su@oracle.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ff4d5ed7e6

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/servicer.py`:
- Around line 309-319: Add a brief in-code comment above the return of
vllm_engine_pb2.GetServerInfoResponse explaining that active_requests,
is_paused, last_receive_timestamp, uptime_seconds, and server_type are
intentionally omitted and left at their proto defaults because SMG strips these
fields during worker metadata discovery (so the servicer only returns
kv_connector/kv_role derived from self.engine.vllm_config.kv_transfer_config);
reference the existing local symbols (GetServerInfoResponse, kv_connector,
kv_role, and self.engine.vllm_config.kv_transfer_config) so future maintainers
understand this is deliberate and not an oversight.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 22e41623-bb61-4d9c-b776-30d5c7a2c35f

📥 Commits

Reviewing files that changed from the base of the PR and between ff4d5ed and 239cae0.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/servicer.py
@CatherineSue
CatherineSue merged commit 34edd2b into main Mar 27, 2026
84 of 90 checks passed
@CatherineSue
CatherineSue deleted the refactor/vllm-servicer-engine-client branch March 27, 2026 19:24
smfirmin pushed a commit to smfirmin/smg that referenced this pull request Apr 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant