Skip to content

feat(grpc): add VllmHealthServicer for standard gRPC health checking (grpc.health.v1) - #885

Merged
slin1237 merged 3 commits into
smg-project:mainfrom
V2arK:honglin/vllm-health-servicer
Mar 24, 2026
Merged

slin1237 merged 3 commits into
smg-project:mainfrom
V2arK:honglin/vllm-health-servicer

Conversation

@V2arK

@V2arK V2arK commented Mar 24, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Problem

smg-grpc-servicer provides SGLangHealthServicer for SGLang's standard gRPC health check (grpc.health.v1.Health), but the vLLM side has no equivalent. vLLM gRPC deployments cannot use Kubernetes native gRPC health probes (livenessProbe.grpc / readinessProbe.grpc, GA since K8s 1.27).

Solution

Add VllmHealthServicer in smg_grpc_servicer/vllm/health_servicer.py, mirroring SGLangHealthServicer in structure and placement. Health status delegates to AsyncLLM.check_health() -- the same EngineClient protocol method used by vLLM's HTTP /health endpoint.

Companion vLLM PR: vllm-project/vllm#38016

Changes

  • New file: smg_grpc_servicer/vllm/health_servicer.py
    • VllmHealthServicer(health_pb2_grpc.HealthServicer) with TYPE_CHECKING type hint for AsyncLLM
    • Check(): calls await async_llm.check_health(), returns SERVING / NOT_SERVING / SERVICE_UNKNOWN. Logs exceptions via logger.exception().
    • Watch(): inlines status computation (avoids Check()'s context.set_code() side effect on streaming responses). Logs exceptions via logger.debug(exc_info=True). Single-yield matching SGLangHealthServicer behavior; persistent streaming deferred to a follow-up PR for both servicers.
    • set_not_serving(): sets _shutting_down flag for graceful shutdown. Called by vLLM's serve_grpc() in its finally block (companion PR).
    • Supports service names "" (liveness) and "vllm.grpc.engine.VllmEngine" (readiness)
  • Updated: smg_grpc_servicer/vllm/__init__.py to export VllmHealthServicer

Test Plan

Tests are in the companion vLLM PR (test file). All verified on NVIDIA H200 MIG (1g.18gb) with facebook/opt-125m.

Unit tests: 13/13 passed

pytest -xvs tests/entrypoints/test_grpc_health.py
# Test Verifies
1 test_check_serving_overall Check(service="") -> SERVING
2 test_check_serving_vllm_service Check(service="vllm.grpc.engine.VllmEngine") -> SERVING
3 test_check_not_serving_engine_errored check_health() raises -> NOT_SERVING
4 test_check_not_serving_shutting_down set_not_serving() -> NOT_SERVING, check_health not called
5 test_check_unknown_service_status Unknown service -> SERVICE_UNKNOWN
6 test_check_unknown_service_grpc_code Unknown service -> gRPC NOT_FOUND code + details
7 test_check_shutting_down_overrides_healthy Shutdown flag takes priority over healthy engine
8 test_check_logs_exception_on_error logger.exception() called on failure
9 test_watch_yields_serving Watch yields SERVING
10 test_watch_yields_not_serving Watch yields NOT_SERVING on error
11 test_watch_yields_exactly_once Watch produces exactly 1 response
12 test_watch_unknown_service Watch unknown service -> SERVICE_UNKNOWN, context.set_code NOT called
13 test_set_not_serving_sets_flag _shutting_down flag set correctly

E2E tests: 10/10 scenarios passed

# Scenario Result
1 SERVING (overall + service + grpc_health_probe + probe with timeout) SERVING, exit 0
2 Reflection includes grpc.health.v1.Health
3 Unknown service Code: NotFound
4 Watch stream SERVING, exit 0
5 Custom HealthCheck coexistence {"healthy":true}
6 SIGTERM shutdown connection refused
7 Pre-established channel + SIGTERM UNAVAILABLE
8 Invalid model server fails, port closed
9 Stability (20x sequential) 20/20 SERVING
10 Concurrency (10x parallel) 10/10 SERVING
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • Python changes only -- no Rust changes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • New Features
    • Added a health-check service for the vLLM engine and overall server: reports serving/not-serving (including during shutdown), supports a streaming watch that emits current status, and returns a clear error for unknown service queries.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a gap in vLLM's gRPC capabilities by introducing a standard health check servicer. This enhancement is crucial for operational environments, particularly those leveraging Kubernetes, as it provides a standardized and robust mechanism for monitoring the health and readiness of vLLM gRPC services, thereby improving reliability and manageability.

Highlights

  • New Feature: VllmHealthServicer: Introduced VllmHealthServicer to provide standard gRPC health checking (grpc.health.v1) for vLLM deployments, enabling native Kubernetes liveness and readiness probes.
  • Kubernetes Compatibility: The new servicer allows vLLM gRPC deployments to integrate seamlessly with Kubernetes native gRPC health probes (livenessProbe.grpc / readinessProbe.grpc), which became GA in K8s 1.27.
  • Health Status Delegation: Health status is determined by delegating to AsyncLLM.check_health(), consistent with vLLM's HTTP /health endpoint.
  • Service Levels Supported: The VllmHealthServicer supports two service levels: overall server health (for liveness probes) and VllmEngine service health (for readiness probes).
  • Graceful Shutdown: A set_not_serving() method was added to mark all services as NOT_SERVING during graceful shutdown.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 24, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: bdb04172-eb98-4218-941c-0c37135833ed

📥 Commits

Reviewing files that changed from the base of the PR and between fef5e88 and b365f6f.

📒 Files selected for processing (2)
  • grpc_servicer/smg_grpc_servicer/vllm/__init__.py
  • grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py

📝 Walkthrough

Walkthrough

Added a new vLLM gRPC health servicer VllmHealthServicer (implements Check and Watch), tracks shutdown via set_not_serving(), reports health for overall server ("") and vllm.grpc.engine.VllmEngine by delegating to async_llm.check_health(), and exported it from the vLLM package.

Changes

Cohort / File(s) Summary
vLLM package export
grpc_servicer/smg_grpc_servicer/vllm/__init__.py
Export list updated to include VllmHealthServicer alongside VllmEngineServicer.
gRPC Health Servicer
grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py
New VllmHealthServicer implementing Check and Watch, shutdown tracking via set_not_serving(), supports services "" and vllm.grpc.engine.VllmEngine, delegates to async_llm.check_health(), returns appropriate gRPC health statuses and NOT_FOUND for unknown services.

Sequence Diagram

sequenceDiagram
    autonumber
    actor Client
    participant VllmHealthServicer as "VllmHealthServicer\n(rpc)"
    participant AsyncLLM as "async_llm"
    participant GRPC_CTX as "gRPC Context"

    Client->>VllmHealthServicer: Check(HealthCheckRequest(service))
    alt _shutting_down == True
        VllmHealthServicer-->>Client: HealthCheckResponse(NOT_SERVING)
    else Supported service ("" or "vllm.grpc.engine.VllmEngine")
        VllmHealthServicer->>AsyncLLM: await check_health()
        alt check_health succeeds
            AsyncLLM-->>VllmHealthServicer: success
            VllmHealthServicer-->>Client: HealthCheckResponse(SERVING)
        else check_health raises
            AsyncLLM-->>VllmHealthServicer: exception
            VllmHealthServicer-->>Client: HealthCheckResponse(NOT_SERVING)
        end
    else Unknown service
        VllmHealthServicer->>GRPC_CTX: set_code(NOT_FOUND), set_details("Unknown service: " + request.service)
        VllmHealthServicer-->>Client: HealthCheckResponse(SERVICE_UNKNOWN)
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related issues

Possibly related PRs

Suggested reviewers

  • slin1237
  • CatherineSue
  • njhill

Poem

🐰 I hopped to check the engine's heart,
I listened close to every part,
I waved a flag when evening came,
I answered health and signed my name,
Then nibbled code and hopped away.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main change: adding VllmHealthServicer for gRPC health checking, which is the primary focus of this PR.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new VllmHealthServicer to implement the standard gRPC health check protocol for vLLM, enabling Kubernetes probes to monitor the health of the AsyncLLM engine. The servicer supports both overall server health and specific VllmEngine service health checks. Feedback includes adding a type hint for the async_llm parameter in the constructor for improved type safety and clarity. Additionally, a critical issue was identified in the Watch method, where its current implementation violates the gRPC Health Checking Protocol by incorrectly setting the RPC status for unknown services; a direct re-implementation of the health checking logic within Watch is suggested to ensure protocol compliance.

Comment thread grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py Outdated
Comment thread grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py`:
- Around line 88-110: The Watch method must not call Check (which sets context
code to NOT_FOUND) and must be converted from a single-yield to a persistent
stream: compute the initial status from your internal status store (e.g. lookup
in self._status_map or equivalent) and if missing send a
HealthCheckResponse(status=health_pb2.HealthCheckResponse.SERVICE_UNKNOWN)
without calling context.set_code(NOT_FOUND); yield that initial
HealthCheckResponse, then enter a loop that polls for status changes (or awaits
notifications) and yields new HealthCheckResponse messages whenever the status
changes, sleeping briefly between polls (e.g. asyncio.sleep(poll_interval)) and
terminating only when the gRPC context is cancelled; ensure you reference Watch,
Check, HealthCheckResponse and avoid invoking Check so the stream stays open for
unknown services.
- Around line 45-48: The VllmHealthServicer.set_not_serving() is never wired
into the server lifecycle so shutdown probes stay SERVING; mirror the sglang
server wiring: create and register a VllmHealthServicer instance with the gRPC
server, pass that instance into VllmEngineServicer on construction/registration,
and ensure the global graceful-shutdown handler calls
VllmHealthServicer.set_not_serving() before stopping the server (i.e., call
set_not_serving() in the same shutdown sequence that currently calls
server.stop/graceful drain as in the sglang example); look for
VllmHealthServicer, VllmEngineServicer and the server startup/shutdown block to
add the instantiation, registration, and shutdown call.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 468496dc-afcd-4e82-90db-684ff3ea1f4d

📥 Commits

Reviewing files that changed from the base of the PR and between 59f49d1 and 300bc3f.

📒 Files selected for processing (2)
  • grpc_servicer/smg_grpc_servicer/vllm/__init__.py
  • grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py
Comment on lines +88 to +110
async def Watch(
self,
request: health_pb2.HealthCheckRequest,
context: grpc.aio.ServicerContext,
) -> AsyncIterator[health_pb2.HealthCheckResponse]:
"""
Streaming health check - sends current status once.

For now, sends current status once (Kubernetes doesn't use Watch).
A full implementation would monitor status changes and stream updates.

Args:
request: Contains service name
context: gRPC context

Yields:
HealthCheckResponse messages
"""
service_name = request.service
logger.debug(f"Health watch request for service: '{service_name}'")

response = await self.Check(request, context)
yield response

@coderabbitai coderabbitai Bot Mar 24, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🌐 Web query:

In grpc.health.v1, what is the expected behavior of Health/Watch for unknown services, and should Watch be a continuous stream of status updates instead of a single response?

💡 Result:

For unknown services in grpc.health.v1, the Health/Watch RPC has a specific expected behavior distinct from the Check RPC. When Watch receives a HealthCheckRequest for an unknown service, the server immediately sends a single HealthCheckResponse with status SERVICE_UNKNOWN (value 3) but does NOT terminate the call. The stream remains open, and the server MUST send subsequent messages if the service becomes known later (e.g., changes to SERVING or NOT_SERVING). If the stream terminates with UNIMPLEMENTED, clients should not retry; other terminations trigger retries with backoff. Yes, Watch is designed as a continuous bidirectional streaming RPC (server-streaming from the proto definition: rpc Watch(HealthCheckRequest) returns (stream HealthCheckResponse)). The server sends an initial status immediately upon request, then additional updates whenever the service's status changes over time. It is explicitly not a single unary response but a persistent stream for ongoing monitoring.

Citations:


🏁 Script executed:

cat -n grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py | head -120

Repository: lightseekorg/smg

Length of output: 4741


Watch implementation violates gRPC health.v1 streaming semantics in two ways.

  1. For unknown services: Line 109 calls Check(), which sets the gRPC context code to NOT_FOUND (line 84). This terminates the stream with a NOT_FOUND error. Per the spec, Watch should send a SERVICE_UNKNOWN response and keep the stream open without terminating.

  2. Single response instead of stream: Line 110 yields once then closes. Watch must be a persistent stream that sends initial status immediately and continues to emit updates when service status changes, not a single unary response.

The suggested fix addresses both issues by bypassing Check() (avoiding the problematic NOT_FOUND context code) and implementing a continuous stream with periodic health checks.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py` around lines 88 -
110, The Watch method must not call Check (which sets context code to NOT_FOUND)
and must be converted from a single-yield to a persistent stream: compute the
initial status from your internal status store (e.g. lookup in self._status_map
or equivalent) and if missing send a
HealthCheckResponse(status=health_pb2.HealthCheckResponse.SERVICE_UNKNOWN)
without calling context.set_code(NOT_FOUND); yield that initial
HealthCheckResponse, then enter a loop that polls for status changes (or awaits
notifications) and yields new HealthCheckResponse messages whenever the status
changes, sleeping briefly between polls (e.g. asyncio.sleep(poll_interval)) and
terminating only when the gRPC context is cancelled; ensure you reference Watch,
Check, HealthCheckResponse and avoid invoking Check so the stream stays open for
unknown services.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Acknowledged. The current single-yield Watch matches SGLangHealthServicer in the same repo (sglang/health_servicer.py:165). Converting Watch to a persistent stream is a larger change that should be done for both SGLang and vLLM servicers together as a follow-up PR. The context.set_code() side effect was already fixed in the previous commit by inlining the status computation.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Check the current state of the vLLM health servicer Watch method
cat -n grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py

echo "---SGLang health servicer Watch---"
# Check the SGLang health servicer Watch method for comparison
grep -n "Watch\|yield\|set_code\|Check" grpc_servicer/smg_grpc_servicer/sglang/health_servicer.py

Length of output: 8352


@V2arK, confirmed — the current branch already addresses both points:

  1. context.set_code() side effect — fixed. Lines 113–130 of the current health_servicer.py inline the status computation in Watch with an explicit comment explaining the rationale, so Check() is no longer called from Watch. This is actually an improvement over SGLangHealthServicer, which still delegates to self.Check() in its Watch (line 164 of sglang/health_servicer.py).

  2. Single-yield Watch — intentional. The pattern is consistent with SGLangHealthServicer, and the docstring documents the limitation and rationale. The follow-up refactor covering both servicers together makes sense.

Noting the design decision for future reviews.


✏️ Learnings added
Learnt from: V2arK
Repo: lightseekorg/smg PR: 885
File: grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py:0-0
Timestamp: 2026-03-24T18:53:57.311Z
Learning: In repo lightseekorg/smg, the single-yield Watch() pattern in grpc health servicers is intentional and consistent across both `grpc_servicer/smg_grpc_servicer/sglang/health_servicer.py` and `grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py`. Kubernetes only uses the Check RPC, so Watch yields one status response then closes. A persistent-stream Watch is a known future improvement to be implemented for both servicers together in a follow-up PR. Do not flag the single-yield Watch as a protocol violation in these files.

Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.


🧠 Learnings used
Learnt from: zhaowenzi
Repo: lightseekorg/smg PR: 807
File: model_gateway/src/routers/openai/responses/streaming.rs:821-855
Timestamp: 2026-03-18T21:57:03.433Z
Learning: In repo lightseekorg/smg, file model_gateway/src/routers/openai/responses/streaming.rs: The early `return` statements inside `handle_streaming_with_tool_interception` (on `tx` send failures in `forward_streaming_event`, `send_mcp_list_tools_events`, and the `is_in_progress`/`mcp_list_tools_sent` branches) are pre-existing behavior that predates PR `#807`. They cause the persistence phase (final response / conversation-backed storage writes at the end of the tool loop) to be skipped when the client disconnects mid-stream with `store=true` or a conversation-backed request. This is a known pre-existing gap in the MCP streaming path, not a regression introduced by the storage context header changes in PR `#807`.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 638
File: grpc_servicer/smg_grpc_servicer/vllm/servicer.py:321-323
Timestamp: 2026-03-05T04:48:40.754Z
Learning: In `grpc_servicer/smg_grpc_servicer/vllm/servicer.py`, the `last_receive_timestamp` field in `GetServerInfoResponse` is intentionally set to `time.time()` at read time (with an inline `# TODO looks wrong?` comment). This is a pre-existing TODO carried over from vLLM and is not to be fixed within the current PR scope.

Learnt from: hyeongyun0916
Repo: lightseekorg/smg PR: 784
File: grpc_servicer/smg_grpc_servicer/vllm/render_servicer.py:65-70
Timestamp: 2026-03-19T10:00:53.513Z
Learning: In repo lightseekorg/smg, the convention across all gRPC servicer implementations (e.g., grpc_servicer/smg_grpc_servicer/vllm/servicer.py, grpc_servicer/smg_grpc_servicer/vllm/render_servicer.py) is to NOT add an explicit `return` after `await context.abort(...)` calls, because `context.abort()` raises `grpc.aio.AbortError` internally making any subsequent return unreachable dead code. This is an intentional consistency decision across the codebase.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 690
File: model_gateway/src/core/worker_registry.rs:667-670
Timestamp: 2026-03-10T05:04:49.809Z
Learning: In repo lightseekorg/smg, file model_gateway/src/core/worker_registry.rs: `any_external_worker_supports_model` intentionally uses `healthy_only = true` for two reasons: (1) The 503 ("service unavailable") path in `select_worker_for_model` is specifically for the circuit-breaker case — healthy workers whose circuit breaker is open — while workers failing health checks fall through to 404. (2) Unhealthy workers have stale model lists (models registered at startup may no longer be accurate) and should not be trusted for model existence checks. This design follows the upstream sglang pattern from sgl-project/sglang#15611. Do not flag `healthy_only = true` in `any_external_worker_supports_model` as a bug.

Learnt from: hyeongyun0916
Repo: lightseekorg/smg PR: 784
File: grpc_servicer/smg_grpc_servicer/vllm/render_servicer.py:98-103
Timestamp: 2026-03-19T10:01:20.396Z
Learning: In repo lightseekorg/smg, the convention across all gRPC servicer implementations (e.g., vllm/servicer.py, vllm/render_servicer.py) is to NOT add an explicit `return` after `await context.abort(...)` calls. The codebase relies on `grpc.aio.AbortError` being raised by `context.abort()` to stop execution, and omitting the explicit `return` is an intentional, consistent design decision.

Learnt from: CatherineSue
Repo: lightseekorg/smg PR: 690
File: model_gateway/src/core/worker_registry.rs:667-670
Timestamp: 2026-03-10T04:59:38.803Z
Learning: In repo lightseekorg/smg, file model_gateway/src/core/worker_registry.rs: `any_external_worker_supports_model` intentionally uses `healthy_only = true`. The 503 ("service unavailable") path in `select_worker_for_model` is specifically for the circuit-breaker case — healthy workers whose circuit breaker is open. Workers that are genuinely unhealthy (failing health checks) are intentionally excluded: the model falls through to 404 for those. This design follows the upstream sglang pattern established in sgl-project/sglang#15611 ("[model-gateway] return 503 when all workers are circuit-broken"). Do not flag the `healthy_only = true` argument in this method as a bug.

Learnt from: pallasathena92
Repo: lightseekorg/smg PR: 687
File: model_gateway/src/routers/openai/realtime/webrtc.rs:238-289
Timestamp: 2026-03-11T01:29:56.655Z
Learning: In repo lightseekorg/smg, file model_gateway/src/routers/openai/realtime/webrtc.rs: `Metrics::record_router_request` is already emitted in router.rs (around line 1152) before `handle_realtime_webrtc` is called, and `Metrics::record_router_error` is emitted inside `handle_realtime_webrtc` for the no-workers case. The missing instrumentation is success/duration recording after `setup_and_spawn_bridge` returns — this is a metrics improvement deferred to a follow-up PR, not a correctness gap. Do not flag missing success/duration metrics as a blocking issue for PR `#687`.

Learnt from: key4ng
Repo: lightseekorg/smg PR: 867
File: tui/src/app.rs:571-596
Timestamp: 2026-03-23T20:26:39.459Z
Learning: In lightseekorg/smg, tui/src/app.rs, the "toggle-health" command uses `worker.is_healthy` (not a `disable_health_check` field) to compute the new `disable_health_check` value because the worker list API does not expose `disable_health_check` in its response. The intended behavior is: healthy worker → disable health checks (maintenance mode), unhealthy worker → re-enable health checks (recovery). Do not flag this as a bug; it is an intentional heuristic given the API limitation.

Learnt from: hyeongyun0916
Repo: lightseekorg/smg PR: 784
File: grpc_servicer/smg_grpc_servicer/vllm/render_servicer.py:94-96
Timestamp: 2026-03-17T09:41:10.752Z
Learning: Enforce the repository-wide convention for gRPC INTERNAL errors in all servicer implementations: log the exception on the server (using logger.exception or logger.error) and then call context.abort(grpc.StatusCode.INTERNAL, str(e)) without replacing the message. Do not substitute a generic message. This applies to all servicer files (e.g., vllm/servicer.py, sglang/servicer.py, vllm/render_servicer.py) and should be treated as an intentional, consistent design decision across the codebase.

If you found this review helpful, would you consider giving us a shout-out on X?

Thank you for using CodeRabbit!

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A follow up PR sounds good. @V2arK lmk if you want to take on this task. Appreciate the contribution

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Seems like the humans are having a chat. I'll hop back into my burrow for now. If you need me again, just tag @coderabbitai in a new comment, and I'll come hopping out!

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can take a look.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CatherineSue For the Watch continuous stream follow-up, I'm considering two approaches:

  1. Inline: each servicer implements its own Watch loop independently (duplicated but self-contained)
  2. Mixin: extract the Watch loop into a shared HealthWatchMixin, servicers only provide_compute_watch_status() and _is_shutting_down()

Given there are only two servicers today, I'm leaning toward (1) for simplicity. Any preference?

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py`:
- Around line 34-42: Update the __init__ signature of the Health servicer to
type-hint the async_llm parameter as "AsyncLLM" (use a string forward reference
to avoid circular imports) so it matches VllmEngineServicer; if you prefer
explicit typing, add a TYPE_CHECKING block to import AsyncLLM (from vllm import
AsyncLLM) guarded by typing.TYPE_CHECKING and annotate self.async_llm
accordingly in the __init__ of the health servicer class.
- Around line 119-120: The Watch method currently swallows all exceptions;
mirror the Check method's behavior by logging the exception instead of silently
ignoring it: inside the except Exception block in HealthServicer.Watch, call
logger.debug (or logger.exception if you prefer full error visibility) with a
brief message and exc_info=True so the traceback is captured when debugging is
enabled; this keeps behavior consistent with Check (which uses logger.exception)
while avoiding noisy logs at higher levels.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 0a60ca28-8fd2-4d2c-a97b-ab9cba214558

📥 Commits

Reviewing files that changed from the base of the PR and between 300bc3f and 77fbafd.

📒 Files selected for processing (1)
  • grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py

Comment thread grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py Outdated
Comment thread grpc_servicer/smg_grpc_servicer/vllm/health_servicer.py
@V2arK

V2arK commented Mar 24, 2026

Copy link
Copy Markdown
Collaborator Author

Note: the 3 CI failures (build-wheel, unit-tests, finish) are unrelated to this PR -- they all fail at "Build WASM test fixtures" with a Rust compilation error in heck 0.4.1 vs unicode-segmentation 1.13.0 (struct UnicodeWords is private). This is a pre-existing infra issue on main. All Python-relevant checks pass: pre-commit, python-lint, grpc-proto-build-check, DCO.

@CatherineSue
CatherineSue force-pushed the honglin/vllm-health-servicer branch from 57d0352 to fef5e88 Compare March 24, 2026 19:11
@CatherineSue

Copy link
Copy Markdown
Member

WASM should be fixed. Rebased the commits.

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for contributing this critical piece. LGTM

V2arK added 3 commits March 24, 2026 15:15
…protocol

Signed-off-by: Honglin Cao <Caohonglin317@hotmail.com>
- Add forward reference type hint for async_llm parameter
- Inline health check logic in Watch() to avoid Check()'s
  context.set_code() side effect on streaming responses

Signed-off-by: Honglin Cao <Caohonglin317@hotmail.com>
- Add TYPE_CHECKING guard for AsyncLLM type hint on __init__
- Add logger.debug(exc_info=True) in Watch except block for
  consistency with Check's logger.exception()

Signed-off-by: Honglin Cao <Caohonglin317@hotmail.com>
@V2arK
V2arK force-pushed the honglin/vllm-health-servicer branch from fef5e88 to b365f6f Compare March 24, 2026 19:16
@V2arK

V2arK commented Mar 24, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for contributing this critical piece. LGTM

My pleasure!

@slin1237
slin1237 merged commit ed383f5 into smg-project:main Mar 24, 2026
20 of 25 checks passed
smfirmin pushed a commit to smfirmin/smg that referenced this pull request Apr 2, 2026
…(grpc.health.v1) (smg-project#885)

Signed-off-by: Honglin Cao <Caohonglin317@hotmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority:high High priority

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants