Skip to content

feat(serve): add unified --connection-mode and fix gRPC health checks - #335

Merged
slin1237 merged 2 commits into
mainfrom
serve-fix
Feb 5, 2026
Merged

slin1237 merged 2 commits into
mainfrom
serve-fix

Conversation

@slin1237

@slin1237 slin1237 commented Feb 5, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Add --connection-mode {grpc,http} argument (default: grpc) for unified connection mode across backends
  • sglang supports both grpc and http modes
  • vllm and trtllm only support grpc (error if http requested)
  • Move health_check and worker_url to base WorkerLauncher class to reduce duplication
  • Add grpcio and grpcio-health-checking dependencies for proper gRPC health checks
  • Remove unused dependencies (uvicorn, fastapi, aiohttp, orjson)
  • Update README with smg serve command documentation

Test plan

  • Test smg serve --backend sglang with default grpc mode
  • Test smg serve --backend sglang --connection-mode http
  • Test smg serve --backend vllm (grpc only)
  • Verify error when --backend vllm --connection-mode http

Summary by CodeRabbit

  • Documentation

    • Updated Python bindings README with Quick Start, installation, usage, serve options, and revised build/testing guidance.
  • New Features

    • Added CLI options: --connection-mode (grpc/http, default grpc), --host (default 127.0.0.1), --port, and --data-parallel-size (default 1).
  • Chores

    • Adjusted Python packaging metadata and dependencies to add gRPC support and remove unused web-framework packages.
  • Tests

    • Test coverage updated to validate new CLI args, defaults, and per-backend connection-mode behavior.

- Add --connection-mode {grpc,http} argument (default: grpc)
- sglang supports both grpc and http modes
- vllm and trtllm only support grpc (error if http requested)
- Move health_check and worker_url to base WorkerLauncher class
- Add grpcio and grpcio-health-checking dependencies
- Remove unused dependencies (uvicorn, fastapi, aiohttp, orjson)
- Update README with serve command documentation
@github-actions github-actions Bot added documentation Improvements or additions to documentation python-bindings Python bindings changes dependencies Dependency updates labels Feb 5, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request streamlines the smg serve command by introducing a unified --connection-mode argument, allowing users to specify either gRPC or HTTP for worker-router communication. It centralizes connection and health-check logic within the base WorkerLauncher class, enhancing code reusability and maintainability. The changes also ensure backend-specific connection constraints are enforced, update necessary dependencies for gRPC health checks, and provide clear documentation for the new functionality.

Highlights

  • Unified Connection Mode: Introduced a --connection-mode argument (grpc or http) for the smg serve command, standardizing how workers connect to the router.
  • Backend Compatibility Enforcement: Implemented checks to ensure vLLM and TensorRT-LLM backends are restricted to gRPC connections, raising an error if HTTP is attempted.
  • Centralized Worker Logic: Refactored health_check and worker_url methods into the base WorkerLauncher class, reducing code duplication and centralizing connection logic.
  • Dependency Management: Updated pyproject.toml to include grpcio and grpcio-health-checking for robust gRPC health checks, while removing unused HTTP-related dependencies (uvicorn, fastapi, aiohttp, orjson).
  • Documentation Update: The README.md has been significantly updated to include detailed "Quick Start" instructions and a comprehensive "Serve Options" table for the smg serve command.
Changelog
  • bindings/python/README.md
    • Added a new "Quick Start" section with installation instructions and usage examples for smg serve across different backends and connection modes.
    • Introduced a "Serve Options" table detailing command-line arguments like --backend, --connection-mode, --host, --port, --dp-size, and other worker-related options.
    • Updated the directory structure to reflect changes in Python source files.
    • Removed verbose comments from build and test instructions, and deprecated "Configuration" and "Notes" sections.
  • bindings/python/pyproject.toml
    • Removed aiohttp, orjson, uvicorn, and fastapi from the project's core dependencies.
    • Added grpcio and grpcio-health-checking as new core dependencies to support gRPC functionality.
  • bindings/python/src/smg/serve.py
    • Modified the WorkerLauncher base class: health_check and worker_url methods are no longer abstract and now contain default implementations that dynamically select between gRPC and HTTP based on the connection_mode argument.
    • Updated SglangWorkerLauncher.build_command to use the new unified --connection-mode argument.
    • Added explicit ValueError checks in VllmWorkerLauncher.build_command and TrtllmWorkerLauncher.build_command to prevent HTTP connection mode usage, as these backends are gRPC-only.
    • Removed redundant health_check and worker_url method overrides from SglangWorkerLauncher, VllmWorkerLauncher, and TrtllmWorkerLauncher subclasses.
    • Introduced --connection-mode, --host, and --port arguments to the add_serve_args function, providing unified control over router and worker connection settings.
    • Renamed the --dp-size argument to --data-parallel-size (with dest="data_parallel_size") while retaining --dp-size as an alias for backward compatibility and clarity.
    • Configured argparse.ArgumentParser with conflict_handler="resolve" in parse_serve_args to allow backend-specific arguments to gracefully override general serve arguments.
    • Updated internal calls within _launch_workers, _wait_healthy, and _build_router_args to pass the args object to health_check and worker_url methods, and to use data_parallel_size.
Activity
  • The pull request description indicates that it was "Generated with [Claude Code]". No other human activity such as comments, reviews, or explicit progress updates are provided in the context.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 5, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Refactors Python bindings: README expanded; pyproject switches HTTP stack to gRPC deps; serve logic updated to accept a connection mode and host/port/data-parallel-size args, converting health_check and worker_url to concrete, arg-aware implementations and propagating connection_mode through worker launch and tests. (50 words)

Changes

Cohort / File(s) Summary
Documentation
bindings/python/README.md
Expanded Quick Start, Installation, Usage and Serve Options; updated directory structure to reflect new CLI/serve entry points and removed legacy build/test manifest notes.
Project Metadata
bindings/python/pyproject.toml
Adjusted Python compatibility (>=3.9) and replaced HTTP-related dependencies (aiohttp, uvicorn, fastapi, orjson) with gRPC (grpcio, grpcio-health-checking); updated classifiers.
Serve Runtime & CLI
bindings/python/src/smg/serve.py
Made WorkerLauncher.health_check and worker_url concrete and accept argparse.Namespace for connection_mode, host, port; added CLI args --connection-mode, --host, --port, --data-parallel-size (alias --dp-size); per-backend enforcement of grpc-only modes and routing of health checks/URLs by mode.
Tests
bindings/python/tests/test_serve.py
Renamed dp_size → data_parallel_size; added connection_mode tests (default grpc and explicit http); updated expected URLs, health-check dispatch, host default, and backend mode restrictions across tests.

Sequence Diagram(s)

sequenceDiagram
  participant CLI as "CLI / User"
  participant Orch as "ServeOrchestrator"
  participant Launcher as "WorkerLauncher"
  participant Worker as "Worker (gRPC / HTTP)"
  participant Health as "Health Endpoint"

  CLI->>Orch: parse args (--connection-mode, --host, --port, --data-parallel-size)
  Orch->>Launcher: launch worker(s) with args, host, port
  Launcher->>Worker: start process (backend-specific)
  par health-check
    Launcher->>Health: perform health_check(args, host, port)
    Health-->>Launcher: healthy / unhealthy
  end
  Launcher-->>Orch: worker URLs (built using args.connection_mode)
  Orch-->>CLI: report ready / worker URLs
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 I hopped from HTTP to gRPC light,
Arguments guiding each worker's flight,
Health-checks tuned by mode and host,
CLI carrots served to the most,
New README crumbs for the devs' delight.

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 22.73% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main changes: adding a --connection-mode CLI option and fixing gRPC health checks, which are the primary objectives of this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch serve-fix

Comment @coderabbitai help to get the list of available commands and usage tips.

@CatherineSue

Copy link
Copy Markdown
Member

I think vLLM and TrtLLM may not have the standard grpc health check implemented yet

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the smg serve command to use a unified --connection-mode argument, which is a great improvement for consistency. The changes correctly move shared logic for health checks and worker URL generation into the base WorkerLauncher class, reducing code duplication. The argument parsing is also made more robust. My review includes a couple of suggestions for the README.md file to improve the accuracy and completeness of the documentation for both users and contributors.

Comment thread bindings/python/README.md Outdated
Comment thread bindings/python/README.md
@slin1237

slin1237 commented Feb 5, 2026

Copy link
Copy Markdown
Member Author

I think vLLM and TrtLLM may not have the standard grpc health check implemented yet

yes, the code fall back to channel being ready

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In `@bindings/python/pyproject.toml`:
- Around line 30-34: The pyproject declares Python 3.8 but the unpinned grpcio
and grpcio-health-checking now require Python >=3.9; either update the package
metadata to drop 3.8 (set requires-python to ">=3.9" and remove 3.8 from
classifiers) OR keep 3.8 support by pinning the dependencies in the dependencies
list to versions below 1.71.0 (e.g., change "grpcio" and
"grpcio-health-checking" entries to version-constrained strings like
grpcio<1.71.0 and grpcio-health-checking<1.71.0) so installs on Python 3.8
succeed.

In `@bindings/python/README.md`:
- Around line 38-48: Update the --connection-mode table entry to correctly state
backend restrictions: clarify that `--connection-mode` supports `grpc` (default)
and `http` for the `sglang` backend, but `vllm` and `trtllm` backends enforce
`grpc`-only (attempting `http` will error); reference the option names
`--connection-mode`, `--backend`, and backend values `sglang`, `vllm`, `trtllm`
so readers can see which backends allow HTTP vs which require gRPC.

In `@bindings/python/src/smg/serve.py`:
- Around line 319-335: The current CLI flag definition for "--host" sets
default="0.0.0.0" which exposes the router; change the default to "127.0.0.1"
(or "localhost") in the add_argument call that defines "--host" so the service
binds to localhost by default, and update the corresponding help string to
instruct users how to opt into public binding (e.g., "--host 0.0.0.0") if
required; locate the "--host" add_argument invocation in serve.py and modify the
default and help text accordingly.

Comment thread bindings/python/pyproject.toml
Comment thread bindings/python/README.md
Comment thread bindings/python/src/smg/serve.py
- Drop Python 3.8 support (grpcio 1.71.0+ requires Python ≥3.9)
- Fix README: clarify vllm/trtllm only support grpc (not sglang)
- Default --host to 127.0.0.1 for security (avoid exposing router)
- Update tests for new API:
  - dp_size → data_parallel_size
  - grpc_mode → connection_mode
  - health_check/worker_url now take args as first parameter
  - Default connection mode is grpc
@github-actions github-actions Bot added the tests Test changes label Feb 5, 2026
@slin1237
slin1237 merged commit 716f9e8 into main Feb 5, 2026
5 checks passed
@slin1237
slin1237 deleted the serve-fix branch February 5, 2026 21:53
@coderabbitai coderabbitai Bot mentioned this pull request Feb 18, 2026
3 tasks
ppraneth pushed a commit that referenced this pull request Feb 18, 2026
…#335)

Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates documentation Improvements or additions to documentation python-bindings Python bindings changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants