Repository navigation
Implement Standalone gRPC Server for SGLang Python Scheduler - #10283
Conversation
- It is uncommon to establish a handshake connection
- Remove Initialize method and run compile proto to update auto-generated code - Remove flush cache - Remove open_session and close_session
- Remove all detailed health metrics - Keep HealthCheck same as /health_generate in http_server - Clean up proto and re-compile
…in grpc_server.py
- Proper dependency injection: implement a GrpcSignalHandler instead of using SignalHanlder in tokenizer_manager.py - Have GrpcSignalHandler to simply gracefully exit with scheduler failure instead of dumping crash info
- When tokenizer is init (skip_tokenizer_init is False), scheduler sends results to detoeknizer_ipc_name
- Only keep tokenized
There was a problem hiding this comment.
Summary of Changes
Hello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!
This pull request introduces a new gRPC interface for the SGLang scheduler, providing a high-performance, standalone alternative to the existing HTTP server. The change enables more flexible deployment and scaling options by decoupling the scheduler's communication layer. It streamlines the request handling process by focusing on tokenized inputs and streaming outputs, while also simplifying the overall architecture by removing certain features deemed unnecessary for this new interface.
Highlights
- Standalone gRPC Server Implementation: A new, fully independent gRPC server has been implemented for SGLang's scheduler, operating as a separate process with its own event loop and state management. This allows for independent scaling from the existing HTTP server.
- Core Components: Key new components include
GrpcRequestManagerfor request orchestration (adapted fromTokenizerManager),SglangSchedulerServicerfor gRPC service implementation, andGrpcReqStatefor tracking request states and metrics. - Streaming-First Communication: All generation requests internally utilize a streaming approach, even for single responses, enhancing real-time interaction.
- ZMQ Integration: Direct communication with the scheduler is established using existing ZMQ PUSH/PULL socket patterns, ensuring efficient inter-process communication.
- Protocol Buffer Updates: The protobuf definitions have been updated to enforce tokenized-only input for requests, simplify health checks, and ensure cross-repository synchronization. Generated protobuf files are now excluded from pre-commit checks.
- Feature Removals: Session management and
FlushCacheoperations have been removed from the gRPC interface for simplicity, with alternatives or future enhancements noted.
Using Gemini Code Assist
The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.
Invoking Gemini
You can request assistance from Gemini at any point in your pull request via creating an issue comment (i.e. comment on the pull request page) using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands.
| Feature | Command | Description |
|---|---|---|
| Code Review | /gemini review |
Performs a code review for the current pull request in its current state. |
| Pull Request Summary | /gemini summary |
Provides a summary of the current pull request in its current state. |
| Comment | @gemini-code-assist | Responds in comments when explicitly tagged, both in issue comments and review comments. |
| Help | /gemini help |
Displays a list of available commands. |
Customization
To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.
Limitations & Feedback
Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.
You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.
Footnotes
-
Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩
There was a problem hiding this comment.
Code Review
This pull request introduces a standalone gRPC server for the SGLang scheduler, which is a significant and well-structured addition. The separation of concerns is clear, with GrpcRequestManager handling orchestration and SglangSchedulerServicer implementing the gRPC service. The code is generally clean and follows existing patterns.
I've identified a few issues that should be addressed:
- A high-severity bug in how default sampling parameters are handled, which could lead to incorrect behavior when clients explicitly set parameters to zero.
- A medium-severity issue where several metrics in the gRPC response are hardcoded to zero, providing incomplete information to the client.
- A couple of minor maintainability improvements in the pre-commit configuration and a copy-paste error in a log message.
Overall, this is a great contribution. Addressing these points will improve the robustness and correctness of the new gRPC server.
Removed Fields from GenerateComplete: - ❌ prompt_tokens - ❌ completion_tokens - ❌ cached_tokens - ❌ total_generation_time - ❌ time_to_first_token - ❌ tokens_per_second - ❌ spec_verify_count
|
tokenization is inside router, but router is stateful. how can i have a multi-process tokenizer to handle with largh RPS |
|
tokenizer inside of the router today is multi threaded by nature |
yes, i forget it is rust. cool. |
|
QQ @CatherineSue @slin1237 What's the progress about this pr? Thanks! |
Motivation
This PR implements a fully standalone gRPC server that provides an alternative interface to SGLang's scheduler, separate from the existing HTTP server. The gRPC server mimics the orchestration behaviors of the existing Tokenizer Manager while operating as an independent process with its own event loop and state management.
Modifications
Key Components
GrpcRequestManager: Core request orchestration (adapted from TokenizerManager)SglangSchedulerServicer: gRPC service implementationGrpcReqState: Request state tracking with metricsStandalone gRPC Server
Protocol Buffer Updates
sglangandsgl-router🔄 gRPC Server vs TokenizerManager
🚫 REMOVED Features (By Design)
Accuracy Tests
gRPC Server Start
Health Check
Benchmarking and Profiling
Checklist