Skip to content

Implement Standalone gRPC Server for SGLang Python Scheduler - #10283

Merged
zhyncs merged 30 commits into
mainfrom
chang/grpc
Sep 12, 2025
Merged

zhyncs merged 30 commits into
mainfrom
chang/grpc

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Sep 10, 2025 •

Copy link
Copy Markdown
Collaborator

Motivation

This PR implements a fully standalone gRPC server that provides an alternative interface to SGLang's scheduler, separate from the existing HTTP server. The gRPC server mimics the orchestration behaviors of the existing Tokenizer Manager while operating as an independent process with its own event loop and state management.

Modifications

Key Components

  • GrpcRequestManager: Core request orchestration (adapted from TokenizerManager)
  • SglangSchedulerServicer: gRPC service implementation
  • GrpcReqState: Request state tracking with metrics
  • Protocol buffer definitions for type-safe communication

Standalone gRPC Server

  • Independent Process: gRPC server runs separately from HTTP server, enabling independent scaling
  • Streaming-First: All generation requests use streaming internally, even for single responses
  • ZMQ Integration: Direct scheduler communication using existing PUSH/PULL socket patterns
  • Event Loop: Async request handling with proper state tracking (adapted from TokenizerManager)

Protocol Buffer Updates

  • Tokenized-Only Input: Removed oneof input with text fields, enforcing TokenizedInput across all requests
  • Simplified Health Checks: Streamlined health check protocol without session management

NOTE: Currently, the health check relies on grpc client to set health check inputs

  • Cross-Repository Sync: Updated protobuf definitions in both sglang and sgl-router
  • Generated File Management: Added .pyi, _pb2.py, and _pb2_grpc.py files to pre-commit exclusions
  • Removed FlushCache and Initialize endpoint

🔄 gRPC Server vs TokenizerManager

Feature TokenizerManager gRPC Server
Request Process Part of HTTP server ✅ Standalone process
Communication HTTP → ZMQ ✅ gRPC → ZMQ
Tokenization Built-in ✅ sgl-router client-side (by design)
Response HTTP request/response ✅ Always streaming
State Tracking ReqState ✅ GrpcReqState
Scheduler Comm ZMQ PUSH/PULL ✅ ZMQ PUSH/PULL (identical)

🚫 REMOVED Features (By Design)

Feature Reason Alternative
Session Management Removed for simplicity Can be re-added if needed
FlushCache Operations Not needed for core functionality Removed from proto and implementations
Model Updates Handled at scheduler level Future enhancement
LoRA Management Handled at scheduler level Future enhancement

Accuracy Tests

gRPC Server Start

Screenshot 2025-09-10 at 2 42 45 PM

Health Check

Screenshot 2025-09-10 at 2 27 22 PM

Benchmarking and Profiling

Checklist

- It is uncommon to establish a handshake connection
- Remove Initialize method and run compile proto to update auto-generated code
- Remove flush cache
- Remove open_session and close_session
- Remove all detailed health metrics
- Keep HealthCheck same as /health_generate in http_server
- Clean up proto and re-compile
- Proper dependency injection: implement a GrpcSignalHandler instead of using SignalHanlder in tokenizer_manager.py
- Have GrpcSignalHandler to simply gracefully exit with scheduler failure instead of dumping crash info
- When tokenizer is init (skip_tokenizer_init is False), scheduler sends results to detoeknizer_ipc_name

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary of Changes

Hello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a new gRPC interface for the SGLang scheduler, providing a high-performance, standalone alternative to the existing HTTP server. The change enables more flexible deployment and scaling options by decoupling the scheduler's communication layer. It streamlines the request handling process by focusing on tokenized inputs and streaming outputs, while also simplifying the overall architecture by removing certain features deemed unnecessary for this new interface.

Highlights

  • Standalone gRPC Server Implementation: A new, fully independent gRPC server has been implemented for SGLang's scheduler, operating as a separate process with its own event loop and state management. This allows for independent scaling from the existing HTTP server.
  • Core Components: Key new components include GrpcRequestManager for request orchestration (adapted from TokenizerManager), SglangSchedulerServicer for gRPC service implementation, and GrpcReqState for tracking request states and metrics.
  • Streaming-First Communication: All generation requests internally utilize a streaming approach, even for single responses, enhancing real-time interaction.
  • ZMQ Integration: Direct communication with the scheduler is established using existing ZMQ PUSH/PULL socket patterns, ensuring efficient inter-process communication.
  • Protocol Buffer Updates: The protobuf definitions have been updated to enforce tokenized-only input for requests, simplify health checks, and ensure cross-repository synchronization. Generated protobuf files are now excluded from pre-commit checks.
  • Feature Removals: Session management and FlushCache operations have been removed from the gRPC interface for simplicity, with alternatives or future enhancements noted.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point in your pull request via creating an issue comment (i.e. comment on the pull request page) using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in issue comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a standalone gRPC server for the SGLang scheduler, which is a significant and well-structured addition. The separation of concerns is clear, with GrpcRequestManager handling orchestration and SglangSchedulerServicer implementing the gRPC service. The code is generally clean and follows existing patterns.

I've identified a few issues that should be addressed:

  1. A high-severity bug in how default sampling parameters are handled, which could lead to incorrect behavior when clients explicitly set parameters to zero.
  2. A medium-severity issue where several metrics in the gRPC response are hardcoded to zero, providing incomplete information to the client.
  3. A couple of minor maintainability improvements in the pre-commit configuration and a copy-paste error in a log message.

Overall, this is a great contribution. Addressing these points will improve the robustness and correctness of the new gRPC server.

Comment thread python/sglang/srt/entrypoints/grpc_server.py
Comment thread .pre-commit-config.yaml
Comment thread python/sglang/srt/entrypoints/grpc_request_manager.py Outdated
Comment thread python/sglang/srt/entrypoints/grpc_server.py
Comment thread python/sglang/srt/entrypoints/grpc_server.py Outdated
  Removed Fields from GenerateComplete:

  - ❌ prompt_tokens
  - ❌ completion_tokens
  - ❌ cached_tokens
  - ❌ total_generation_time
  - ❌ time_to_first_token
  - ❌ tokens_per_second
  - ❌ spec_verify_count
@jimmy-evo

Copy link
Copy Markdown
Contributor

tokenization is inside router, but router is stateful. how can i have a multi-process tokenizer to handle with largh RPS

@slin1237

Copy link
Copy Markdown
Collaborator

tokenizer inside of the router today is multi threaded by nature
what do u mean by router is stateful
if we are talking about radix tree
the synchronization of those aren't dealt with yet, it can be solved by a mesh of connections between routers or introduce a storage among them for multiple replicas
the worker management can be done in the same way or use service discovery

@jimmy-evo

Copy link
Copy Markdown
Contributor

tokenizer inside of the router today is multi threaded by nature

yes, i forget it is rust.

cool.

@zhyncs

zhyncs commented Sep 11, 2025

Copy link
Copy Markdown
Contributor

QQ @CatherineSue @slin1237 What's the progress about this pr? Thanks!

@zhyncs
zhyncs merged commit 53ca155 into main Sep 12, 2025
10 of 61 checks passed
@zhyncs
zhyncs deleted the chang/grpc branch September 12, 2025 03:57
@slin1237 slin1237 mentioned this pull request Sep 11, 2025
36 of 48 tasks
0826joyce pushed a commit to 0826joyce/sglang-perf-opt that referenced this pull request May 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants