Skip to content

feat support max_total_num_tokens in getserverinforesponse to keep align with get_server_info in HTTP server - #817

Merged
slin1237 merged 2 commits into
smg-project:mainfrom
Huixxi:feature_add_max_total_num_tokens_into_getserverinforesponse
Mar 21, 2026
Merged

slin1237 merged 2 commits into
smg-project:mainfrom
Huixxi:feature_add_max_total_num_tokens_into_getserverinforesponse

Conversation

@Huixxi

@Huixxi Huixxi commented Mar 19, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

Add the max_total_num_tokens keyword in gRCP mode's GetServerInfoResponse to keep align with the origin get_server_info in HTTP server's response.

Solution

Changes

Test Plan

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • New Features
    • Server info now includes a reported maximum total token capacity, so clients and dashboards can display and reason about system capacity limits.
    • Improves visibility into server capacity for monitoring and decision-making, with no changes to service endpoints or existing client workflows.

Copilot AI review requested due to automatic review settings March 19, 2026 07:34
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the gRPC server's information reporting by introducing a new field to its server info response. This addition ensures consistency with the HTTP server's capabilities, providing a more comprehensive view of the server's token capacity across different interfaces. The update facilitates better resource management and client-side decision-making by making crucial configuration details accessible via gRPC.

Highlights

  • gRPC Protocol Update: The GetServerInfoResponse message in the sglang_scheduler.proto file has been updated to include a new field, max_total_num_tokens.
  • Feature Alignment: This change aims to align the gRPC server information with the existing HTTP server's get_server_info functionality by exposing the maximum total number of tokens.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 19, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Added an int32 max_total_num_tokens = 10 field to GetServerInfoResponse in the scheduler protobuf and populated that field in the servicer implementation from self.scheduler_info (defaults to 0).

Changes

Cohort / File(s) Summary
Protobuf Schema Update
crates/grpc_client/proto/sglang_scheduler.proto
Appended int32 max_total_num_tokens = 10; to message GetServerInfoResponse.
Server Implementation
grpc_servicer/smg_grpc_servicer/sglang/servicer.py
Set max_total_num_tokens in SGLangSchedulerServicer.GetServerInfo using self.scheduler_info.get("max_total_num_tokens", 0) when building the GetServerInfoResponse.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Poem

🐰 I dug a line, a tiny token nest,
A new small number now takes its rest.
From proto to servicer, snug and neat,
Max tokens reported, a soft little beat.
Cheers and carrots for this tidy feat! 🥕

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title directly describes the main change: adding max_total_num_tokens support to GetServerInfoResponse to align with HTTP server behavior.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@mergify

mergify Bot commented Mar 19, 2026

Copy link
Copy Markdown
Contributor

Hi @Huixxi, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

…ign with get_server_info in HTTP server

Signed-off-by: 桓希 <huxiguo.hxg@taobao.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request introduces the max_total_num_tokens field to the GetServerInfoResponse protobuf message, which is a positive step towards aligning the gRPC server information with the HTTP server's capabilities. However, for this feature to be fully functional and provide accurate data, the corresponding Python servicer's GetServerInfo method needs to be updated to populate this newly added field. Without this update, gRPC clients will receive a default value (0), potentially leading to misleading server information.

// Server metadata
string server_type = 8; // "grpc"
google.protobuf.Timestamp start_time = 9;
int32 max_total_num_tokens = 10;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The addition of max_total_num_tokens to GetServerInfoResponse is a good step towards aligning gRPC server information with the HTTP server. However, please ensure that the corresponding Python servicer's GetServerInfo method (in grpc_servicer/smg_grpc_servicer/sglang/servicer.py) is updated to populate this new field with the correct value. Currently, gRPC clients will receive the default value (0), which might lead to incorrect information being reported.

@Huixxi
Huixxi force-pushed the feature_add_max_total_num_tokens_into_getserverinforesponse branch from 81708c0 to b3b73d1 Compare March 19, 2026 07:36

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the SGLang scheduler gRPC protobuf to include max_total_num_tokens in GetServerInfoResponse, aligning the gRPC server-info payload with the token-capacity information available elsewhere (e.g., loads snapshots / HTTP server info).

Changes:

  • Add max_total_num_tokens field to GetServerInfoResponse in sglang_scheduler.proto.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

You can also share your feedback on Copilot code review. Take the survey.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 81708c0573

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

// Server metadata
string server_type = 8; // "grpc"
google.protobuf.Timestamp start_time = 9;
int32 max_total_num_tokens = 10;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Populate max_total_num_tokens before advertising it

The schema now exposes max_total_num_tokens, but the only SGLang gRPC implementation I inspected still never sets it: grpc_servicer/smg_grpc_servicer/sglang/servicer.py:491-500 builds GetServerInfoResponse without this field. Because proto3 defaults unset int32 fields to 0, every GetServerInfo reply from the gRPC server will report a bogus token limit until the servicer is updated, which breaks any client that starts using this new field for capacity/context-length decisions.

Useful? React with 👍 / 👎.

// Server metadata
string server_type = 8; // "grpc"
google.protobuf.Timestamp start_time = 9;
int32 max_total_num_tokens = 10;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Regenerate the checked-in Go bindings for this proto change

This API change is not propagated to the repository's Go SDK. bindings/golang/internal/proto/sglang_scheduler.pb.go:2820-2931 still defines GetServerInfoResponse with fields only through start_time, so Go clients built from this commit will silently drop field 10 on decode and have no accessor for it. If this proto stays source-controlled, the generated Go files need to be refreshed in the same change.

Useful? React with 👍 / 👎.

…ign with get_server_info in HTTP server 2

Signed-off-by: 桓希 <huxiguo.hxg@taobao.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d6b65b6bed

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

sglang_version=sglang.__version__,
server_type="grpc",
start_time=start_timestamp,
max_total_num_tokens=self.scheduler_info.get("max_total_num_tokens", 0),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reuse the existing context-length fallback here

If the scheduler bootstrap dict does not contain max_total_num_tokens—a case the same startup path already anticipates in grpc_servicer/smg_grpc_servicer/sglang/server.py:81-85 by falling back to server_args.context_length or 8192—GetServerInfo will now advertise 0 for the new field. Any client that starts using this field for capacity/context-length decisions will treat that worker as having no usable context, even though model_info.max_context_length from the same process remains non-zero.

Useful? React with 👍 / 👎.

@Huixxi

Huixxi commented Mar 20, 2026

Copy link
Copy Markdown
Contributor Author

@slin1237 hi~ whenever you have some time, could you please take a look at this PR? Thanks!

@slin1237
slin1237 merged commit 6ed9e1b into smg-project:main Mar 21, 2026
30 of 33 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants