Skip to content

feat(workflow): surface max_running_requests on SGLang HTTP /server_info - #1529

Merged
slin1237 merged 1 commit into
mainfrom
feat/sglang-server-info-max-running-requests
May 24, 2026
Merged

slin1237 merged 1 commit into
mainfrom
feat/sglang-server-info-max-running-requests

Conversation

@slin1237

@slin1237 slin1237 commented May 24, 2026 •

Copy link
Copy Markdown
Member

Summary

  • SMG's curated ServerInfo struct in discover_metadata.rs deserializes a subset of SGLang's /server_info response (which has ~800 fields). max_running_requests — the per-instance batch-capacity cap from --max-running-requests — was being silently dropped.
  • Adds pub max_running_requests: Option<usize> to the struct. flat_labels is serde-based and picks it up automatically — no other changes needed.
  • The SGLang gRPC label pipeline (SGLANG_GRPC_KEYS in routers/grpc/client.rs) already extracts this field; this closes the HTTP-only path so capacity-aware consumers see the same label regardless of transport.

Test plan

  • cargo test --package smg --lib workflow::steps::local::discover_metadata::tests::test_sglang_server_info — 2 new unit tests pass (present-and-absent cases).
  • Existing #[ignore] integration tests against a live SGLang server still pass.
  • No regression in the rest of the workflow test suite.

Notes

  • Independent of feat(worker): add max_running_requests accessor on Worker trait #1526. Can merge in any order.
  • vLLM's HTTP server doesn't expose an equivalent endpoint and vLLM's vllm_engine.proto GetServerInfoResponse has no batch-capacity field, so this fix is SGLang-HTTP-specific. TRT-LLM uses max_batch_size (different name); that's a separate, larger change.

Summary by CodeRabbit

  • New Features

    • System now surfaces per-instance concurrency limits so users can see how many requests an instance will run concurrently.
  • Tests

    • Added unit tests to verify correct detection and handling when the concurrency-limit information is present or omitted.

Review Change Stack

@slin1237
slin1237 requested a review from CatherineSue as a code owner May 24, 2026 06:14
@coderabbitai

coderabbitai Bot commented May 24, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: cf3bf8d2-4cc3-4176-bb11-de2715bf6dc6

📥 Commits

Reviewing files that changed from the base of the PR and between 02421cb and bf42012.

📒 Files selected for processing (1)
  • model_gateway/src/workflow/steps/local/discover_metadata.rs

📝 Walkthrough

Walkthrough

ServerInfo gains an optional max_running_requests field to expose SGLang's per-instance request concurrency cap. Unit tests confirm deserialization when present and omission from flattened labels when absent.

Changes

Server Metadata Enhancement

Layer / File(s) Summary
ServerInfo concurrency field
model_gateway/src/workflow/steps/local/discover_metadata.rs
ServerInfo struct adds optional max_running_requests: Option<usize> field to capture SGLang's per-instance concurrency cap.
Field validation tests
model_gateway/src/workflow/steps/local/discover_metadata.rs
Unit tests validate that the field deserializes when present in /server_info and is absent from flattened labels when not provided.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • lightseekorg/smg#817: Both PRs extend GetServerInfo/server metadata with new “max-*” concurrency/limit fields (max_running_requests vs max_total_num_tokens) sourced from server info payloads.

Suggested reviewers

  • CatherineSue

Poem

🐰 A field hops in, small and bright,
Counting requests through day and night,
Tests nibble gently, numbers clear,
Concurrency whispers, now we hear. ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat(workflow): surface max_running_requests on SGLang HTTP /server_info' directly and accurately summarizes the main change: adding max_running_requests field to ServerInfo struct to surface SGLang's capacity information.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/sglang-server-info-max-running-requests

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds the max_running_requests field to the ServerInfo struct in discover_metadata.rs to surface the per-instance concurrency cap for SGLang. It also includes unit tests to verify the deserialization and label extraction of this new field, ensuring it handles both present and missing values correctly. I have no feedback to provide.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, well-scoped change. The new max_running_requests field on ServerInfo is correctly typed as Option<usize>, matches the SGLang JSON field name (no rename needed), and is auto-surfaced by the existing generic flat_labels pipeline. Tests cover both the present and absent cases. No issues found.

@github-actions github-actions Bot added the model-gateway Model gateway crate changes label May 24, 2026
@slin1237
slin1237 force-pushed the feat/sglang-server-info-max-running-requests branch from ab192ae to 02421cb Compare May 24, 2026 06:27

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, minimal change. The new max_running_requests: Option<usize> field correctly closes the HTTP-only gap (gRPC path already had it at SGLANG_GRPC_KEYS). flat_labels picks it up automatically via serde — no manual wiring needed. Tests cover both present and absent cases. No issues found.

SMG's curated ServerInfo struct deserializes a subset of SGLang's
/server_info response (~800 fields). max_running_requests — the CLI
flag --max-running-requests, which is the natural per-instance batch
capacity cap — was being dropped silently.

The SGLang gRPC label pipeline already extracts this field (see
SGLANG_GRPC_KEYS in routers/grpc/client.rs). Adding it here closes
the HTTP-only path so capacity-aware consumers see the same label
regardless of transport.

Two unit tests verify both the present-and-absent cases; flat_labels
picks the field up automatically via its serde-based serialization.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237
slin1237 force-pushed the feat/sglang-server-info-max-running-requests branch from 02421cb to bf42012 Compare May 24, 2026 18:21
@slin1237
slin1237 merged commit 781bbdd into main May 24, 2026
29 of 39 checks passed
@slin1237
slin1237 deleted the feat/sglang-server-info-max-running-requests branch May 24, 2026 18:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant