Skip to content

[gRPC] Expose multimodal flag + tokenizer metadata via get_init_info - #25170

Closed
laudney wants to merge 2 commits into
sgl-project:mainfrom
laudney:feat/grpc-scheduler-init-info
Closed

laudney wants to merge 2 commits into
sgl-project:mainfrom
laudney:feat/grpc-scheduler-init-info

Conversation

@laudney

@laudney laudney commented May 13, 2026

Copy link
Copy Markdown

Summary

  • Scheduler.get_init_info() now emits is_generation, supports_vision, vocab_size, eos_token_ids, pad_token_id, and bos_token_id alongside the existing three keys, sourced from self.model_config and self.is_generation.
  • DataParallelController.run_data_parallel_controller_process now spreads the first child scheduler's full init dict into the upstream pipe send instead of rebuilding it from scratch, so DP-mode workers don't strip the new fields.

Why

GetModelInfo over --grpc-mode was reporting supports_vision: false for every model — including genuinely multimodal models like nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8 (architecture NemotronH_Nano_Omni_Reasoning_V3, which is registered in multimodal_model_archs at python/sglang/srt/configs/model_config.py:1518).

Root cause is a long-standing dialect mismatch between SGLang's scheduler and the smg-grpc-servicer it talks to:

  • smg_grpc_servicer/sglang/server.py:81-97 builds the model_info dict that GetModelInfo returns by reading scheduler_info.get("supports_vision", False), .get("vocab_size", 128256), .get("eos_token_ids", []), .get("pad_token_id", 0), .get("bos_token_id", 1).
  • smg_grpc_servicer/sglang/servicer.py:325-329 also reads scheduler_info.get("is_generation").
  • But Scheduler.get_init_info() only ever returned {status, max_total_num_tokens, max_req_input_len}. Every other key was absent, so all six reads silently fell back to their hardcoded defaults.

The bug has been present since #10283 (the original gRPC server PR) and was carried forward through #20478 (the standalone-package extraction). It went unnoticed because most --grpc-mode deployments to date have been text-only LLMs where supports_vision: false happened to be the correct answer.

The vLLM servicer counterpart at smg_grpc_servicer/vllm/servicer.py:268-275 reads these same fields directly from model_config.is_multimodal_model, etc.; the SGLang bridge was written expecting them via scheduler_info, but SGLang's scheduler never put them there.

What the patch does

Scheduler.get_init_info() (scheduler.py)

Adds six new keys to the dict it sends up the IPC pipe, sourced from objects that already exist on the scheduler at the time the pipe send runs:

Key Source Notes
is_generation self.is_generation set in init_tokenizer() before scheduler is exposed
supports_vision self.model_config.is_multimodal mirrors vLLM's is_multimodal_model
vocab_size self.model_config.vocab_size
eos_token_ids sorted(self.model_config.hf_eos_token_id or []) sorted() for deterministic wire order; set→list conversion needed for proto repeated int32
pad_token_id getattr(hf_config, "pad_token_id", None) (coerced) None-coerced to 0 (proto int32 can't carry None)
bos_token_id getattr(hf_config, "bos_token_id", None) (coerced) None-coerced to 1

The None-coercion uses x if x is not None else default rather than x or default, which correctly passes through a legitimate pad_token_id = 0 (common in many models).

DataParallelController (data_parallel_controller.py)

When dp_size > 1, the SGLang gRPC launcher spawns run_data_parallel_controller_process instead of per-rank schedulers. The controller previously rebuilt the upstream handshake dict from scratch with just {status, max_total_num_tokens, max_req_input_len, SCHEDULER_PIDS_ARG}, which would silently strip every field this PR adds to get_init_info(). The fix:

  1. Cache the first child scheduler's full init dict on the controller as self.scheduler_init_info.
  2. Replace the hand-built dict in run_data_parallel_controller_process with {**controller.scheduler_init_info, SCHEDULER_PIDS_ARG: scheduler_pids}.

The existing self.max_total_num_tokens / self.max_req_input_len attributes are kept in place because python/sglang/srt/ray/engine.py:285-286 reads them.

Note: the Ray DP path at ray/engine.py:283-288 has the same bug shape (it also constructs a stripped-down dict by hand) but isn't on the --grpc-mode critical path. Left for a follow-up.

Test plan

  • Verified live against sglang serve --grpc-mode on a DGX Spark worker serving nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8. GetModelInfo before vs after:

    Field Before (defaults) After (real values)
    supports_vision false true
    vocab_size 128256 131072
    eos_token_ids [] [2, 11]
    is_generation implicit fallback true
    pad_token_id 0 (default) 0 (real, coincidence)
    bos_token_id 1 (default) 1 (real, coincidence)
  • Pre-commit (isort / ruff / black / codespell) clean on the two changed files.

  • Patch applies cleanly against main (5227b0766) and the v0.5.11 release tag.

  • No unit tests added — get_init_info returns a snapshot of objects that themselves have coverage elsewhere; the meaningful end-to-end signal is the gRPC GetModelInfo round-trip, which lives outside this repo in the smg-grpc-servicer package. Happy to add a unit test in test/registered/unit/managers/test_scheduler_init_info.py if reviewers prefer.

  • HTTP mode regression — the tokenizer_manager only reads max_req_input_len from the init dict (http_server.py:222), so the additional keys are no-ops for HTTP consumers.

Related

GetModelInfo over gRPC always reports `supports_vision: false`
(and falls back to default vocab_size / eos_token_ids / pad /
bos / is_generation) for every model. The smg-grpc-servicer
SGLang bridge reads these fields out of `scheduler_info` via
`.get(..., default)` (see smg_grpc_servicer/sglang/server.py
and servicer.py), but `Scheduler.get_init_info()` only ever
returned `{status, max_total_num_tokens, max_req_input_len}`,
so every other key silently falls back to its hardcoded
default. For multimodal models like Nemotron-3-Nano-Omni this
makes the worker invisible to image/audio routing through the
gateway. This patch populates `is_generation`, `supports_vision`,
`vocab_size`, `eos_token_ids`, `pad_token_id`, and `bos_token_id`
from `self.model_config` and `self.is_generation`, mirroring
what the vLLM servicer reads from `model_config.is_multimodal_model`
at smg_grpc_servicer/vllm/servicer.py:273. A second hunk fixes
`DataParallelController`, which rebuilds the upstream handshake
dict from scratch at line 651 and would otherwise strip the new
fields when `dp_size > 1`; the controller now caches the first
child scheduler's full init dict and spreads it through the
pipe send. Verified on a live `--grpc-mode` Spark worker serving
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8: GetModelInfo
now correctly returns supports_vision=True, vocab_size=131072,
and eos_token_ids=[2, 11] where they previously came back as
False / 128256 / []. HTTP-mode consumers are unaffected.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the scheduler initialization information to include essential model metadata such as vision support, vocabulary size, and token IDs, ensuring these fields are correctly propagated through the data parallel controller. The review identified a potential crash risk in the handling of eos_token_ids due to incorrect type handling and falsy value evaluation, which requires a more robust implementation to handle both integer and list types correctly.

Comment on lines +1503 to 1519
hf_cfg = self.model_config.hf_config
pad_id = getattr(hf_cfg, "pad_token_id", None)
bos_id = getattr(hf_cfg, "bos_token_id", None)
result_dict = {
"status": "ready",
"max_total_num_tokens": self.max_total_num_tokens,
"max_req_input_len": self.max_req_input_len,
"is_generation": self.is_generation,
"supports_vision": self.model_config.is_multimodal,
"vocab_size": self.model_config.vocab_size,
"eos_token_ids": sorted(self.model_config.hf_eos_token_id or []),
# int32 proto field; coerce missing IDs to the smg-grpc-servicer
# int defaults (0 for pad, 1 for bos) so encoding doesn't crash
# on `None`.
"pad_token_id": pad_id if pad_id is not None else 0,
"bos_token_id": bos_id if bos_id is not None else 1,
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The logic for eos_token_ids has two issues:

  1. If hf_eos_token_id is a single integer (common in many models), sorted() will raise a TypeError as integers are not iterable.
  2. The use of or [] will treat a token ID of 0 as falsy and return an empty list instead of [0].

It's safer to explicitly handle the integer case and check for None.

References
  1. Defensive programming: ensure appropriate handling of different types (int vs list) and edge cases (0 as a valid ID).

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the careful read. Both concerns are technically defused by an upstream invariant, but you're right that the diff doesn't show that — so I've pushed a defensive rewrite in 83c2f89 that makes the safety locally obvious.

Details on the original code, for the record:

  1. Bare int → TypeError — model_config.hf_eos_token_id is populated by ModelConfig._get_hf_eos_token_id, which normalizes HF's int | list[int] | None into a Set[int] (line 1307 wraps a bare int with {eos_ids}, line 1309 substitutes set() for None). By the time sorted() sees it, it's always a set.

  2. or [] drops 0 — or was operating on the set, not a scalar. bool({0}) is True, so {0} or [] == {0} and sorted({0}) == [0]. The falsy-zero footgun is real for scalar int fields but not for sets.

That said, the producer's type annotation is Optional[Set[int]], which advertises a None return the function never actually produces. The handshake shouldn't depend on a cross-file invariant the type system doesn't enforce, so the new commit coerces inline:

eos_ids = self.model_config.hf_eos_token_id
eos_token_ids = sorted(
    [eos_ids] if isinstance(eos_ids, int) else eos_ids or []
)

isinstance(int) runs first so token ID 0 survives (isinstance(0, int) is True, [0] or [] == [0]). Same behavior today, robust against future refactors.

sorted(hf_eos_token_id or []) silently relied on
ModelConfig._get_hf_eos_token_id normalizing int|list|None into
a Set[int] before this site runs. That contract is invisible
from this hunk — the producer's annotation is Optional[Set[int]]
— so a future refactor could legitimately break the handshake.
Coerce inline (wrap a bare int, treat None as empty) so the
serialization is robust without depending on a cross-file
invariant that the type system doesn't enforce.
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Thanks @laudney. Closing this because it has had no updates in 112 days.

Reopen it if the work is still relevant.

Some directories moved recently, so an older branch may need retargeting:
sgl-kernel/ -> python/sglang/kernels/aot/, python/sglang/jit_kernel/
-> python/sglang/kernels/jit/, docs/ -> docs/docs/ (.mdx),
bench_serving.py -> benchmark/serving.py, test/srt/ -> test/registered/.

@github-actions github-actions Bot closed this Sep 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant