Skip to content

fix(serve): use correct argparse attr for sglang tp_size - #665

Merged
slin1237 merged 2 commits into
mainfrom
chang/fix-arg-parse
Mar 7, 2026
Merged

slin1237 merged 2 commits into
mainfrom
chang/fix-arg-parse

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 7, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

SglangWorkerLauncher._get_tp_size reads args.tp_size, but sglang's ServerArgs.add_cli_args registers --tensor-parallel-size as the primary flag (with --tp-size as alias). Argparse derives the dest from the first flag name, storing the value as args.tensor_parallel_size. This means _get_tp_size always returns the default 1, causing CUDA_VISIBLE_DEVICES to be set to a single GPU — regardless of what --tp-size or --tensor-parallel-size value is passed on the CLI.

This results in CUDA error: invalid device ordinal when running multi-GPU TP with smg serve --backend sglang --tp-size N.

Solution

  • Change SglangWorkerLauncher._get_tp_size to read args.tensor_parallel_size (matching VllmWorkerLauncher which already does this correctly).
  • Rewrite the gpu_env tests to go through parse_serve_args (integration-style) so CLI flag → attribute name mismatches are caught by tests, rather than fabricating argparse.Namespace objects with arbitrary attribute names.

Changes

  • bindings/python/src/smg/serve.py: Fix SglangWorkerLauncher._get_tp_size to read tensor_parallel_size instead of tp_size.
  • bindings/python/tests/test_serve.py: Add integration tests that parse CLI args through parse_serve_args before checking gpu_env output for sglang, vllm, and trtllm. Update existing unit tests to use the correct attribute name.

Test Plan

  • New integration tests test_sglang_tp_from_cli, test_vllm_tp_from_cli, test_trtllm_tp_from_cli verify that --tp-size 4 on the CLI produces CUDA_VISIBLE_DEVICES=0,1,2,3.
  • These tests would have caught the original bug — test_sglang_tp_from_cli fails on the old code because parse_serve_args stores the value as tensor_parallel_size, not tp_size.
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Changes

    • Adjusted which tensor-parallel CLI parameter controls per-GPU allocation for Sglang workers, which may change GPU assignment behavior for setups that used the previous parameter. No public API changes.
  • Tests

    • Added and expanded integration and unit tests to verify CLI parameter propagation and GPU device assignment across multiple launchers.

SglangWorkerLauncher._get_tp_size was reading `args.tp_size` but
sglang's ServerArgs.add_cli_args registers --tensor-parallel-size as
the primary flag (with --tp-size as alias), so argparse stores the
value as `args.tensor_parallel_size`. This caused _get_tp_size to
always return the default of 1, setting CUDA_VISIBLE_DEVICES to a
single GPU regardless of the --tp-size value passed on the CLI.

Also rewrites the gpu_env tests to go through parse_serve_args so
that CLI flag → attribute name mismatches are caught by tests.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@github-actions github-actions Bot added python-bindings Python bindings changes tests Test changes labels Mar 7, 2026
@coderabbitai

coderabbitai Bot commented Mar 7, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 35719a09-7a43-414f-8a2d-b3e2cf7c3257

📥 Commits

Reviewing files that changed from the base of the PR and between b596264 and cbf859f.

📒 Files selected for processing (1)
  • bindings/python/tests/test_serve.py

📝 Walkthrough

Walkthrough

switched _get_tp_size to read args.tensor_parallel_size instead of args.tp_size; updated and added tests to validate CLI flag propagation to GPU assignment across Sglang, VLLM, and TRT-LLM launchers.

Changes

Cohort / File(s) Summary
Tensor-parallel attribute change
bindings/python/src/smg/serve.py
_get_tp_size now returns args.tensor_parallel_size instead of args.tp_size, changing which CLI/config field controls per-worker tensor-parallel sizing for Sglang workers.
Tests — CLI, launcher integration, TRT-LLM precedence
bindings/python/tests/test_serve.py
Expanded integration tests to cover --tp-size and --tensor-parallel-size through parse_serve_args into CUDA_VISIBLE_DEVICES for Sglang, VLLM, TRT-LLM; added unit tests for TRT-LLM _get_tp_size precedence; updated tests to use tensor_parallel_size where applicable.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • key4ng
  • slin1237

Poem

🐰 A hop of code, a tiny tweak in view,
From tp_size to tensor_parallel_size anew,
GPUs shuffle into a fresher dance,
Tests listen closely, giving each step a chance,
I nibble a carrot — CI passes, hooray!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 35.71% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically describes the main fix: correcting which argparse attribute (tensor_parallel_size instead of tp_size) is read in the sglang tensor-parallel size logic.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/fix-arg-parse

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request resolves a critical bug affecting multi-GPU tensor parallelism with the sglang backend in the smg serve command. Previously, the system incorrectly parsed the tensor parallel size, defaulting to a single GPU and causing runtime errors. The fix ensures that the tensor_parallel_size is correctly identified and applied, enabling proper multi-GPU utilization. Additionally, the changes significantly improve test coverage by adding integration tests that validate the end-to-end parsing of CLI arguments, safeguarding against future attribute mismatch issues.

Highlights

  • Bug Fix: Corrected an issue where SglangWorkerLauncher failed to properly read the tensor_parallel_size argument, leading to incorrect GPU allocation and CUDA error: invalid device ordinal during multi-GPU operations.
  • Attribute Alignment: Modified SglangWorkerLauncher._get_tp_size to read args.tensor_parallel_size, aligning it with how argparse stores the value from --tensor-parallel-size or --tp-size.
  • Enhanced Testing: Introduced new integration tests that parse CLI arguments via parse_serve_args to ensure correct mapping of CLI flags to internal attribute names for sglang, vllm, and trtllm, preventing similar regressions.
  • Test Refinement: Updated existing unit tests in test_serve.py to use the correct tensor_parallel_size attribute when fabricating argparse.Namespace objects.
Changelog
  • bindings/python/src/smg/serve.py
    • Changed SglangWorkerLauncher._get_tp_size to retrieve tensor_parallel_size instead of tp_size from args.
  • bindings/python/tests/test_serve.py
    • Added test_sglang_tp_from_cli, test_vllm_tp_from_cli, and test_trtllm_tp_from_cli to verify CLI flag parsing for tensor parallelism.
    • Updated test_gpu_env_dp_rank_0_tp_2, test_gpu_env_dp_rank_1_tp_2, test_gpu_env_preserves_existing_env, test_gpu_env_does_not_mutate_input, test_gpu_env_none_copies_os_environ, and test_launch_workers_passes_gpu_env to use tensor_parallel_size in argparse.Namespace objects.
    • Removed SglangWorkerLauncher and VllmWorkerLauncher specific tests from test_get_tp_size and introduced test_get_tp_size_trtllm for TRT-LLM specific behavior.
Activity
  • CatherineSue opened this pull request to fix a bug related to sglang tensor parallelism.
  • The pull request description provides a detailed problem statement, solution, changes, and test plan.
  • The author has confirmed that cargo +nightly fmt and cargo clippy checks pass.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request correctly fixes a bug in SglangWorkerLauncher where the tensor parallel size was always defaulting to 1 due to reading the wrong attribute from the argparse namespace. The fix aligns sglang's behavior with vllm by using tensor_parallel_size. Furthermore, the pull request significantly improves the test suite by replacing brittle unit tests that used fabricated Namespace objects with more robust integration-style tests that parse command-line arguments. This ensures that mismatches between CLI flags and their corresponding attribute names are caught by tests, preventing similar issues in the future. The changes are correct and the test improvements are excellent.

The sglang/vllm integration tests called parse_serve_args without
mocking _import_backend_args, causing failures in environments where
sglang/vllm are not installed (including CI unit test job).

Mock the backend arg adders to simulate the --tensor-parallel-size
and --tp-size flags that each backend would register.

Signed-off-by: Chang Su <chang.s.su@oracle.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b596264d0e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread bindings/python/tests/test_serve.py Outdated
@slin1237
slin1237 merged commit cb25b2f into main Mar 7, 2026
58 of 63 checks passed
@slin1237
slin1237 deleted the chang/fix-arg-parse branch March 7, 2026 03:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

python-bindings Python bindings changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants