Skip to content

fix(serve): respect user-set CUDA_VISIBLE_DEVICES in gpu_env - #750

Closed
paxiaatucsdedu wants to merge 2 commits into
smg-project:mainfrom
paxiaatucsdedu:pan/GPU-assignment-control
Closed

paxiaatucsdedu wants to merge 2 commits into
smg-project:mainfrom
paxiaatucsdedu:pan/GPU-assignment-control

Conversation

@paxiaatucsdedu

@paxiaatucsdedu paxiaatucsdedu commented Mar 13, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

When running smg serve inside a Docker container with CUDA_VISIBLE_DEVICES set
(e.g. docker run -e CUDA_VISIBLE_DEVICES=4 ...), the gpu_env method in
WorkerLauncher unconditionally overrides the variable with sequential IDs starting
from 0. This causes workers to land on the wrong GPU, leading to OOM errors when
another process is already using GPU 0.

Solution

Modify gpu_env to check whether CUDA_VISIBLE_DEVICES is already set in the
environment. If it is, treat it as the available GPU pool and slice into it by
dp_rank and tp_size. If it is not set (or empty), fall back to the original
sequential assignment (0,1,...). A bounds check raises a clear ValueError when
the pool has fewer GPUs than the requested dp_rank * tp_size range.

Changes

  • bindings/python/src/smg/serve.py: Rewrite WorkerLauncher.gpu_env to index
    into an existing CUDA_VISIBLE_DEVICES pool instead of overriding it. Add bounds
    check with descriptive error message.
  • bindings/python/tests/test_serve.py: Add unit tests for pool indexing
    (parametrized across multiple GPU layouts), cross-backend verification (vllm),
    bounds check (ValueError), and empty-string fallback.

Test Plan

Manual verification (Docker):

# Before fix: worker ignores CUDA_VISIBLE_DEVICES=4 and launches on GPU 0
docker run -itd --gpus all -e CUDA_VISIBLE_DEVICES=4 \
  <smg-image> smg serve --backend sglang --model-path /models/test
# Logs show: torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 150.00 MiB. GPU 0 has a total capacity of 79.44 GiB of which 19.00 MiB is free.
 
# After fix: worker correctly uses GPU 4
docker run -itd --gpus all -e CUDA_VISIBLE_DEVICES=4 \
  -v ./serve.py:/usr/local/lib/python3.12/dist-packages/smg/serve.py:ro \
  <smg-image> smg serve --backend sglang --model-path /models/test
# Logs show: CUDA_VISIBLE_DEVICES=4

Added unit test at bindings/python/tests/test_serve.py

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • New Features

    • GPU environment now respects externally provided CUDA_VISIBLE_DEVICES pool, automatically selecting the appropriate device subset for each worker. Includes validation to ensure sufficient GPUs are available, with graceful fallback to default behavior when not specified.
  • Tests

    • Added comprehensive test coverage for GPU environment handling, including device pool selection, validation of edge cases, and error handling scenarios.

@github-actions github-actions Bot added python-bindings Python bindings changes tests Test changes labels Mar 13, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a critical issue where smg serve workers incorrectly ignored user-defined CUDA_VISIBLE_DEVICES settings, leading to GPU assignment conflicts and out-of-memory errors, particularly in containerized environments. The changes ensure that the system now properly interprets and utilizes the CUDA_VISIBLE_DEVICES environment variable as the primary pool for GPU allocation, enhancing resource management and preventing runtime failures.

Highlights

  • GPU Assignment Logic: Modified the gpu_env method in WorkerLauncher to respect and index into an existing CUDA_VISIBLE_DEVICES environment variable, preventing workers from being assigned to incorrect GPUs.
  • Error Handling: Introduced a bounds check that raises a ValueError if the requested GPU range (based on dp_rank and tp_size) exceeds the number of available GPUs specified in CUDA_VISIBLE_DEVICES.
  • Test Coverage: Added comprehensive unit tests for gpu_env to verify correct indexing into the CUDA_VISIBLE_DEVICES pool, cross-backend compatibility (vllm), bounds checking, and fallback behavior for empty CUDA_VISIBLE_DEVICES.
Changelog
  • bindings/python/src/smg/serve.py
    • Updated the gpu_env method to prioritize an existing CUDA_VISIBLE_DEVICES environment variable for GPU allocation.
    • Implemented logic to slice into the existing CUDA_VISIBLE_DEVICES pool based on dp_rank and tp_size.
    • Added a ValueError check for insufficient available GPUs when CUDA_VISIBLE_DEVICES is set.
    • Ensured fallback to sequential GPU assignment if CUDA_VISIBLE_DEVICES is not set or empty.
  • bindings/python/tests/test_serve.py
    • Added parameterized tests to verify gpu_env correctly indexes into various CUDA_VISIBLE_DEVICES configurations.
    • Included a test case for vllm backend to confirm cross-backend compatibility.
    • Added tests to confirm ValueError is raised when CUDA_VISIBLE_DEVICES specifies fewer GPUs than required.
    • Verified that gpu_env falls back to default sequential assignment when CUDA_VISIBLE_DEVICES is empty.
Activity
  • No specific activity (comments, reviews, progress updates) has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 13, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@paxiaatucsdedu has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 11 minutes and 58 seconds before requesting another review.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 95b4b2ec-8419-4f1b-b495-64679cdac01c

📥 Commits

Reviewing files that changed from the base of the PR and between b598982 and ed6f625.

📒 Files selected for processing (1)
  • bindings/python/src/smg/serve.py
📝 Walkthrough

Walkthrough

The PR enhances GPU environment setup in the Python serve module by adding support for respecting externally provided CUDA_VISIBLE_DEVICES as a GPU pool. When set, it partitions this list according to tensor parallelism size and distributed parallelism rank to assign appropriate GPU subsets to workers, with bounds validation and fallback behavior for legacy setups.

Changes

Cohort / File(s) Summary
GPU environment handling logic
bindings/python/src/smg/serve.py
Enhanced gpu_env function to partition externally provided CUDA_VISIBLE_DEVICES based on tp_size and dp_rank, with validation that raises ValueError on insufficient GPUs. Falls back to contiguous assignment when CUDA_VISIBLE_DEVICES is unset.
GPU environment tests
bindings/python/tests/test_serve.py
Added comprehensive unit tests covering CUDA_VISIBLE_DEVICES slicing for SglangWorkerLauncher and VllmWorkerLauncher, validation of insufficient GPU availability, and fallback behavior when pool is empty.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 Hop hop, the GPUs now align,
With CUDA pools sliced by rank so fine,
Each worker finds its GPU suite,
Validation checks before too late! ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'fix(serve): respect user-set CUDA_VISIBLE_DEVICES in gpu_env' clearly and specifically summarizes the main change: modifying the gpu_env function to respect pre-configured CUDA_VISIBLE_DEVICES instead of overwriting it.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request correctly addresses the issue of respecting an existing CUDA_VISIBLE_DEVICES environment variable when launching workers. The implementation is solid: it properly parses the existing variable, slices the available GPU pool based on data and tensor parallelism ranks, and includes a helpful bounds check with a clear error message. The new unit tests are thorough and effectively validate the changes. I have one minor suggestion to simplify the code slightly.

Comment thread bindings/python/src/smg/serve.py Outdated
When CUDA_VISIBLE_DEVICES is already set (e.g. via Docker -e),
treat it as the available GPU pool instead of overriding it.
Adds bounds check and unit tests."

Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
@paxiaatucsdedu
paxiaatucsdedu force-pushed the pan/GPU-assignment-control branch from 3984d1a to b598982 Compare March 13, 2026 02:41

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@bindings/python/src/smg/serve.py`:
- Around line 67-83: Validate tp_size returned by _get_tp_size in serve.py
before using it in GPU index math: ensure tp_size is a positive integer (tp_size
> 0) and raise a clear ValueError if not; update the logic around tp_size,
base_idx = dp_rank * tp_size, and the available_gpus slicing to assume a valid
tp_size only after this check (refer to _get_tp_size, tp_size, base_idx,
available_gpus, and gpu_ids in the shown block).

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: b9d0d2c5-03de-4f7e-984d-92a46842c871

📥 Commits

Reviewing files that changed from the base of the PR and between 50aa6dc and b598982.

📒 Files selected for processing (2)
  • bindings/python/src/smg/serve.py
  • bindings/python/tests/test_serve.py

Comment thread bindings/python/src/smg/serve.py
Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
@paxiaatucsdedu
paxiaatucsdedu force-pushed the pan/GPU-assignment-control branch from 7cdd2bb to ed6f625 Compare March 13, 2026 02:57
@paxiaatucsdedu
paxiaatucsdedu deleted the pan/GPU-assignment-control branch March 13, 2026 03:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

python-bindings Python bindings changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant