Skip to content

fix(docker): install smg-grpc-servicer from source in engine images - #1086

Merged
slin1237 merged 2 commits into
smg-project:mainfrom
key4ng:fix/engine-image-grpc-servicer-source-install
Apr 10, 2026
Merged

slin1237 merged 2 commits into
smg-project:mainfrom
key4ng:fix/engine-image-grpc-servicer-source-install

Conversation

@key4ng

@key4ng key4ng commented Apr 10, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

Engine images (ghcr.io/lightseekorg/smg:*-sglang-v0.5.10, etc.) shipped a stale smg_grpc_servicer inherited from whatever the base image (lmsysorg/sglang:v0.5.10) happened to preinstall from PyPI at its own build time — not the version in the SMG source tree being baked into the image.

When sglang v0.5.10 refactored srt/utils.py into the srt/utils/ subpackage and moved get_zmq_socket into srt/utils/network.py without re-exporting it at package level, the bundled stale servicer broke at import time:

File "/usr/local/lib/python3.12/dist-packages/smg_grpc_servicer/sglang/request_manager.py", line 38
    from sglang.srt.utils import get_or_create_event_loop, get_zmq_socket, kill_process_tree
ImportError: cannot import name 'get_zmq_socket' from 'sglang.srt.utils'
    (/sgl-workspace/sglang/python/sglang/srt/utils/__init__.py)

This crash-looped engine pods in production despite the v1.4.1 source tree already containing the correct import (commit 5321bce, released as smg-grpc-servicer 0.5.2 in #1078).

Why CI didn't catch it: scripts/ci_install_sglang.sh installs the servicer from source (uv pip install -e grpc_servicer/), so CI ran against the fixed code. The published Docker image ran against a stale PyPI snapshot baked into the base image. CI and the image were testing two different versions of the same package.

Solution

Eliminate the PyPI path entirely for engine images. Always install both smg-grpc-proto and smg-grpc-servicer from the cloned source tree, using --force-reinstall to override anything the base image preinstalled. This makes the image a single source of truth that tracks the repo.

Changes

  • scripts/installation/install-smg.sh: after installing the smg wheel from bindings/python, also pip install --no-cache-dir --force-reinstall crates/grpc_client/python and grpc_servicer from the cloned source. No extras activated — engines are already in the base image and we don't want to reinstall vllm/sglang.
  • docker/engine.Dockerfile: drop the now-redundant ENGINE=vllm-only pip install smg-grpc-servicer[vllm] block (covered uniformly by the source install).

Test Plan

Reproduction before this PR (from a crashing pod running ghcr.io/lightseekorg/smg:1.4.1-sglang-v0.5.10):

$ kubectl run smg-debug --image=ghcr.io/lightseekorg/smg:1.4.1-sglang-v0.5.10 \
    --restart=Never --command -- python3 -c \
    "from smg_grpc_servicer.sglang.servicer import SGLangSchedulerServicer"
...
ImportError: cannot import name 'get_zmq_socket' from 'sglang.srt.utils'

After this PR, install-smg.sh reinstalls smg-grpc-servicer from ${SMG_SRC}/grpc_servicer (which contains the fixed from sglang.srt.utils.network import get_zmq_socket) with --force-reinstall, overriding the base-image version.

Confirmation that the fix is actually live in the built image comes from the release-sglang-docker.yml PR dry-run on this branch: the initial build attempt (commit f091a01) successfully installed from source and got past request_manager.py:38 — the line the original get_zmq_socket ImportError surfaced at — and only failed deeper in the chain at sgl_kernel loading (which requires libcuda.so.1 and a real GPU driver, unavailable on a build runner). If the stale base-image servicer had still been active, the failure would have matched the production traceback exactly. That it didn't is proof --force-reinstall from source is in effect.

Note on defense-in-depth

I initially added a RUN python -c "from smg_grpc_servicer.sglang.servicer import ..." smoke test to the Dockerfile as defense-in-depth, but reverted it (commit 644980f) once CI proved it unrunnable: the servicer's import chain transitively loads sglang's sgl_kernel, which dlopens libcuda.so.1 at module init time and requires an NVIDIA driver that doesn't exist on CPU-only docker build runners. The appropriate place for a full import smoke test is a post-build job in release-sglang-docker.yml that runs the built image on a GPU runner — that's a separate workflow change and out of scope for this PR.

Local verification

  • bash -n scripts/installation/install-smg.sh — syntax OK
  • shellcheck scripts/installation/install-smg.sh — clean
  • pre-commit run --files docker/engine.Dockerfile scripts/installation/install-smg.sh — all hooks pass
  • cargo +nightly fmt --all -- --check — clean (Rust workspace unperturbed; no Rust files in diff)
Checklist
  • cargo +nightly fmt passes (no-op — zero Rust changes)
  • cargo clippy --all-targets --all-features -- -D warnings (no-op — zero Rust changes)
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Build & Installation
    • Removed a previous special-case dependency install for the vLLM engine from the final Docker stage.
    • Added a forced reinstall of gRPC-related Python packages from source after building bindings to ensure installed packages match the current codebase and avoid stale versions.

The engine images inherited smg-grpc-servicer from whatever the base image
(e.g. lmsysorg/sglang:v0.5.10) happened to preinstall from PyPI at its own
build time, rather than from the SMG source tree being baked into the image.
When sglang v0.5.10 refactored srt/utils into a subpackage and moved
get_zmq_socket into srt/utils/network, the stale bundled servicer broke at
import time with "cannot import name 'get_zmq_socket' from 'sglang.srt.utils'",
crash-looping the engine pod despite v1.4.1 source already containing the fix.

CI never caught this because scripts/ci_install_sglang.sh installs the
servicer from source (`uv pip install -e grpc_servicer/`), so CI and the
published image were testing two different versions of the same package.

Fix: install smg-grpc-proto and smg-grpc-servicer from the cloned source
tree in install-smg.sh, using --force-reinstall to override any stale
version present in the base image. Drop the now-redundant vLLM-only PyPI
install from the Dockerfile. Add an import smoke test at build time so an
upstream engine refactor that breaks our servicer fails the image build
immediately instead of surfacing as a CrashLoopBackOff in production.

Signed-off-by: key4ng <rukeyang@gmail.com>
@github-actions github-actions Bot added the docker Docker configuration changes label Apr 10, 2026
@coderabbitai

coderabbitai Bot commented Apr 10, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6e065973-f696-4ca2-8b4e-8453d9383165

📥 Commits

Reviewing files that changed from the base of the PR and between f091a01 and 644980f.

📒 Files selected for processing (1)
  • docker/engine.Dockerfile
💤 Files with no reviewable changes (1)
  • docker/engine.Dockerfile

📝 Walkthrough

Walkthrough

The PR removes a vLLM-specific pip install step from docker/engine.Dockerfile and adds a forced pip reinstall of smg-grpc-proto and smg-grpc-servicer from source in scripts/installation/install-smg.sh. Engine install control flow and ENGINE validation remain unchanged.

Changes

Cohort / File(s) Summary
Docker Engine Build
docker/engine.Dockerfile
Removed the vLLM-specific pip install --no-cache-dir smg-grpc-proto "smg-grpc-servicer[vllm]" step in the final stage; left ENGINE validation and engine-source install script invocation unchanged.
SMG Installation Script
scripts/installation/install-smg.sh
Added pip install --no-cache-dir --force-reinstall calls to install smg-grpc-proto and smg-grpc-servicer from local source paths after building Python wheels, ensuring repo source versions are installed.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~12 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • gongwei-130

Poem

🐰
Fresh wheels spun in morning light,
Reinstalling to make things right,
No stale bits left in the den,
Builds hop forward, try again—
A rabbit cheers: new source, new flight! 🥕✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly describes the main change: installing smg-grpc-servicer from source in engine images, which is the core fix addressing the stale package issue documented in the PR objectives.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the Docker build and installation scripts to install smg-grpc-proto and smg-grpc-servicer directly from source, ensuring the container image remains synchronized with the repository. Additionally, a smoke-test step was added to the Dockerfile to verify successful imports of engine-specific servicers. Feedback suggests using python3 instead of python in the smoke-test to ensure compatibility across different base images.

Comment thread docker/engine.Dockerfile Outdated
The smoke test added alongside the source install cannot run at docker
build time: importing smg_grpc_servicer.sglang.servicer transitively
loads sglang's sgl_kernel, which dlopens libcuda.so.1 and requires an
NVIDIA driver. Build runners don't have GPUs, so it always fails:

    ImportError: libcuda.so.1: cannot open shared object file
    ModuleNotFoundError: No module named 'common_ops'

The source-install fix is confirmed working by the same failure — the
build got past request_manager.py (where the original get_zmq_socket
ImportError would have surfaced if the stale base-image servicer were
still active) and only failed deeper in the chain when sgl_kernel tried
to load. That's the bug this PR fixes, and removing the unrunnable
smoke test leaves the actual fix intact.

A proper defense-in-depth smoke test needs to run against the built
image on a GPU runner, which is a separate workflow change.

Signed-off-by: key4ng <rukeyang@gmail.com>
@slin1237
slin1237 merged commit 6c39621 into smg-project:main Apr 10, 2026
34 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docker Docker configuration changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants