Skip to content

test(zmq): direct-backend e2e tests and CI lanes - #2060

Merged
slin1237 merged 1 commit into
mainfrom
zmq/07-e2e-tests
Aug 11, 2026
Merged

slin1237 merged 1 commit into
mainfrom
zmq/07-e2e-tests

Conversation

@slin1237

@slin1237 slin1237 commented Aug 5, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

The direct-ZMQ backend and its e2e harness are not covered by tests or run in CI,
so regressions in ZMQ dispatch, capability validation, or the headless command
builders would go undetected. Nothing has ever exercised the ZMQ wire end to end
on a GPU lane.

Solution

Add unit tests for the ZMQ harness and wire the ZMQ lanes into CI: a connection_mode
input on the reusable GPU e2e job, and an e2e-1gpu-chat-zmq lane that runs the
existing single-worker chat suite over the ZMQ wire for both vLLM and TokenSpeed.

Changes

  • e2e_test/infra/{test_connection_mode.py,test_zmq_cmd_builders.py} —
    connection-mode parsing, engine-capability validation, headless command builders.
  • e2e_test/fixtures/{test_connection_mode_validation.py,test_hooks_zmq_filter.py}
    — wire-family dedup/deselect hook coverage.
  • .github/workflows/e2e-gpu-job.yml — connection_mode input, folded into the
    log-artifact name so the gRPC and ZMQ legs don't collide.
  • .github/workflows/pr-test-rust.yml — the e2e-1gpu-chat-zmq lane, gated on the
    same change filters as the gRPC chat lane and added to finish.
  • scripts/ci_install_tokenspeed.sh — pin bump to a TokenSpeed revision with the
    ZMQ entrypoint.

Test Plan

  • pytest e2e_test/infra/test_connection_mode.py e2e_test/infra/test_zmq_cmd_builders.py e2e_test/fixtures/test_connection_mode_validation.py e2e_test/fixtures/test_hooks_zmq_filter.py
    → 33 passed, 3 skipped (the 3 skips need the smg wheel installed, which CI has).
  • The GPU e2e lane needs real hardware + models; it runs on this PR.
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Rebuilt on current main. The rest of the original 7-PR stack (core dispatch,
structured outputs, multimodal, EOS forwarding, harmony stop matcher, e2e infra)
has already landed in reworked form, so this is now a standalone tests+CI change.

@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 49e27355-2692-4aa8-9ca6-53d2e2aaa148

📥 Commits

Reviewing files that changed from the base of the PR and between 8fc1bd9 and 27281ab.

📒 Files selected for processing (1)
  • e2e_test/fixtures/hooks.py

📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added end-to-end GPU chat testing over ZMQ.
    • Added configurable connection modes for reusable test workflows.
    • Improved diagnostic artifact naming for parallel connection-mode tests.
  • Bug Fixes

    • Improved validation of supported connection modes.
    • Prevented duplicate or incompatible ZMQ test runs.
    • Skipped affected TokenSpeed tests on unsupported H100 configurations.
  • Tests

    • Added coverage for connection validation, command generation, filtering, and environment parsing.
  • Chores

    • Updated the TokenSpeed revision used during CI installation.

Walkthrough

The PR adds optional connection-mode workflow wiring, separates diagnostic artifacts by mode, adds ZMQ validation and command-builder tests, filters redundant ZMQ fixture cases, skips unsupported TokenSpeed tests, and runs vLLM and TokenSpeed ZMQ chat tests in CI.

Changes

ZMQ E2E coverage

Layer / File(s) Summary
Connection-mode workflow wiring
.github/workflows/e2e-gpu-job.yml, e2e_test/fixtures/test_connection_mode_validation.py, e2e_test/infra/test_connection_mode.py
The reusable GPU workflow accepts connection_mode, exports E2E_CONNECTION_MODE, and adds the mode to worker artifact names. Tests cover valid, blank, unset, and invalid values across engines.
ZMQ command and fixture behavior
e2e_test/infra/test_zmq_cmd_builders.py, e2e_test/fixtures/test_hooks_zmq_filter.py, e2e_test/fixtures/hooks.py
Tests cover IPC URLs, headless vLLM and TokenSpeed commands, handshake ports, gRPC URL preservation, deduplication, and worker-topology filtering. TokenSpeed gpt-oss tests are skipped.
ZMQ CI execution and gating
.github/workflows/pr-test-rust.yml, scripts/ci_install_tokenspeed.sh
CI adds a single-GPU ZMQ chat job for vLLM and TokenSpeed, includes it in final status gating, and updates the default TokenSpeed revision.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant CIWorkflow
  participant E2EGPUWorkflow
  participant E2EWorker
  CIWorkflow->>E2EGPUWorkflow: pass connection_mode
  E2EGPUWorkflow->>E2EWorker: export E2E_CONNECTION_MODE
  E2EWorker->>E2EGPUWorkflow: write mode-specific diagnostic artifact
Loading

Possibly related PRs

Suggested reviewers: catherinesue

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main changes: direct-backend ZMQ end-to-end tests and CI lanes.
Description check ✅ Passed The description directly explains the ZMQ testing gaps, CI changes, affected files, and test plan.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch zmq/07-e2e-tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Aug 5, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean PR — thorough unit tests for the ZMQ harness (connection-mode parsing, engine validation, headless command builders, wire-family dedup/deselect hooks) and well-structured CI integration. All test assertions verified against source implementations. No issues found.

0 🔴 Important · 0 🟡 Nit · 0 🟣 Pre-existing

Base automatically changed from zmq/06-e2e-infra to main August 6, 2026 06:18
@slin1237 slin1237 changed the title test(zmq): direct-backend e2e tests and CI lanes (7/7) test(zmq): direct-backend e2e tests and CI lanes Aug 8, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
e2e_test/fixtures/test_hooks_zmq_filter.py (1)

73-78: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

🟡 Nit — Test gRPC preference in both collection orders.

The test only checks [grpc, http]. Add [http, grpc] and still require that gRPC is kept. This detects an order-dependent filter that keeps HTTP when collection order changes.

As per coding guidelines, use the pr-test-analyzer agent to verify tests adequately cover new or changed functionality.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e_test/fixtures/test_hooks_zmq_filter.py` around lines 73 - 78, Extend
test_grpc_http_twins_collapse_to_grpc to exercise both collection orders,
including [http, grpc], and assert that grpc remains kept while http is
deselected in each case. Use the pr-test-analyzer agent to verify the tests
adequately cover this filtering behavior.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@e2e_test/fixtures/test_connection_mode_validation.py`:
- Around line 27-34: Expand the engine parameter list in
test_non_zmq_modes_accept_any_engine to include Runtime.TRTLLM.value and
Runtime.MLX.value alongside the existing runtimes, ensuring both GRPC and HTTP
modes are validated against every supported engine.
- Line 9: Remove the test-collection dependency on setup_backend in
test_connection_mode_validation.py so missing anthropic/openai SDKs cannot
silently skip _validate_connection_mode coverage. Move or reuse
_validate_connection_mode from a dependency-free location such as
infra.constants, and keep the validation tests runnable through the Python
unit-test path without cloud SDK imports.

---

Nitpick comments:
In `@e2e_test/fixtures/test_hooks_zmq_filter.py`:
- Around line 73-78: Extend test_grpc_http_twins_collapse_to_grpc to exercise
both collection orders, including [http, grpc], and assert that grpc remains
kept while http is deselected in each case. Use the pr-test-analyzer agent to
verify the tests adequately cover this filtering behavior.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 8871e97b-dd96-4223-9e5f-8028a51e913e

📥 Commits

Reviewing files that changed from the base of the PR and between 73da68c and 410a8b1.

📒 Files selected for processing (7)
  • .github/workflows/e2e-gpu-job.yml
  • .github/workflows/pr-test-rust.yml
  • e2e_test/fixtures/test_connection_mode_validation.py
  • e2e_test/fixtures/test_hooks_zmq_filter.py
  • e2e_test/infra/test_connection_mode.py
  • e2e_test/infra/test_zmq_cmd_builders.py
  • scripts/ci_install_tokenspeed.sh

from infra.constants import ConnectionMode, Runtime

# setup_backend pulls in the cloud SDKs; skip if the env lacks them.
setup_backend = pytest.importorskip("fixtures.setup_backend")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Locate the jobs that collect this test and the dependency-install steps.
rg -n -C 4 \
  'test_connection_mode_validation|python-unit-tests|ci_install_e2e_deps|fixtures\.setup_backend' \
  .github scripts e2e_test 2>/dev/null || true

Repository: smg-project/smg

Length of output: 9462


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== candidate workflow names =="
git ls-files .github/workflows scripts e2e_test | sort

echo
echo "== setup_backend contents =="
fd -a 'setup_backend\.py$|setup_backend/py$' e2e_test 2>/dev/null | sed 's#^\./##'
setup_files=$(fd 'setup_backend\.py$' e2e_test 2>/dev/null || true)
for f in $setup_files; do
  echo "--- $f ---"
  wc -l "$f"
  sed -n '1,220p' "$f"
done

echo
echo "== test_connection_mode_validation outline and contents =="
wc -l e2e_test/fixtures/test_connection_mode_validation.py
sed -n '1,220p' e2e_test/fixtures/test_connection_mode_validation.py

echo
echo "== dependency install script =="
sed -n '1,220p' scripts/ci_install_e2e_deps.sh

echo
echo "== pytest config and invocation references =="
rg -n -C 3 'pytest|e2e_test/fixtures/test_connection_mode_validation|fixtures/test_connection_mode_validation|test_connection_mode_validation|setup_backend|python-unit-tests|ci_install_e2e_deps' pyproject.toml pytest.ini setup.cfg tox.ini .github scripts pyproject*.toml 2>/dev/null || true

Repository: smg-project/smg

Length of output: 40171


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== setup_backend imports section =="
sed -n '1,70p' e2e_test/fixtures/setup_backend.py

echo
echo "== cloud SDK / setup_backend dependent imports =="
python3 - <<'PY'
import ast
from pathlib import Path
path = Path("e2e_test/fixtures/setup_backend.py")
tree = ast.parse(path.read_text(), path=path)
imports = []
functions = {}
classes = {}
for node in ast.walk(tree):
    if isinstance(node, ast.FunctionDef):
        functions[node.name] = (node.first_arg_line := node.lineno)
    elif isinstance(node, ast.ClassDef):
        classes[node.name] = node.lineno
for node in tree.body:
    if isinstance(node, ast.Import):
        for alias in node.names:
            imports.append(alias.name)
    elif isinstance(node, ast.ImportFrom):
        imports.append((node.module, {alias.name for alias in node.names}))
for name in ("_setup_cloud", "_make_client", "_start_worker", "_start_gateway", "setup_backend"):
    print(f"{name}: imports={functions.get(name, [])}")
print("top-level imports:", imports)
print("cloud packages in imports:", [x for x in imports if isinstance(x, str) and x in {"anthropic", "openai"}])
PY

echo
echo "== e2e_test metadata =="
for f in pyproject.toml setup.py setup.cfg poetry.lock uv.lock requirements*.txt; do
  [ -f "e2e_test/$f" ] && { echo "--- e2e_test/$f ---"; sed -n '1,220p' "e2e_test/$f"; }
done
for f in e2e_test/pyproject.toml setup.py setup.cfg; do [ -f "$f" ] && { echo "--- $f ---"; sed -n '1,220p' "$f"; }; done

echo
echo "== changed files hint =="
git status --short 2>/dev/null | sed -n '1,120p' || true

Repository: smg-project/smg

Length of output: 2456


🔴 Do not make connection-mode validation dependent on fixtures.setup_backend.

test_connection_mode_validation.py is collected with the Python unit-test path, but importing fixtures.setup_backend imports anthropic and openai. A missing cloud SDK makes these unit tests successful skips instead of running the _validate_connection_mode coverage. Move the validator to infra.constants, a dependency-free test helper, or make the skip fail/noise loudly instead of silencing the suite.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e_test/fixtures/test_connection_mode_validation.py` at line 9, Remove the
test-collection dependency on setup_backend in
test_connection_mode_validation.py so missing anthropic/openai SDKs cannot
silently skip _validate_connection_mode coverage. Move or reuse
_validate_connection_mode from a dependency-free location such as
infra.constants, and keep the validation tests runnable through the Python
unit-test path without cloud SDK imports.

Source: Coding guidelines

Comment on lines +27 to +34
@pytest.mark.parametrize("mode", [ConnectionMode.GRPC, ConnectionMode.HTTP])
@pytest.mark.parametrize(
"engine",
[Runtime.SGLANG.value, Runtime.VLLM.value, Runtime.TOKENSPEED.value],
)
def test_non_zmq_modes_accept_any_engine(mode, engine):
# Non-ZMQ wires impose no engine restriction.
setup_backend._validate_connection_mode(mode, engine)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

🟡 Nit — Cover all engines for non-ZMQ modes.

test_non_zmq_modes_accept_any_engine omits Runtime.TRTLLM and Runtime.MLX, although lines 20-24 establish that both are valid engine inputs for the same validator. Add them to prevent an accidental gRPC or HTTP restriction from passing untested.

Proposed test expansion
     "engine",
-    [Runtime.SGLANG.value, Runtime.VLLM.value, Runtime.TOKENSPEED.value],
+    [
+        Runtime.SGLANG.value,
+        Runtime.VLLM.value,
+        Runtime.TOKENSPEED.value,
+        Runtime.TRTLLM.value,
+        Runtime.MLX.value,
+    ],
 )

As per coding guidelines, use the pr-test-analyzer agent to verify tests adequately cover new or changed functionality.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
@pytest.mark.parametrize("mode", [ConnectionMode.GRPC, ConnectionMode.HTTP])
@pytest.mark.parametrize(
"engine",
[Runtime.SGLANG.value, Runtime.VLLM.value, Runtime.TOKENSPEED.value],
)
def test_non_zmq_modes_accept_any_engine(mode, engine):
# Non-ZMQ wires impose no engine restriction.
setup_backend._validate_connection_mode(mode, engine)
`@pytest.mark.parametrize`("mode", [ConnectionMode.GRPC, ConnectionMode.HTTP])
`@pytest.mark.parametrize`(
"engine",
[
Runtime.SGLANG.value,
Runtime.VLLM.value,
Runtime.TOKENSPEED.value,
Runtime.TRTLLM.value,
Runtime.MLX.value,
],
)
def test_non_zmq_modes_accept_any_engine(mode, engine):
# Non-ZMQ wires impose no engine restriction.
setup_backend._validate_connection_mode(mode, engine)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e_test/fixtures/test_connection_mode_validation.py` around lines 27 - 34,
Expand the engine parameter list in test_non_zmq_modes_accept_any_engine to
include Runtime.TRTLLM.value and Runtime.MLX.value alongside the existing
runtimes, ensuring both GRPC and HTTP modes are validated against every
supported engine.

Source: Coding guidelines

- Add unit tests for the ZMQ test harness: connection-mode parsing and
  engine-capability validation, headless ZMQ command builders, and the
  pytest wire-family dedup/deselect hooks.
- Wire the ZMQ lanes into CI: run the Rust ZMQ tests on the PR lane
  (installing TokenSpeed) and add the direct-ZMQ e2e job to the GPU lane.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
(cherry picked from commit 6a9a08a)
@slin1237
slin1237 merged commit d676365 into main Aug 11, 2026
46 checks passed
@slin1237
slin1237 deleted the zmq/07-e2e-tests branch August 11, 2026 03:34
slin1237 added a commit that referenced this pull request Aug 11, 2026
E2E_ZMQ_ENGINE_COUNT runs a ZMQ lane with grouped workers, riding the
lanes #2060 landed: the vLLM worker builder appends the engine-level
--data-parallel-size (flowing through the same smg serve launcher as
production), the gateway gains --zmq-engine-count so its handshake
awaits every engine, and start_workers sizes the worker's GPU slice as
tp x engine count. e2e-2gpu-chat-zmq-dp runs the chat suite with dp=2
vLLM groups on the 2-GPU runner, exercising the grouped handshake, the
connector's least-loaded selection, and the wave protocol live.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
slin1237 added a commit that referenced this pull request Aug 11, 2026
E2E_ZMQ_ENGINE_COUNT runs a ZMQ lane with grouped workers, riding the
lanes #2060 landed: the vLLM worker builder appends the engine-level
--data-parallel-size (flowing through the same smg serve launcher as
production), the gateway gains --zmq-engine-count so its handshake
awaits every engine, and start_workers sizes the worker's GPU slice as
tp x engine count. e2e-2gpu-chat-zmq-dp runs the chat suite with dp=2
vLLM groups on the 2-GPU runner, exercising the grouped handshake, the
connector's least-loaded selection, and the wave protocol live.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants