Skip to content

feat(e2e): EPD multimodal smoke CI on 4-gpu-h100 - #1924

Merged
slin1237 merged 13 commits into
mainfrom
feat/epd-e2e-smoke-ci
Jul 15, 2026
Merged

slin1237 merged 13 commits into
mainfrom
feat/epd-e2e-smoke-ci

Conversation

@key4ng

@key4ng key4ng commented Jul 14, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

TokenSpeed EPD (encode-prefill-decode) disaggregation is live in the gateway (--epd-disaggregation, --encode, RoutingMode::EncodePrefillDecode) and the TokenSpeed gRPC encode servicer, but there is no automated coverage — nothing exercises the encode→prefill→decode multimodal path on change.

Solution

Extend the e2e harness with EPD support and add a 4-gpu-h100 smoke that runs a TokenSpeed EPD vision model across four worker-count topologies (1e1p1d, 1e2p1d, 2e1p1d, 1e1p2d, all tp=1). The test asserts real EPD participation (the encode worker's own per-request accept log), not just that a plausible answer returned — a single-worker fallback fails it.

Changes

  • WorkerType.ENCODE + a per-request EPD encode: accepted marker in the TokenSpeed encode servicer.
  • worker.py: TokenSpeed disaggregation launch flags (--disaggregation-mode {encode|prefill|decode}, bootstrap port for encode+prefill, --disaggregation-transfer-backend mooncake, unique --dist-init-addr per worker, prefix-caching/enforce-eager per role, --skip-server-warmup) + NVLink/warmup env.
  • gateway.py: build_epd_mode_args() + an --epd-disaggregation launch mode (--encode/--prefill/--decode, --encode-policy consistent_hashing, --multimodal-tensor-transport inline).
  • setup_backend.py: an epd_grpc fixture mode (_setup_epd) launching N encode + N prefill + N decode workers with sequential GPU offsets.
  • model_specs.py: Qwen/Qwen3.5-9B (multimodal, tp=1, FA3, skip_tier_download).
  • test_epd_multimodal.py: the smoke test over the four topologies (request-scoped encode-acceptance assertion).
  • CI: e2e-4gpu-epd job in pr-test-rust.yml; skip_tier_download + extra_models model provisioning in e2e-gpu-job.yml / ci_download_model.sh.

Note: build_epd_mode_args intentionally leaves prefill/decode routing at the gateway defaults and only sets --encode-policy.

Test Plan

  • Unit tests (no GPU): pytest e2e_test/infra/test_epd_cmd_builders.py -v — 10 tests cover the disaggregation flags, EPD gateway args, and the model spec.
  • Full EPD path runs on the 4-gpu-h100 runner via e2e-4gpu-epd across all four topologies. First-run watch-items: Mooncake transport under non-privileged CI (de-risk 1e1p1d first) and encode-marker visibility under the worker --log-level warning.
Checklist
  • cargo +nightly fmt passes (N/A — no Rust changes beyond a one-line log)
  • cargo clippy --all-targets --all-features -- -D warnings passes (N/A)
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added encode–prefill–decode (EPD) routing for multimodal TokenSpeed workloads, including encode-worker support and EPD startup configuration.
    • Introduced a new 4-GPU E2E lane to validate EPD multimodal behavior.
    • E2E GPU workflows can now optionally download extra models via an extra_models input.
  • Bug Fixes
    • Improved tier-based model download selection to respect models marked to skip tier downloads.
  • Tests
    • Added EPD multimodal E2E coverage and unit tests for EPD command/worker configuration and model specs.
  • Chores
    • Updated TokenSpeed and CUDA 13.0 CI install/build setup for toolchain compatibility.

@coderabbitai

coderabbitai Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: b662dcd1-b74f-42bd-be7f-fa67458990d0

📥 Commits

Reviewing files that changed from the base of the PR and between 56f9891 and 96f91a7.

📒 Files selected for processing (1)
  • .github/workflows/pr-test-rust.yml

📝 Walkthrough

Walkthrough

Adds TokenSpeed EPD worker and gateway support, multimodal routing tests, a Qwen3.5-9B model specification, and CI workflow and CUDA installation updates for dedicated four-GPU EPD testing.

Changes

TokenSpeed EPD multimodal execution

Layer / File(s) Summary
EPD worker and gateway orchestration
e2e_test/infra/constants.py, e2e_test/infra/worker.py, e2e_test/infra/gateway.py, e2e_test/fixtures/setup_backend.py
Adds ENCODE workers, distributed initialization ports, TokenSpeed disaggregation commands and environment settings, EPD gateway arguments, and EPD fixture lifecycle management.
EPD model and command validation
e2e_test/infra/model_specs.py, e2e_test/infra/test_epd_cmd_builders.py
Defines the Qwen3.5-9B multimodal model configuration and tests EPD worker commands, gateway arguments, and model properties.
Multimodal E2E coverage
e2e_test/chat_completions/test_epd_multimodal.py, grpc_servicer/.../encoder_servicer.py
Adds color and animal image tests across worker topologies and logs accepted encode requests for dispatch verification.
EPD CI installation and workflow
.github/workflows/*, scripts/ci_download_model.sh, scripts/ci_install_tokenspeed.sh
Adds extra model downloads, a dedicated four-GPU EPD job, skip-tier model filtering, and CUDA 13/cu130 TokenSpeed installation handling.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant E2E_Test
  participant Gateway
  participant EncodeWorker
  participant PrefillWorker
  participant DecodeWorker
  E2E_Test->>Gateway: Send multimodal chat completion
  Gateway->>EncodeWorker: Dispatch image encode request
  EncodeWorker->>PrefillWorker: Transfer encoded data
  PrefillWorker->>DecodeWorker: Start token generation
  DecodeWorker-->>E2E_Test: Return completion response
Loading

Possibly related PRs

Suggested labels: multimodal, dependencies

Suggested reviewers: catherinesue, slin1237, gongwei-130, xinyuezhang369

Poem

I’m a rabbit routing tokens with care,
Through encode, prefill, decode in the air.
Red paws, blue paws, a pug in sight,
CUDA moons make kernels compile just right.
Hop, hop—EPD tests take flight!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 51.72% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title is concise and accurately highlights the main change: EPD multimodal smoke CI on 4-gpu-h100.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/epd-e2e-smoke-ci

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added ci CI/CD configuration changes grpc gRPC client and router changes tests Test changes labels Jul 14, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces EPD (Encode-Prefill-Decode) multimodal Chat Completions end-to-end tests and adds support for EPD disaggregation topologies. It includes updates to backend setup, gateway argument building, and worker management to support separate encode, prefill, and decode workers using TokenSpeed. Additionally, unit tests for the EPD command-builder logic are added, and the Qwen3.5-9B model is integrated. Feedback is provided regarding a potential security improvement: using tempfile.mkdtemp instead of a predictable path in the shared temporary directory to avoid permission conflicts or hijacking vulnerabilities.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

FIXTURES_DIR = Path(__file__).parent.parent / "fixtures" / "images"
DOG_IMAGE_PATH = FIXTURES_DIR / "dog.jpg" # Black labrador puppy (checked in)

_LOG_DIR = Path(tempfile.gettempdir()) / f"smg-e2e-epd-{os.getpid()}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-medium medium

Using a predictable path in the shared temporary directory (tempfile.gettempdir()) can lead to permission conflicts or hijacking vulnerabilities on shared systems or multi-user CI environments. It is safer and more robust to use tempfile.mkdtemp to create a uniquely named, securely permissioned directory at import time.

Suggested change
_LOG_DIR = Path(tempfile.gettempdir()) / f"smg-e2e-epd-{os.getpid()}"
_LOG_DIR = Path(tempfile.mkdtemp(prefix="smg-e2e-epd-"))

@key4ng
key4ng marked this pull request as ready for review July 15, 2026 03:44
key4ng added 11 commits July 14, 2026 20:45
Signed-off-by: key4ng <rukeyang@gmail.com>
…kers

Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
…alse pass)

Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
…d fixes

Signed-off-by: key4ng <rukeyang@gmail.com>
…pts reasoning_content

Signed-off-by: key4ng <rukeyang@gmail.com>
@key4ng
key4ng force-pushed the feat/epd-e2e-smoke-ci branch from 57608c2 to c8ab2e0 Compare July 15, 2026 03:45

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 57608c2d02

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +735 to +737
e2e-4gpu-epd:
name: e2e-4gpu-epd (tokenspeed)
needs: [e2e-1gpu-chat]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Include EPD job in the finish gate

Adding this workflow job does not make its result part of the aggregate finish check: I inspected the finish job in this same workflow, and its needs list and failure condition still omit e2e-4gpu-epd. If branch protection relies on finish, this new smoke can fail or still be running while finish reports success, so the EPD coverage added here would not actually block merges.

Useful? React with 👍 / 👎.

Comment thread e2e_test/infra/gateway.py Outdated
if pf.bootstrap_port is not None:
args.append(str(pf.bootstrap_port))
for dc in decode_workers:
args += ["--decode", dc.worker_url]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: This line uses dc.worker_url while the encode/prefill loops above use en.base_url / pf.base_url. Since worker_url is just an alias for base_url (worker.py:62-64), this is functionally identical — but mixing both names in the same 10-line function reads like they might be different properties. Using dc.base_url here would make the intent clearer.

Suggested change
args += ["--decode", dc.worker_url]
args += ["--decode", dc.base_url]

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
e2e_test/infra/worker.py (1)

579-600: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Release reserved ports when worker startup fails.

worker.start() can fail before establishing a live process, but stop_workers() calls Worker.stop(), which returns early when process is None or already exited. The newly reserved dist_init_port, bootstrap port, and service port therefore leak after failed starts.

Move reservation cleanup outside the process-liveness early return, or explicitly release these ports in the startup exception path.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e_test/infra/worker.py` around lines 579 - 600, Ensure reserved service,
bootstrap, and dist_init ports are released when Worker.start() fails before a
live process exists. Update Worker.stop() to perform port cleanup before its
process-liveness early return, or invoke equivalent cleanup from the startup
exception path, while preserving normal process termination behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/e2e-gpu-job.yml:
- Around line 134-136: Update the “Download extra models” workflow step so
inputs.extra_models is passed through the step’s env configuration rather than
interpolated into run. In the shell command, safely split the environment value
into arguments before invoking scripts/ci_download_model.sh, preserving support
for multiple model names without allowing input text to become shell syntax.

In `@e2e_test/chat_completions/test_epd_multimodal.py`:
- Around line 120-124: Replace the broad substring assertions in the multimodal
image tests around the answer validation with normalized exact-answer checks and
disjoint expected classifications for each image. Require the pug case to return
“pug” and ensure color cases accept only their intended normalized color,
preventing generic or negated text from passing.

In `@grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py`:
- Around line 201-205: Make the acceptance event in
grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py lines 201-205 use
a logger and level preserved in worker logs. Update the assertion in
e2e_test/chat_completions/test_epd_multimodal.py lines 76-78 to match the
current request ID rather than counting generic router dispatch markers, so
delayed or unrelated requests cannot satisfy it.

---

Outside diff comments:
In `@e2e_test/infra/worker.py`:
- Around line 579-600: Ensure reserved service, bootstrap, and dist_init ports
are released when Worker.start() fails before a live process exists. Update
Worker.stop() to perform port cleanup before its process-liveness early return,
or invoke equivalent cleanup from the startup exception path, while preserving
normal process termination behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 532489ff-9008-4809-b1c0-3b0ef7daa1ea

📥 Commits

Reviewing files that changed from the base of the PR and between 90803e7 and c8ab2e0.

📒 Files selected for processing (12)
  • .github/workflows/e2e-gpu-job.yml
  • .github/workflows/pr-test-rust.yml
  • e2e_test/chat_completions/test_epd_multimodal.py
  • e2e_test/fixtures/setup_backend.py
  • e2e_test/infra/constants.py
  • e2e_test/infra/gateway.py
  • e2e_test/infra/model_specs.py
  • e2e_test/infra/test_epd_cmd_builders.py
  • e2e_test/infra/worker.py
  • grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py
  • scripts/ci_download_model.sh
  • scripts/ci_install_tokenspeed.sh

Comment thread .github/workflows/e2e-gpu-job.yml Outdated
Comment on lines +120 to +124
# (2) The answer is correct about the image — only possible if the encoder's
# embeddings reached prefill+decode (EPD-only gateway; no fallback path).
assert any(k in text.lower() for k in keywords), (
f"expected one of {keywords} in the answer, got: {text!r}"
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Require answers that distinguish each image.

Substring matching allows false passes: both animal cases accept "dog"/"puppy", and generic or negated text containing a color also passes. The test can therefore succeed without proving that different images produced different classifications.

Use normalized exact answers for colors and disjoint expectations—such as requiring "pug" for the pug image—or use images from different species.

Also applies to: 136-152

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e_test/chat_completions/test_epd_multimodal.py` around lines 120 - 124,
Replace the broad substring assertions in the multimodal image tests around the
answer validation with normalized exact-answer checks and disjoint expected
classifications for each image. Require the pug case to return “pug” and ensure
color cases accept only their intended normalized color, preventing generic or
negated text from passing.

Comment on lines +201 to +205
logger.info(
"EPD encode: accepted request_id=%s room=%s",
request.request_id,
bootstrap_room,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Make the acceptance assertion request-scoped and observable.

The new acceptance event is logged at INFO, while the E2E test states that TokenSpeed logging suppresses it and instead counts an unscoped router marker. An unrelated or delayed dispatch can therefore satisfy the assertion.

  • grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py#L201-L205: emit the acceptance event through a logger/level preserved in worker logs.
  • e2e_test/chat_completions/test_epd_multimodal.py#L76-L78: correlate the assertion with the current request ID instead of counting generic dispatch lines.
📍 Affects 2 files
  • grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py#L201-L205 (this comment)
  • e2e_test/chat_completions/test_epd_multimodal.py#L76-L78
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py` around lines
201 - 205, Make the acceptance event in
grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py lines 201-205 use
a logger and level preserved in worker logs. Update the assertion in
e2e_test/chat_completions/test_epd_multimodal.py lines 76-78 to match the
current request ID rather than counting generic router dispatch markers, so
delayed or unrelated requests cannot satisfy it.

…args, mkdtemp log dir

Signed-off-by: key4ng <rukeyang@gmail.com>
@key4ng

key4ng commented Jul 15, 2026

Copy link
Copy Markdown
Member Author

Thanks for the reviews! e2e-4gpu-epd (tokenspeed) is now green across all four topologies (1e1p1d / 1e2p1d / 2e1p1d / 1e1p2d) on the rebased branch, and e2e-1gpu-chat (tokenspeed) passed too. Addressed the findings in 56f98911:

Fixed

  • Template injection — e2e-gpu-job.yml (extra_models) (@coderabbitai): the input is now passed through env: EXTRA_MODELS and referenced as $EXTRA_MODELS, so it can never become shell syntax; unquoted expansion still word-splits into multiple model ids.
  • Predictable temp dir (@gemini-code-assist): _LOG_DIR now uses tempfile.mkdtemp(prefix="smg-e2e-epd-").
  • dc.worker_url vs base_url (@claude): build_epd_mode_args now uses dc.base_url for --decode, consistent with encode/prefill.

Intentionally deferred

  • Not in the finish gate (@chatgpt-codex-connector): correct, and deliberate — this is a new GPU smoke and gating merges on it before it's proven reliably green would block unrelated PRs. I'll add e2e-4gpu-epd to finish's needs + result check in a follow-up now that it's passing.

Kept as-is (with rationale)

  • Substring matching allows false passes (@coderabbitai): the red vs blue cases are already disjoint — the red image must yield "red" and the blue image "blue" — so the suite already proves input-specific output, not a generic pass. Since Qwen3.5 is a thinking model (the answer can arrive in reasoning_content), tightening the animal case to require the exact breed ("pug") risks smoke-test flakiness for little added signal, so I kept robust keyword matching. Happy to tighten if you'd prefer.

Note: the worker-side EPD encode: accepted marker doesn't survive TokenSpeed's logging reconfiguration, so per-request EPD participation is asserted from the router-side EPD encode dispatch issued log instead.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits. You can see your limits in the Codex usage dashboard.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Great to hear that the e2e-4gpu-epd suite is passing reliably across all topologies! Your approach to using disjoint keyword matching for the red/blue image tests is a sound way to ensure input-specific validation without introducing unnecessary flakiness from strict breed matching in the thinking model's output. Given that you've confirmed the EPD dispatch logging is now captured correctly via the router-side marker, this looks like a robust smoke test for the EPD disaggregation path. Regarding the finish gate, adding it as a needs dependency once it's proven stable is definitely the right move to maintain CI reliability.

@coderabbitai

coderabbitai Bot commented Jul 15, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

@key4ng Thanks for the detailed rundown — glad to hear e2e-4gpu-epd is green across all four topologies now.

The fixes all sound solid:

  • Passing extra_models through env: EXTRA_MODELS correctly closes the template-injection vector while still allowing word-splitting for multiple model ids.
  • tempfile.mkdtemp(prefix="smg-e2e-epd-") is the right call for a non-predictable temp dir.
  • Using dc.base_url for --decode keeps it consistent with the encode/prefill argument construction.

On the deferred items:

  • Makes sense to hold off on adding e2e-4gpu-epd to the finish gate until it's proven stable over more runs — thanks for flagging that you'll follow up once it's reliably green.
  • On the substring/keyword matching for the animal case: your reasoning holds up — the red/blue color test already demonstrates input-specific output, and given Qwen3.5 can emit answers via reasoning_content, tightening to an exact breed match risks flakiness for marginal signal. I'm fine keeping it as-is.

Good call also documenting that the router-side EPD encode dispatch issued log is the durable signal for per-request EPD participation, since the worker-side marker doesn't survive TokenSpeed's logging reconfiguration — that's useful context for anyone touching this test later.

(ノ◕‿◕)ノ*:・゚✧


✏️ Learnings added
Learnt from: key4ng
URL: https://github.com/lightseekorg/smg/pull/1924

Timestamp: 2026-07-15T05:45:21.098Z
Learning: In lightseekorg/smg, for TokenSpeed EPD (encode-prefill-decode) e2e tests, the worker-side "EPD encode: accepted" log marker in grpc_servicer/smg_grpc_servicer/tokenspeed/encoder_servicer.py does not survive TokenSpeed's logging reconfiguration. Per-request EPD encode-worker participation should instead be asserted from the router-side "EPD encode dispatch issued" log (in e2e_test/chat_completions/test_epd_multimodal.py), not the worker-side marker.

Learnt from: key4ng
URL: https://github.com/lightseekorg/smg/pull/1924

Timestamp: 2026-07-15T05:45:21.098Z
Learning: In lightseekorg/smg, the new `e2e-4gpu-epd (tokenspeed)` CI job (added in .github/workflows/pr-test-rust.yml, covering EPD multimodal smoke tests) is intentionally NOT included in the `finish` job's `needs`/result gate initially, since it's a new GPU smoke test and gating unrelated PR merges on it before it's proven reliably green would be disruptive. The author (key4ng) plans to add it to the finish gate in a follow-up once it's stable.

Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.

Signed-off-by: key4ng <rukeyang@gmail.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
e2e_test/infra/gateway.py (1)

136-148: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Reject incomplete EPD/PD topologies before launch. encode_workers=[] or a missing prefill_workers/decode_workers side is treated as a valid mode, converted to empty lists, and started with only a health check. That can leave the gateway reporting healthy until the first request. Require all worker lists for EPD, and both lists for PD, or fail fast on partial configs.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e_test/infra/gateway.py` around lines 136 - 148, Validate the worker-list
contents before mode selection in the gateway initialization flow: EPD requires
non-empty encode_workers, prefill_workers, and decode_workers, while PD requires
non-empty prefill_workers and decode_workers. Reject partial or empty
configurations with ValueError before launch, and ensure only complete
topologies contribute to modes_specified.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/e2e-gpu-job.yml:
- Around line 134-141: Update the “Download extra models” step to parse
EXTRA_MODELS into a shell array using space-separated words, then invoke
scripts/ci_download_model.sh with the array elements quoted via "${models[@]}".
Preserve the existing environment-variable-based input handling and conditional
execution while preventing pathname expansion.

In `@e2e_test/chat_completions/test_epd_multimodal.py`:
- Line 48: Replace the module-level tempfile.mkdtemp call used for _LOG_DIR with
pytest-managed temporary storage, or add explicit teardown that removes the
created directory after the E2E tests complete. Preserve _LOG_DIR’s existing
Path-based usage while ensuring every smg-e2e-epd-* directory is cleaned up.

---

Outside diff comments:
In `@e2e_test/infra/gateway.py`:
- Around line 136-148: Validate the worker-list contents before mode selection
in the gateway initialization flow: EPD requires non-empty encode_workers,
prefill_workers, and decode_workers, while PD requires non-empty prefill_workers
and decode_workers. Reject partial or empty configurations with ValueError
before launch, and ensure only complete topologies contribute to
modes_specified.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 85443d0a-39e2-45e6-a5b5-5a1a370d70ce

📥 Commits

Reviewing files that changed from the base of the PR and between c8ab2e0 and 56f9891.

📒 Files selected for processing (3)
  • .github/workflows/e2e-gpu-job.yml
  • e2e_test/chat_completions/test_epd_multimodal.py
  • e2e_test/infra/gateway.py

Comment thread .github/workflows/e2e-gpu-job.yml
Comment thread e2e_test/chat_completions/test_epd_multimodal.py
@slin1237
slin1237 merged commit a4ea581 into main Jul 15, 2026
52 checks passed
@slin1237
slin1237 deleted the feat/epd-e2e-smoke-ci branch July 15, 2026 15:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes grpc gRPC client and router changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants