Skip to content

fix(ci): add reusable composite actions for backend setup - #319

Merged
CatherineSue merged 6 commits into
mainfrom
chang/ci
Feb 4, 2026
Merged

CatherineSue merged 6 commits into
mainfrom
chang/ci

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Feb 4, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

After migrating to ephemeral k8s-runner-gpu runners, multiple jobs independently install inference backends (SGLang, vLLM, TRT-LLM) from scratch every run. The setup logic is duplicated across the workflow with no pip caching for SGLang jobs, and only partial caching (~/.cache/pip/wheels) for vLLM.

Solution

Create reusable composite actions (following the existing setup-rust pattern) that encapsulate venv creation, pip/wheel caching, and backend installation:

  • setup-sglang — venv + pip cache (~/.cache/pip) + ci_install_sglang.sh
  • setup-vllm — venv + pip cache (~/.cache/pip) + ci_install_vllm.sh
  • setup-trtllm — venv + TRT-LLM wheel cache restore + ci_install_trtllm.sh + wheel cache save

Changes

  • Add .github/actions/setup-sglang/action.yml composite action
  • Add .github/actions/setup-vllm/action.yml composite action
  • Add .github/actions/setup-trtllm/action.yml composite action (includes restore/save of /tmp/trtllm-wheel)
  • Replace inline backend setup in gateway-e2e, go-bindings-e2e, and go-bindings-benchmark jobs with composite action calls
  • SGLang jobs gain pip caching they didn't have before
  • vLLM pip cache expanded from ~/.cache/pip/wheels to all of ~/.cache/pip

Test Plan

  • gateway-e2e matrix entries without setup_vllm/setup_trtllm use setup-sglang
  • chat-completions-vllm and vllm-pd entries use setup-vllm
  • chat-completions-trtllm uses setup-trtllm
  • go-bindings-e2e and go-bindings-benchmark use setup-sglang
  • Second CI run benefits from pip cache (check for Cache restored in logs)
  • All tests pass
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Chores
    • Consolidated CI backend setup into reusable setup actions for vLLM, TRT-LLM, and SGLang to simplify pipeline steps.
    • Standardized cache restore/save behavior per backend to improve build efficiency and reduce redundant logic.
    • Replaced inline installation sequences with dedicated setup actions and unified conditional backend selection across jobs for more consistent CI runs.

Replace duplicated inline backend installation steps with reusable
composite actions (setup-sglang, setup-vllm, setup-trtllm) that
encapsulate venv creation, pip/wheel caching, and backend install.
@github-actions github-actions Bot added the ci CI/CD configuration changes label Feb 4, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly refactors the continuous integration (CI) setup for inference backends by introducing reusable GitHub composite actions. The primary goal is to optimize CI run times and reduce resource consumption by centralizing installation logic, implementing more effective caching strategies for Python dependencies and pre-built wheels, and eliminating redundant setup steps across various jobs. This change improves the efficiency and maintainability of the CI pipeline.

Highlights

  • New Reusable Composite Actions: Introduced three new GitHub composite actions: setup-sglang, setup-vllm, and setup-trtllm. These actions centralize the setup logic for their respective inference backends, including Python virtual environment creation, dependency installation, and robust caching mechanisms.
  • Improved Caching Strategy: Enhanced caching for SGLang and vLLM installations. SGLang jobs now benefit from a comprehensive pip cache, which was previously absent. The vLLM pip cache has been expanded from only ~/.cache/pip/wheels to the entire ~/.cache/pip directory, improving cache hit rates and reducing build times.
  • Streamlined CI Workflows: Replaced repetitive, inline backend setup steps in gateway-e2e, go-bindings-e2e, and go-bindings-benchmark jobs with calls to the new composite actions. This reduces duplication and simplifies maintenance of CI configurations.
  • TRT-LLM Wheel Caching: The setup-trtllm action now includes specific logic to restore and save TRT-LLM wheel caches in /tmp/trtllm-wheel, ensuring that pre-built wheels are reused across CI runs when available.
Changelog
  • .github/actions/setup-sglang/action.yml
    • Added a new composite action to encapsulate the setup of the SGLang backend, including Python venv creation, pip cache restoration, and SGLang installation.
  • .github/actions/setup-trtllm/action.yml
    • Added a new composite action for TRT-LLM backend setup, which handles Python venv creation, restoration and saving of TRT-LLM wheel caches, and TRT-LLM installation.
  • .github/actions/setup-vllm/action.yml
    • Added a new composite action to streamline the setup of the vLLM backend, covering Python venv creation, pip cache restoration, and vLLM installation.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/pr-test-rust.yml
Activity
  • The author has confirmed that cargo +nightly fmt and cargo clippy --all-targets --all-features -- -D warnings checks pass.
  • A detailed test plan is provided, outlining how the new composite actions are integrated into existing CI jobs (gateway-e2e, chat-completions-vllm, chat-completions-trtllm, go-bindings-e2e, go-bindings-benchmark) and how cache benefits can be verified by checking for 'Cache restored' in logs.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds three new GitHub composite actions to set up SGLang, TRT-LLM, and vLLM backends in CI, and updates the PR test workflow to invoke these actions instead of inline installation and cache steps across jobs. (30 words)

Changes

Cohort / File(s) Summary
Backend Setup Actions
.github/actions/setup-sglang/action.yml, .github/actions/setup-trtllm/action.yml, .github/actions/setup-vllm/action.yml
Added three composite actions that create a Python virtualenv and run backend-specific install scripts. TRT-LLM action includes explicit wheel cache restore/save; vLLM and SGLang actions restore/save pip/uv caches.
CI Workflow Consolidation
.github/workflows/pr-test-rust.yml
Replaced inline backend installation and cache steps with calls to the new composite actions; added conditional to run SGLang setup only when neither vLLM nor TRT-LLM is selected; removed duplicate caching and inline install logic.

Sequence Diagram(s)

sequenceDiagram
    participant WF as Workflow Job
    participant ACT as Composite Action
    participant RUN as Runner Shell
    participant CACHE as Cache Service
    participant SCRIPTS as Repo Scripts

    WF->>ACT: invoke setup-<backend> composite
    ACT->>RUN: run scripts/ci_setup_python_venv.sh
    ACT->>CACHE: restore cache (pip/uv or trtllm wheel)
    ACT->>RUN: run scripts/ci_install_<backend>.sh
    alt cache save required
        ACT->>CACHE: save cache (if cache miss)
    end
    ACT-->>WF: finish
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Poem

🐇 I hopped through CI at dawn's first light,

Virtualenv snug, caches held tight.
Three tiny actions, each clears the way,
SGLang, TRT, vLLM — ready to play.
🥕

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding reusable composite GitHub Actions for centralizing backend setup (SGLang, vLLM, TRT-LLM) across CI jobs.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/ci

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces reusable composite actions for setting up SGLang, vLLM, and TRT-LLM backends, which is a great step towards reducing duplication and improving maintainability in the CI workflows. The approach of encapsulating setup logic is solid.

I've found a couple of issues related to caching that could prevent the new actions from working as expected. Specifically, the setup-vllm action appears to be caching the wrong directory for uv, and the setup-trtllm action uses a static cache key which could lead to using stale artifacts. My review includes suggestions to fix these issues. Once these are addressed, the CI should be more robust and efficient.

Comment thread .github/actions/setup-trtllm/action.yml Outdated
Comment thread .github/actions/setup-trtllm/action.yml Outdated
Comment thread .github/actions/setup-vllm/action.yml Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In @.github/actions/setup-trtllm/action.yml:
- Around line 11-27: The cache key for the TRT-LLM cache (used in the
trtllm-cache step and the save step) is too static and can serve stale wheels;
update the key used in both places (currently trtllm-wheel-${{ runner.os
}}-cuda13-v2) to include a hash of the install script and/or related artifacts
so the cache invalidates when scripts/ci_install_trtllm.sh changes (e.g., add
${{ hashFiles('scripts/ci_install_trtllm.sh') }} into the key expression),
ensuring both the restore and save steps use the identical new key.

Comment thread .github/actions/setup-trtllm/action.yml Outdated
Invalidate the TRT-LLM wheel cache when ci_install_trtllm.sh changes,
consistent with the sglang and vllm composite actions.
ci_install_vllm.sh uses uv, which caches to ~/.cache/uv not ~/.cache/pip.
coderabbitai[bot]
coderabbitai Bot previously approved these changes Feb 4, 2026
slin1237
slin1237 previously approved these changes Feb 4, 2026
- vLLM: switch from actions/cache@v4 to split cache/restore + cache/save
  so the save happens right after install instead of during post-job
  cleanup (avoids 9+ min tar+upload of ~/.cache/uv at job end)
- TRT-LLM: add restore-keys prefix fallback so the old cached wheel
  (under the previous static key) can still be found
Same pattern as vllm and trtllm: save immediately after install instead
of relying on post-job cleanup.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants