Repository navigation
fix(ci): add reusable composite actions for backend setup - #319
Conversation
Replace duplicated inline backend installation steps with reusable composite actions (setup-sglang, setup-vllm, setup-trtllm) that encapsulate venv creation, pip/wheel caching, and backend install.
Summary of ChangesHello @CatherineSue, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly refactors the continuous integration (CI) setup for inference backends by introducing reusable GitHub composite actions. The primary goal is to optimize CI run times and reduce resource consumption by centralizing installation logic, implementing more effective caching strategies for Python dependencies and pre-built wheels, and eliminating redundant setup steps across various jobs. This change improves the efficiency and maintainability of the CI pipeline. Highlights
Changelog
Ignored Files
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
📝 WalkthroughWalkthroughAdds three new GitHub composite actions to set up SGLang, TRT-LLM, and vLLM backends in CI, and updates the PR test workflow to invoke these actions instead of inline installation and cache steps across jobs. (30 words) Changes
Sequence Diagram(s)sequenceDiagram
participant WF as Workflow Job
participant ACT as Composite Action
participant RUN as Runner Shell
participant CACHE as Cache Service
participant SCRIPTS as Repo Scripts
WF->>ACT: invoke setup-<backend> composite
ACT->>RUN: run scripts/ci_setup_python_venv.sh
ACT->>CACHE: restore cache (pip/uv or trtllm wheel)
ACT->>RUN: run scripts/ci_install_<backend>.sh
alt cache save required
ACT->>CACHE: save cache (if cache miss)
end
ACT-->>WF: finish
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces reusable composite actions for setting up SGLang, vLLM, and TRT-LLM backends, which is a great step towards reducing duplication and improving maintainability in the CI workflows. The approach of encapsulating setup logic is solid.
I've found a couple of issues related to caching that could prevent the new actions from working as expected. Specifically, the setup-vllm action appears to be caching the wrong directory for uv, and the setup-trtllm action uses a static cache key which could lead to using stale artifacts. My review includes suggestions to fix these issues. Once these are addressed, the CI should be more robust and efficient.
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Fix all issues with AI agents
In @.github/actions/setup-trtllm/action.yml:
- Around line 11-27: The cache key for the TRT-LLM cache (used in the
trtllm-cache step and the save step) is too static and can serve stale wheels;
update the key used in both places (currently trtllm-wheel-${{ runner.os
}}-cuda13-v2) to include a hash of the install script and/or related artifacts
so the cache invalidates when scripts/ci_install_trtllm.sh changes (e.g., add
${{ hashFiles('scripts/ci_install_trtllm.sh') }} into the key expression),
ensuring both the restore and save steps use the identical new key.
Invalidate the TRT-LLM wheel cache when ci_install_trtllm.sh changes, consistent with the sglang and vllm composite actions.
ci_install_vllm.sh uses uv, which caches to ~/.cache/uv not ~/.cache/pip.
- vLLM: switch from actions/cache@v4 to split cache/restore + cache/save so the save happens right after install instead of during post-job cleanup (avoids 9+ min tar+upload of ~/.cache/uv at job end) - TRT-LLM: add restore-keys prefix fallback so the old cached wheel (under the previous static key) can still be found
59494a8
Same pattern as vllm and trtllm: save immediately after install instead of relying on post-job cleanup.
Description
Problem
After migrating to ephemeral
k8s-runner-gpurunners, multiple jobs independently install inference backends (SGLang, vLLM, TRT-LLM) from scratch every run. The setup logic is duplicated across the workflow with no pip caching for SGLang jobs, and only partial caching (~/.cache/pip/wheels) for vLLM.Solution
Create reusable composite actions (following the existing
setup-rustpattern) that encapsulate venv creation, pip/wheel caching, and backend installation:setup-sglang— venv + pip cache (~/.cache/pip) +ci_install_sglang.shsetup-vllm— venv + pip cache (~/.cache/pip) +ci_install_vllm.shsetup-trtllm— venv + TRT-LLM wheel cache restore +ci_install_trtllm.sh+ wheel cache saveChanges
.github/actions/setup-sglang/action.ymlcomposite action.github/actions/setup-vllm/action.ymlcomposite action.github/actions/setup-trtllm/action.ymlcomposite action (includes restore/save of/tmp/trtllm-wheel)gateway-e2e,go-bindings-e2e, andgo-bindings-benchmarkjobs with composite action calls~/.cache/pip/wheelsto all of~/.cache/pipTest Plan
gateway-e2ematrix entries withoutsetup_vllm/setup_trtllmusesetup-sglangchat-completions-vllmandvllm-pdentries usesetup-vllmchat-completions-trtllmusessetup-trtllmgo-bindings-e2eandgo-bindings-benchmarkusesetup-sglangCache restoredin logs)Checklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit