Repository navigation
Conversation
Signed-off-by: key4ng <rukeyang@gmail.com>
|
Note Gemini is unable to generate a summary for this pull request due to the file types involved not being currently supported. |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughWorkflow and Kubernetes runner configs updated: CI steps switched model/cache env from Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
📝 Coding Plan
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In @.github/workflows/nightly-benchmark.yml:
- Around line 201-202: The nightly benchmark workflow still references the old
PVC path for H200 jobs via the ROUTER_LOCAL_MODEL_PATH variable in the
single-worker-h200 job; update those occurrences to use the migrated path
(/models) consistently (e.g., replace ROUTER_LOCAL_MODEL_PATH="/raid/models"
with ROUTER_LOCAL_MODEL_PATH="/models" or switch the job to use
HF_HOME="/models" like the H100 jobs) so the matrix is consistent and runners
that expect the mounted PVC path will find models; search for
ROUTER_LOCAL_MODEL_PATH in the workflow (single-worker-h200 job and the other
H200 invocation mentioned) and make the same replacement in both places.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 7893e161-d3ad-451d-80ff-d9cef9c23709
📒 Files selected for processing (2)
.github/workflows/nightly-benchmark.yml.github/workflows/pr-test-rust.yml
| HF_HOME="/models" \ | ||
| pytest e2e_test/benchmarks/test_nightly_perf.py \ |
There was a problem hiding this comment.
Complete the model-path migration for H200 nightly job as well.
These updates move H100 jobs to HF_HOME="/models", but Line 429-431 in single-worker-h200 still uses ROUTER_LOCAL_MODEL_PATH="/raid/models". This leaves the nightly benchmark matrix inconsistent and can still fail on runners expecting the mounted PVC path.
Suggested patch
- E2E_RUNTIME=${{ matrix.variant.runtime }} \
- ROUTER_LOCAL_MODEL_PATH="/raid/models" \
+ E2E_RUNTIME=${{ matrix.variant.runtime }} \
+ HF_HOME="/models" \
pytest e2e_test/benchmarks/test_nightly_perf.py \
-k "$K_FILTER" \
-s -vv -o log_cli=true --log-cli-level=INFOAlso applies to: 314-315
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In @.github/workflows/nightly-benchmark.yml around lines 201 - 202, The nightly
benchmark workflow still references the old PVC path for H200 jobs via the
ROUTER_LOCAL_MODEL_PATH variable in the single-worker-h200 job; update those
occurrences to use the migrated path (/models) consistently (e.g., replace
ROUTER_LOCAL_MODEL_PATH="/raid/models" with ROUTER_LOCAL_MODEL_PATH="/models" or
switch the job to use HF_HOME="/models" like the H100 jobs) so the matrix is
consistent and runners that expect the mounted PVC path will find models; search
for ROUTER_LOCAL_MODEL_PATH in the workflow (single-worker-h200 job and the
other H200 invocation mentioned) and make the same replacement in both places.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f3ae25d073
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| - name: Run benchmarks | ||
| env: | ||
| ROUTER_LOCAL_MODEL_PATH: /home/ubuntu/models | ||
| HF_HOME: /models |
There was a problem hiding this comment.
Set ROUTER_LOCAL_MODEL_PATH for local model fixtures
This replacement drops the only variable that e2e_test/infra/model_specs.py uses to resolve mounted local model directories (_resolve_model_path reads ROUTER_LOCAL_MODEL_PATH only). In these jobs, setup_backend will now pass HF model IDs instead of PVC-backed local paths, so workers can fall back to remote Hub resolution; that breaks or times out on runners without Hub access/tokens or when /models is a local mirror rather than HF cache metadata. Keep ROUTER_LOCAL_MODEL_PATH=/models here (you can still set HF_HOME in addition).
Useful? React with 👍 / 👎.
Signed-off-by: key4ng <rukeyang@gmail.com>
There was a problem hiding this comment.
♻️ Duplicate comments (1)
.github/workflows/nightly-benchmark.yml (1)
420-433:⚠️ Potential issue | 🟠 MajorComplete H200 migration to
HF_HOME=/modelsand remove old path variable usage.
single-worker-h200still uses/raid/modelsandROUTER_LOCAL_MODEL_PATH, so the workflow is only partially migrated. This keeps H200 behavior inconsistent with H100 and can still fail on runners expecting/models.Suggested patch
- name: Run benchmark if: steps.filter.outputs.skip != 'true' env: GPU_TYPE: H200 E2E_NIGHTLY: "1" E2E_LOG_DIR: nightly_gateway_logs HF_TOKEN: ${{ secrets.HF_TOKEN }} - HF_HOME: /raid/models + HF_HOME: /models run: | mkdir -p nightly_gateway_logs bash scripts/ci_killall_sglang.sh "nuk_gpus" @@ E2E_RUNTIME=${{ matrix.variant.runtime }} \ - ROUTER_LOCAL_MODEL_PATH="/raid/models" \ + HF_HOME="/models" \ pytest e2e_test/benchmarks/test_nightly_perf.py \ -k "$K_FILTER" \ -s -vv -o log_cli=true --log-cli-level=INFO🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In @.github/workflows/nightly-benchmark.yml around lines 420 - 433, The workflow still sets HF_HOME=/raid/models and passes ROUTER_LOCAL_MODEL_PATH="/raid/models" for the H200 job; update the H200/nightly job to complete the H200 migration by setting HF_HOME=/models (replace any `/raid/models` occurrences) and remove usage of ROUTER_LOCAL_MODEL_PATH and any references to that env var (e.g., the assignment ROUTER_LOCAL_MODEL_PATH="/raid/models" and any consumer of ROUTER_LOCAL_MODEL_PATH); ensure the matrix entry for single-worker-h200 and any variant.grpc_only logic use HF_HOME=/models so H200 matches H100 behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Duplicate comments:
In @.github/workflows/nightly-benchmark.yml:
- Around line 420-433: The workflow still sets HF_HOME=/raid/models and passes
ROUTER_LOCAL_MODEL_PATH="/raid/models" for the H200 job; update the H200/nightly
job to complete the H200 migration by setting HF_HOME=/models (replace any
`/raid/models` occurrences) and remove usage of ROUTER_LOCAL_MODEL_PATH and any
references to that env var (e.g., the assignment
ROUTER_LOCAL_MODEL_PATH="/raid/models" and any consumer of
ROUTER_LOCAL_MODEL_PATH); ensure the matrix entry for single-worker-h200 and any
variant.grpc_only logic use HF_HOME=/models so H200 matches H100 behavior.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 8998fde1-3e3f-4598-a661-7740feba3f4e
📒 Files selected for processing (2)
.github/workflows/e2e-gpu-job.yml.github/workflows/nightly-benchmark.yml
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cec4bb830a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| E2E_RUNTIME: ${{ inputs.engine }} | ||
| E2E_GPU_TIER: ${{ inputs.gpu_tier }} | ||
| ROUTER_LOCAL_MODEL_PATH: /home/ubuntu/models | ||
| HF_HOME: /models |
There was a problem hiding this comment.
Restore ROUTER_LOCAL_MODEL_PATH in shared E2E GPU job
This change replaces ROUTER_LOCAL_MODEL_PATH with HF_HOME, but E2E model selection still resolves local fixtures only via ROUTER_LOCAL_MODEL_PATH (e2e_test/infra/model_specs.py, _resolve_model_path). With only HF_HOME set, workers receive HF model IDs instead of mounted PVC paths, so this workflow stops using /models/<org>/<model> and can fail or stall on runners without Hub access or complete HF cache metadata. Keep ROUTER_LOCAL_MODEL_PATH=/models here (you can set HF_HOME in addition).
Useful? React with 👍 / 👎.
|
|
||
| E2E_RUNTIME=${{ matrix.variant.runtime }} \ | ||
| ROUTER_LOCAL_MODEL_PATH="/raid/models" \ | ||
| HF_HOME="/models" \ |
There was a problem hiding this comment.
Preserve local model path env in nightly H100 benchmark runs
Switching the benchmark invocation to HF_HOME drops the variable that nightly benchmark plumbing uses for local model mounts: e2e_test/infra/model_specs.py resolves model paths from ROUTER_LOCAL_MODEL_PATH, and e2e_test/benchmarks/conftest.py uses that same env var to mount local model/tokenizer directories into genai-bench. As written, these H100 jobs no longer target mounted local model dirs and may regress to remote Hub resolution/download behavior. Set ROUTER_LOCAL_MODEL_PATH=/models for these pytest commands (optionally alongside HF_HOME).
Useful? React with 👍 / 👎.
|
Hi @key4ng, the DCO sign-off check has failed. All commits must include a To fix existing commits: # Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-leaseTo sign off future commits automatically:
|
There was a problem hiding this comment.
♻️ Duplicate comments (1)
.github/workflows/nightly-benchmark.yml (1)
422-435:⚠️ Potential issue | 🟠 MajorComplete the H200 migration away from
/raid/models.The H200 benchmark step still sets
HF_HOMEto/raid/models(Line 423) and still injectsROUTER_LOCAL_MODEL_PATH="/raid/models"(Line 434), so this remains inconsistent with the/modelsmigration.Suggested patch
- name: Run benchmark if: steps.filter.outputs.skip != 'true' env: GPU_TYPE: H200 E2E_NIGHTLY: "1" E2E_LOG_DIR: nightly_gateway_logs HF_TOKEN: ${{ secrets.HF_TOKEN }} - HF_HOME: /raid/models + HF_HOME: /models run: | mkdir -p nightly_gateway_logs bash scripts/ci_killall_sglang.sh "nuk_gpus" K_FILTER="${{ matrix.model.test_class }}" if [ "${{ matrix.variant.grpc_only }}" == "true" ]; then K_FILTER="${{ matrix.model.test_class }} and grpc" fi E2E_RUNTIME=${{ matrix.variant.runtime }} \ - ROUTER_LOCAL_MODEL_PATH="/raid/models" \ pytest e2e_test/benchmarks/test_nightly_perf.py \ -k "$K_FILTER" \ -s -vv -o log_cli=true --log-cli-level=INFO#!/bin/bash # Verify model-path migration consistency in workflows. rg -n 'ROUTER_LOCAL_MODEL_PATH|HF_HOME:\s*/raid/models|HF_HOME="/raid/models"' .github/workflows -C2Expected result after fix: no
/raid/modelsreferences in migrated benchmark/e2e paths.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In @.github/workflows/nightly-benchmark.yml around lines 422 - 435, Update the nightly benchmark workflow to finish the H200 migration by replacing hardcoded /raid/models references: change the HF_HOME environment variable (HF_HOME) and the injected ROUTER_LOCAL_MODEL_PATH value (ROUTER_LOCAL_MODEL_PATH) from "/raid/models" to the migrated path "/models" (or to the appropriate matrix-provided path if configurable), ensuring the pytest invocation for e2e_test/benchmarks/test_nightly_perf.py continues to receive the correct E2E_RUNTIME and model path; search for HF_HOME and ROUTER_LOCAL_MODEL_PATH in the job step and update both occurrences so no /raid/models remains.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Duplicate comments:
In @.github/workflows/nightly-benchmark.yml:
- Around line 422-435: Update the nightly benchmark workflow to finish the H200
migration by replacing hardcoded /raid/models references: change the HF_HOME
environment variable (HF_HOME) and the injected ROUTER_LOCAL_MODEL_PATH value
(ROUTER_LOCAL_MODEL_PATH) from "/raid/models" to the migrated path "/models" (or
to the appropriate matrix-provided path if configurable), ensuring the pytest
invocation for e2e_test/benchmarks/test_nightly_perf.py continues to receive the
correct E2E_RUNTIME and model path; search for HF_HOME and
ROUTER_LOCAL_MODEL_PATH in the job step and update both occurrences so no
/raid/models remains.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 053cc266-45d6-435c-9d53-9dcb34f832d0
📒 Files selected for processing (2)
.github/workflows/e2e-gpu-job.yml.github/workflows/nightly-benchmark.yml
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 339d123786
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| export CGO_LDFLAGS="-L$(pwd)/bindings/golang/target/release" | ||
| export LD_LIBRARY_PATH="$(pwd)/bindings/golang/target/release:$LD_LIBRARY_PATH" | ||
| SHOW_WORKER_LOGS=0 SHOW_ROUTER_LOGS=1 ROUTER_LOCAL_MODEL_PATH="/home/ubuntu/models" \ | ||
| SHOW_WORKER_LOGS=0 SHOW_ROUTER_LOGS=1 HF_HOME="/models" \ |
There was a problem hiding this comment.
Restore ROUTER_LOCAL_MODEL_PATH for Go bindings E2E
This step now sets only HF_HOME, but the Go bindings E2E path resolution still depends on ROUTER_LOCAL_MODEL_PATH: e2e_test/infra/model_specs.py only rewrites model IDs to mounted local paths when that variable is present, and e2e_test/bindings_go/conftest.py uses the resolved value for SGL_TOKENIZER_PATH. Without ROUTER_LOCAL_MODEL_PATH, the tokenizer path becomes a HF model ID and can fall back to Hub resolution/download (this job does not set HF_TOKEN), which can fail or hang on runners that rely on pre-mounted local models.
Useful? React with 👍 / 👎.
Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
339d123 to
9a3267e
Compare
There was a problem hiding this comment.
Actionable comments posted: 1
♻️ Duplicate comments (1)
.github/workflows/nightly-benchmark.yml (1)
422-424:⚠️ Potential issue | 🟠 MajorComplete the H200 model-path migration to avoid mixed behavior.
Line 423 sets
HF_HOME, but Line 434 still exportsROUTER_LOCAL_MODEL_PATH, so the H200 path migration remains inconsistent with H100 jobs and can still follow the legacy variable/path.Suggested patch
- HF_HOME: /raid/models + HF_HOME: /models ... - E2E_RUNTIME=${{ matrix.variant.runtime }} \ - ROUTER_LOCAL_MODEL_PATH="/raid/models" \ + E2E_RUNTIME=${{ matrix.variant.runtime }} \ + HF_HOME="/models" \ pytest e2e_test/benchmarks/test_nightly_perf.py \ -k "$K_FILTER" \ -s -vv -o log_cli=true --log-cli-level=INFOAlso applies to: 433-435
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In @.github/workflows/nightly-benchmark.yml around lines 422 - 424, The H200 job still exports the legacy ROUTER_LOCAL_MODEL_PATH causing mixed behavior; update the workflow so H200 uses the same HF_HOME-based path as H100 instead of exporting ROUTER_LOCAL_MODEL_PATH. Locate where ROUTER_LOCAL_MODEL_PATH is exported and either remove that export or set ROUTER_LOCAL_MODEL_PATH to reference HF_HOME (e.g., ROUTER_LOCAL_MODEL_PATH=${{ env.HF_HOME }} or equivalent) so all jobs consistently use HF_HOME for model location.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@scripts/ci_install_vllm.sh`:
- Line 23: The pip install invocation "uv pip install vllm --extra-index-url
https://wheels.vllm.ai/nightly/cu129 --index-strategy unsafe-best-match" is
using the unsafe index strategy; change it to a deterministic approach by either
pinning to an exact nightly wheel URL (e.g., replace the package spec with a
vllm @ https://wheels.vllm.ai/nightly/cu129/vllm-<date>-...whl), or switch the
pip flag to "--index-strategy first-index" so the primary index is preferred and
the nightly index is only a fallback, or add an explicit comment documenting why
"unsafe-best-match" is an acceptable CI trade-off for the vllm nightly
requirement.
---
Duplicate comments:
In @.github/workflows/nightly-benchmark.yml:
- Around line 422-424: The H200 job still exports the legacy
ROUTER_LOCAL_MODEL_PATH causing mixed behavior; update the workflow so H200 uses
the same HF_HOME-based path as H100 instead of exporting
ROUTER_LOCAL_MODEL_PATH. Locate where ROUTER_LOCAL_MODEL_PATH is exported and
either remove that export or set ROUTER_LOCAL_MODEL_PATH to reference HF_HOME
(e.g., ROUTER_LOCAL_MODEL_PATH=${{ env.HF_HOME }} or equivalent) so all jobs
consistently use HF_HOME for model location.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 3f8e87d8-bc90-46d7-a8d6-0f13b12b8370
📒 Files selected for processing (3)
.github/workflows/e2e-gpu-job.yml.github/workflows/nightly-benchmark.ymlscripts/ci_install_vllm.sh
|
|
||
| echo "Installing vLLM (nightly for smg-grpc-servicer support)..." | ||
| uv pip install vllm --extra-index-url https://wheels.vllm.ai/nightly/cu129 | ||
| uv pip install vllm --extra-index-url https://wheels.vllm.ai/nightly/cu129 --index-strategy unsafe-best-match |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# 1) Confirm the exact install line in repo
fd "ci_install_vllm.sh" --exec sed -n '20,26p' {}
# 2) Verify supported index strategy values in the CI toolchain image
if command -v uv >/dev/null 2>&1; then
uv --version
uv pip install --help | sed -n '/--index-strategy/,+12p'
else
printf 'uv is not installed in this sandbox. Re-run these two commands in the CI image:\n'
printf ' uv --version\n'
printf ' uv pip install --help | sed -n '\''/--index-strategy/,+12p'\''\n'
fiRepository: lightseekorg/smg
Length of output: 519
🌐 Web query:
uv pip install --index-strategy unsafe-best-match documentation
💡 Result:
uv pip install --index-strategy unsafe-best-match is documented in uv’s pip compatibility docs under the section about packages that exist on multiple indexes. It explains that unsafe-best-match is the closest behavior to pip when you use multiple indexes (e.g., --extra-index-url), but it is unsafe because it can enable dependency confusion risks. [1]
Configuration docs also list index-strategy (including unsafe-best-match) as a setting you can put under [tool.uv.pip] (and you can set it via the UV_INDEX_STRATEGY environment variable as well). [2]
Sources: uv “Compatibility with pip” docs (index strategies + warning). [1] uv “Settings” reference (the index-strategy setting / config location). [2]
🏁 Script executed:
# Check if there are any comments or git history around this line
git log --oneline -n 20 scripts/ci_install_vllm.sh 2>/dev/null || echo "Git not available"
# Check for any documentation or comments in the script
cat scripts/ci_install_vllm.sh
# Search for other uses of index-strategy in the codebase
rg "index-strategy" --type sh --type py --type yaml --type txtRepository: lightseekorg/smg
Length of output: 1314
🏁 Script executed:
# Check if vllm is pinned elsewhere or if there's a requirements file
fd "requirements" --type f | head -20
fd "pyproject" --type f | head -20
rg "vllm" --type py --type txt -A 2 -B 2 | head -50Repository: lightseekorg/smg
Length of output: 3866
Reconsider unsafe-best-match strategy even for nightly wheels.
Line 23 uses --index-strategy unsafe-best-match, which uv's documentation explicitly marks as unsafe for dependency-confusion risks. This can also cause non-deterministic CI resolution (different runs may resolve different nightly wheels).
Since your script requires nightly vllm for smg-grpc-servicer support, pinning an exact version isn't practical. Instead, consider:
- Pin the nightly wheel URL or build date (e.g.,
vllm @ https://wheels.vllm.ai/nightly/cu129/vllm-<date>-...whl) to ensure determinism. - Or use
--index-strategy first-index(only resolve from primary index, with extra-index as fallback) if the nightly index is reliable. - Or document why the unsafe strategy trade-off is acceptable in your CI context.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@scripts/ci_install_vllm.sh` at line 23, The pip install invocation "uv pip
install vllm --extra-index-url https://wheels.vllm.ai/nightly/cu129
--index-strategy unsafe-best-match" is using the unsafe index strategy; change
it to a deterministic approach by either pinning to an exact nightly wheel URL
(e.g., replace the package spec with a vllm @
https://wheels.vllm.ai/nightly/cu129/vllm-<date>-...whl), or switch the pip flag
to "--index-strategy first-index" so the primary index is preferred and the
nightly index is only a fallback, or add an explicit comment documenting why
"unsafe-best-match" is an acceptable CI trade-off for the vllm nightly
requirement.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9a3267e923
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| echo "Installing vLLM (nightly for smg-grpc-servicer support)..." | ||
| uv pip install vllm --extra-index-url https://wheels.vllm.ai/nightly/cu129 | ||
| uv pip install vllm --extra-index-url https://wheels.vllm.ai/nightly/cu129 --index-strategy unsafe-best-match |
There was a problem hiding this comment.
Keep safe index resolution for vLLM install
uv help pip install documents that the default first-index strategy is what prevents dependency-confusion attacks, while unsafe-best-match searches all indexes and picks the “best” version across them. In this CI path we install from both PyPI and wheels.vllm.ai, so switching to unsafe-best-match can silently pull a different package build (or transitive dependency) than the intended CUDA-specific/nightly source, making installs less predictable and weakening supply-chain safety for PR jobs.
Useful? React with 👍 / 👎.
Signed-off-by: key4ng <rukeyang@gmail.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bfc2a8013b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| E2E_NIGHTLY: "1" | ||
| E2E_LOG_DIR: nightly_gateway_logs | ||
| HF_TOKEN: ${{ secrets.HF_TOKEN }} | ||
| HF_HOME: /raid/models |
There was a problem hiding this comment.
Avoid duplicating /raid/models bind mount in H200 benchmark
single-worker-h200 now exports HF_HOME: /raid/models while the pytest command still sets ROUTER_LOCAL_MODEL_PATH="/raid/models". In e2e_test/benchmarks/conftest.py::_build_command, those two env vars each add a -v /raid/models:/raid/models mount, creating duplicate destination mounts in the generated docker run command. Docker treats duplicate mount destinations as an error (Duplicate mount point), so this can fail nightly H200 benchmark runs before tests execute.
Useful? React with 👍 / 👎.
Signed-off-by: key4ng <rukeyang@gmail.com>
Description
Problem
K8s GPU runners have a 2Ti
model-cachePVC mounted at/models, but:ROUTER_LOCAL_MODEL_PATH="/home/ubuntu/models"(bare-metal path) or"/raid/models"which don't exist on k8s runners, causing models to fall back to HuggingFace download every run/modelsPVC is root-owned (drwxr-xr-x root root), so the runner user (UID 1001) cannot write to it — downloads to~/.cache/huggingfaceon ephemeral node disk insteadHF_TOKENfor download, which was not passed to workflowspull_requesttrigger was incorrectly nested underworkflow_dispatch.inputsuv pip install vllmfails due to index strategy not matching platform tagsSolution
HF_HOME=/modelson k8s workflows so HuggingFace downloads persist to the 2Ti PVCfsGroup: 1001+fsGroupChangePolicy: OnRootMismatchto all GPU runner pod specs so the PVC is writable by the runner userHF_TOKENfrom GitHub secrets to all GPU workflowsROUTER_LOCAL_MODEL_PATH="/raid/models"for H200 bare-metal runners (unchanged)pull_requesttrigger indentation--index-strategy unsafe-best-matchChanges
.github/workflows/e2e-gpu-job.yml—HF_HOME=/models,HF_TOKEN.github/workflows/nightly-benchmark.yml—HF_HOME=/modelsfor H100 jobs,HF_TOKEN+HF_HOME=/raid/modelsfor H200, fix PR trigger.github/workflows/pr-test-rust.yml—HF_HOME=/modelsfor k8s GPU jobsscripts/k8s-runner-resources/runner-values-*.yaml— addfsGroup: 1001,fsGroupChangePolicy: OnRootMismatchscripts/ci_install_vllm.sh— add--index-strategy unsafe-best-matchTest Plan
fsGroupis applied:kubectl exec <pod> -- ls -la /modelsshould show group1001/models/hub/has model filese2e_test/benchmarks/**Checklist
cargo +nightly fmtpasses (no Rust changes)cargo clippy --all-targets --all-features -- -D warningspasses (no Rust changes)Summary by CodeRabbit