Skip to content

[CI] Avoid Hugging Face Hub API calls when the local model cache is complete - #43229

Merged
hnyls2002 merged 1 commit into
mainfrom
lsyin/ci-hf-offline
Oct 8, 2026
Merged

hnyls2002 merged 1 commit into
mainfrom
lsyin/ci-hf-offline

Conversation

@hnyls2002

@hnyls2002 hnyls2002 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator
  • _validate_sharded_model: derive the index name from shards named model.safetensors-0000N-of-0000M.safetensors (Qwen3.5) as model.safetensors.index.json; CI cache validation of these models always failed and fell through to Hub API calls.
  • find_local_repo_dir and the custom cache_dir lookup: resolve a branch or tag revision (e.g. the draft model's default "main") through refs/<revision>, so draft models hit the local snapshot instead of going online on every load.
  • download_weights_from_hf: when listing the repo on the Hub raises HfHubHTTPError (e.g. 429), pick the weight format from the local snapshot instead of failing startup.
  • CI tests: in-process Engine startups (CustomTestCase, CI only) run with HF_HUB_OFFLINE=1 when the model, tokenizer, and draft repos all validate in the local cache, sharing the per-model validation with popen_launch_server. LoRA startups stay online, as in popen_launch_server.
  • CI install: drop the per-job kernels download / kernels lock step; the Hub FA3 kernel is only used with SGLANG_USE_SGL_FA3_KERNEL=0.
  • Add a regression test for the shard index name in test/manual/test_weight_validation.py.

No behavior change outside CI, except that an explicit branch or tag revision now loads from the local cache when refs/<revision> is present.


CI States

Latest PR Test (Base): 🚫 Run #37856132853
Latest PR Test (Extra): ❌ Run #37856132358
Latest PR Test (AMD ROCm 10): 🚫 Run #37856132473

…nd branch revision cache lookup; fail open on hub listing errors; drop ci kernels lock
@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Oct 8, 2026
@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/rerun-test test_lora_qwen3_5_35b_a3b_logprob_diff.py test_eagle_dp_attention.py test_deepep_small.py test_vlm_tp4.py test_score_engine.py test_embedding_models.py test_qwen35_deterministic.py test_lora_qwen3_5_4b_logprob_diff.py

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

🚀 4-gpu-h100 (5 tests): ❌ View workflow run

cd test/ && python3 registered/lora/test_lora_qwen3_5_35b_a3b_logprob_diff.py
cd test/ && python3 registered/spec/eagle/test_eagle_dp_attention.py
cd test/ && python3 registered/ep/test_deepep_small.py
cd test/ && python3 registered/vlm/test_vlm_tp4.py
cd test/ && python3 registered/attention/test_qwen35_deterministic.py

🚀 1-gpu-5090 (2 tests): ✅ View workflow run

cd test/ && python3 registered/prefill_only/test_score_engine.py
cd test/ && python3 registered/prefill_only/test_embedding_models.py

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/lora/test_lora_qwen3_5_4b_logprob_diff.py

@hnyls2002
hnyls2002 merged commit 2e64740 into main Oct 8, 2026
220 of 297 checks passed
@hnyls2002
hnyls2002 deleted the lsyin/ci-hf-offline branch October 8, 2026 23:27
nvpohanh added a commit to thanhhao98/sglang that referenced this pull request Oct 9, 2026
Picks up sgl-project#43229, which fixes the unrelated HF Hub 429 CI failures (e.g. https://github.com/sgl-project/sglang/actions/runs/37738591702/job/113443229627)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant