Repository navigation
[CI] Avoid Hugging Face Hub API calls when the local model cache is complete - #43229
Merged
Merged
Conversation
…nd branch revision cache lookup; fail open on hub listing errors; drop ci kernels lock
Collaborator
Author
|
/tag-and-rerun-ci |
Collaborator
Author
|
/rerun-test test_lora_qwen3_5_35b_a3b_logprob_diff.py test_eagle_dp_attention.py test_deepep_small.py test_vlm_tp4.py test_score_engine.py test_embedding_models.py test_qwen35_deterministic.py test_lora_qwen3_5_4b_logprob_diff.py |
Contributor
|
🚀 🚀 🚀 |
nvpohanh
added a commit
to thanhhao98/sglang
that referenced
this pull request
Oct 9, 2026
Picks up sgl-project#43229, which fixes the unrelated HF Hub 429 CI failures (e.g. https://github.com/sgl-project/sglang/actions/runs/37738591702/job/113443229627) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
_validate_sharded_model: derive the index name from shards namedmodel.safetensors-0000N-of-0000M.safetensors(Qwen3.5) asmodel.safetensors.index.json; CI cache validation of these models always failed and fell through to Hub API calls.find_local_repo_dirand the customcache_dirlookup: resolve a branch or tag revision (e.g. the draft model's default"main") throughrefs/<revision>, so draft models hit the local snapshot instead of going online on every load.download_weights_from_hf: when listing the repo on the Hub raisesHfHubHTTPError(e.g. 429), pick the weight format from the local snapshot instead of failing startup.Enginestartups (CustomTestCase, CI only) run withHF_HUB_OFFLINE=1when the model, tokenizer, and draft repos all validate in the local cache, sharing the per-model validation withpopen_launch_server. LoRA startups stay online, as inpopen_launch_server.kernels download/kernels lockstep; the Hub FA3 kernel is only used withSGLANG_USE_SGL_FA3_KERNEL=0.test/manual/test_weight_validation.py.No behavior change outside CI, except that an explicit branch or tag
revisionnow loads from the local cache whenrefs/<revision>is present.CI States
Latest PR Test (Base): 🚫 Run #37856132853
Latest PR Test (Extra): ❌ Run #37856132358
Latest PR Test (AMD ROCm 10): 🚫 Run #37856132473