Repository navigation
AutoWeightLoader support Sglang native models 1: demo - #28671
Conversation
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
| """Decorator to register an architecture-specific weight remap function. | ||
|
|
||
| The decorated function receives a model instance and returns a WeightsMapper | ||
| (or None if no remap is needed for this configuration). |
There was a problem hiding this comment.
Do we need it to return None? If no remap is needed, we'd better not use the decorator at all (to avoid confusion)
| # --------------------------------------------------------------------------- | ||
|
|
||
|
|
||
| @dataclass |
There was a problem hiding this comment.
Can we use msgspec.Struct instead (codebase rule)
b8zhong
left a comment
There was a problem hiding this comment.
Can we show the use of WeightRemapRegistry in this PR (maybe modifying another model that requires this)
Document PR1–PR9 coverage plan, post_load protocol, unsupported models, and link from auto_loader module docstring. Co-authored-by: Cursor <cursoragent@cursor.com>
Replace P2P/streaming framing with verify_complete asserting all expected runtime tensors were updated. Co-authored-by: Cursor <cursoragent@cursor.com>
- Use msgspec.Struct for StackedParamsDispatch; RemapRegistry requires WeightsMapper - Demo RemapRegistry on Llama v2 path with submodule stacked loaders - Parametrize stacked/remap unit tests with real Llama checkpoint names - Tighten v1/v2 equiv test: randomize init, compare all named_parameters - Drop SGLANG_ENABLE_WEIGHT_LOADER_V2 env comment Co-authored-by: Cursor <cursoragent@cursor.com>
Add registered SRTRunner tests for Qwen2 v1/v2 and transformers backend; restore manual get_model state_dict equiv test; drop test_auto_loader.py. Co-authored-by: Cursor <cursoragent@cursor.com>
Add wave-1 parallelism, sequencing for PR7–PR9, and infra conflict hotspots. Co-authored-by: Cursor <cursoragent@cursor.com>
Remove design.md from the PR; point auto_loader docstring at the migration tracker issue linked from RFC sgl-project#24703. Co-authored-by: Cursor <cursoragent@cursor.com>
|
/rerun-test test/manual/test_weight_loader_v2_equiv.py test/registered/model_loading/test_weight_loader_v2_e2e.py |
|
Results for 🚀 ⛔ |
…gl-project#28671) sgl-project#28671 refactored LlamaForCausalLM/Qwen2ForCausalLM.load_weights to dispatch to self._legacy_load_weights / self._load_weights_v2. The reward and classification wrappers (llama_reward, llama_classification, qwen2_rm, qwen2_classification) borrow the base loader via an unbound call `Base.load_weights(self, ...)` but inherit only nn.Module, so `self` has no `_legacy_load_weights` and weight load fails with AttributeError (default path, since v2 is opt-in). Call `_legacy_load_weights` directly, restoring the pre-sgl-project#28671 behavior. These wrapped models are not AutoWeightLoader v2 citizens, so legacy is the correct path.
|
i will look into supporting the wrappers methods as well. |
…sgl-project#28671) (sgl-project#31988) Co-authored-by: Alex Nails <alex.nails@radixark.ai>
…sgl-project#28671) (sgl-project#31988) Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Motivation
AutoWeightLoader supports sglang models, better structural separation and guarantee.
Following discussion #24703 (comment)
Design and migration plan (PR1–PR9): #31051
Modifications
Speed Tests and Profiling
Checklist
CI States
Latest PR Test (Base): ❌ Run #29841391934
Latest PR Test (Extra): ❌ Run #29841391548