Repository navigation
Conversation
…le dispatch Signed-off-by: Tokha233 <61346912+Tokha233@users.noreply.github.com>
|
Merged current upstream Validation on new head Please run the maintainer-authorized CI against this updated head. |
Motivation
BaseLayerWithLoRAkeeps its disabled/merged pass-through eager and compiles only the dynamic adapter path, because layerwise offload can rebind weights.LinearWithLoRAinstead decorates its entireforwardwithtorch.compile: wrapping an ordinarynn.Linearchanges the base dispatch even when no adapter is active or after it is disabled. Compilers may choose a different GEMM/bias fusion; this also defeats the shared class's offload precaution.The AMD Qwen Image 2.1 LoRA tests in this run fail exact base-output restoration by 1.1920928955078125e-7. This PR addresses the concrete dispatch inconsistency; whether it resolves that AMD numerical failure still needs native CI confirmation.
Modifications
BaseLayerWithLoRA: directly invoke the original base layer when disabled or merged, and apply merged output offsets eagerly when present.Accuracy Tests
Isolated RTX 5090 D v2 host, Python 3.12 / PyTorch 2.13.0+cu130; upstream
f6fcda8:Speed Tests and Profiling
No throughput improvement claimed. Dynamic LoRA remains compiled; merged/disabled linear layers now use the base dispatch, matching the existing parallel-linear wrapper policy. A full-model AMD latency/accuracy comparison requires upstream hardware CI.
Checklist
CI States
Latest PR Test (Base): ❌ Run #37665506619
Latest PR Test (Extra): ❌ Run #37665505945
Latest PR Test (AMD ROCm 10): ❌ Run #37665506527