Skip to content

[diffusion] Preserve eager base dispatch for disabled and merged Linear LoRA - #42257

Open
Tokha233 wants to merge 2 commits into
sgl-project:mainfrom
Tokha233:fix/diffusion-lora-eager-base
Open

Tokha233 wants to merge 2 commits into
sgl-project:mainfrom
Tokha233:fix/diffusion-lora-eager-base

Conversation

@Tokha233

@Tokha233 Tokha233 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

BaseLayerWithLoRA keeps its disabled/merged pass-through eager and compiles only the dynamic adapter path, because layerwise offload can rebind weights. LinearWithLoRA instead decorates its entire forward with torch.compile: wrapping an ordinary nn.Linear changes the base dispatch even when no adapter is active or after it is disabled. Compilers may choose a different GEMM/bias fusion; this also defeats the shared class's offload precaution.

The AMD Qwen Image 2.1 LoRA tests in this run fail exact base-output restoration by 1.1920928955078125e-7. This PR addresses the concrete dispatch inconsistency; whether it resolves that AMD numerical failure still needs native CI confirmation.

Modifications

  • Mirror BaseLayerWithLoRA: directly invoke the original base layer when disabled or merged, and apply merged output offsets eagerly when present.
  • Keep compilation on the active dynamic LoRA path.
  • Add disabled/merged × absent/present offset coverage, exact output checks and repeated weight rebinds. A compiler-sensitive value witness verifies that the pass-through really uses eager dispatch (a Python assertion alone can graph-break and miss this regression).

Accuracy Tests

Isolated RTX 5090 D v2 host, Python 3.12 / PyTorch 2.13.0+cu130; upstream f6fcda8:

  • New regression file against unchanged production code: 4 failed / 2 passed; the four pass-through dispatch cases fail.
  • Candidate plus LoRA pipeline, merge cache, commit-as-base, fused LoRA compose and Qwen Image 2.1 CUDA regression files: 58 passed.
  • Qwen's real CUDA prefill/cache/graph and adapter restore checks pass without changing tolerances. Native AMD output restoration remains unverified.
python -m pytest -q \
  python/sglang/multimodal_gen/test/unit/test_lora_inference_mode.py \
  python/sglang/multimodal_gen/test/unit/test_lora_pipeline.py \
  python/sglang/multimodal_gen/test/unit/test_lora_commit_as_base.py \
  python/sglang/multimodal_gen/test/unit/test_lora_merge_cache.py \
  python/sglang/multimodal_gen/test/unit/test_fused_lora_compose.py \
  python/sglang/multimodal_gen/test/unit/test_qwen_image21_cuda.py

Speed Tests and Profiling

No throughput improvement claimed. Dynamic LoRA remains compiled; merged/disabled linear layers now use the base dispatch, matching the existing parallel-linear wrapper policy. A full-model AMD latency/accuracy comparison requires upstream hardware CI.

Checklist

  • Changed-file pre-commit checks pass.
  • Regression tests added and related CUDA tests pass.
  • Accuracy results and platform limits documented.
  • Follow SGLang code style.
  • Native AMD CI (requires maintainer authorization).

CI States

Latest PR Test (Base): ❌ Run #37665506619
Latest PR Test (Extra): ❌ Run #37665505945
Latest PR Test (AMD ROCm 10): ❌ Run #37665506527

…le dispatch

Signed-off-by: Tokha233 <61346912+Tokha233@users.noreply.github.com>
@Tokha233

Tokha233 commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Merged current upstream 0b635266 to resolve the conflict with the newly added request-level LoRA scale and FP8 storage paths. The upstream snapshot restoration and dynamic scale behavior are retained. The zero-scale bypass now also stays outside the compiled delta helper, so temporarily disabling a request adapter preserves the original base GEMM dispatch and layerwise parameter rebinding.

Validation on new head 5d32348a, RTX 5090 D v2, Torch 2.13.0+cu130: 25 passed across inference-mode, merge-cache and commit-as-base tests (CPU and CUDA cases). Four new cases cover runtime scale zero with both standard and replicated linears, including nonzero output offsets and a parameter rebind. Changed-file pre-commit and diff checks pass. These are dispatch/correctness regressions; no new inference-throughput claim.

Please run the maintainer-authorized CI against this updated head.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion lora

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant