Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
66b3d89
[diffusion] Add CUDA IPC refs for local scheduler tensor hops
niehen6174 Aug 21, 2026
c73c6b8
[diffusion] Keep local Req GPU tensors off the gloo pickle path
niehen6174 Aug 21, 2026
57dd164
[diffusion] Drive ComfyUI from native pipelines plus a checkpoint spec
niehen6174 Aug 21, 2026
9f8f1d2
[diffusion] Cache ComfyUI conditioning in a run-scoped worker session
niehen6174 Aug 21, 2026
bc89dc0
[diffusion] Collapse ComfyUI plugin executors onto shared adapters
niehen6174 Aug 21, 2026
fa70707
[diffusion] Add ComfyUI adapter/profile unit tests and plugin e2e tweaks
niehen6174 Aug 21, 2026
bc75547
[diffusion] Fold ComfyUI profile, session, and checkpoint spec into t…
niehen6174 Aug 22, 2026
f8cbe83
[diffusion] Keep ComfyUI multi-rank CUDA hops off gloo without touchi…
niehen6174 Aug 22, 2026
22670bf
remote test_req_cuda_broadcast and FakeScheduler test
niehen6174 Aug 22, 2026
c99e77f
[diffusion] docs: document ComfyUI native-pipeline architecture
niehen6174 Aug 22, 2026
0bb5977
[diffusion] Keep ComfyUI CUDA IPC retains until the peer has mapped them
niehen6174 Aug 22, 2026
fdc3c1d
[diffusion] Read Flux guidance from the ComfyUI pack path
niehen6174 Aug 22, 2026
cb24c7f
[diffusion] Set Qwen ComfyUI prompt seq lens for later sampler steps
niehen6174 Aug 22, 2026
374d0bf
[diffusion] Drop leftover ComfyUI pack helpers and pickle-era device …
niehen6174 Aug 22, 2026
248e6bc
fix:lint
niehen6174 Aug 22, 2026
ab6dac6
[diffusion] Move ComfyUI CUDA IPC helpers under runtime.distributed
niehen6174 Aug 23, 2026
a89d031
[diffusion] Sort ComfyUI CUDA IPC imports for isort
niehen6174 Aug 23, 2026
4cad8e7
Merge origin/main into feat/comfyui-plugin-refactor
niehen6174 Aug 31, 2026
3540631
Merge remote-tracking branch 'origin/main' into pr-35960
niehen6174 Sep 24, 2026
359b9a1
Merge remote-tracking branch 'origin/main' into pr-35960
niehen6174 Sep 26, 2026
020cca5
[diffusion] Set comfyui_mode on the component-loader stub
niehen6174 Sep 26, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -1648,6 +1648,7 @@
"pages": [
"docs/sglang-diffusion/api/cli",
"docs/sglang-diffusion/api/openai_api",
"docs/sglang-diffusion/comfyui",
"docs/sglang-diffusion/realtime_models",
"docs/sglang-diffusion/models_with_ar",
"docs/sglang-diffusion/models_with_pe",
Expand Down
48 changes: 48 additions & 0 deletions docs/docs/sglang-diffusion/comfyui.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
---
title: ComfyUI plugin
description: Use SGLang Diffusion from ComfyUI in server mode or as a per-step DiT backend.
---

The [ComfyUI SGLDiffusion plugin](https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/apps/ComfyUI_SGLDiffusion) has two modes.

**Server mode** talks to a standalone `sglang serve` process over HTTP. ComfyUI sends prompts and receives images or video. This is the same path as the [OpenAI-compatible API](/docs/sglang-diffusion/api/openai_api).

**Integrated mode** keeps ComfyUI's CLIP, VAE, and sampler loop. SGLang loads only the DiT and runs one forward per sampler step. The worker starts the native Flux / Qwen-Image / Z-Image pipeline under `--comfyui-mode`. A single-file ComfyUI `.safetensors` is loaded through a checkpoint spec. Each `apply_model` call is translated by a per-model adapter.

Integrated mode does not ship a separate `comfyui_*` pipeline class per model.

## Install the plugin

1. Install `sglang[diffusion]`. See [Installation](/docs/sglang-diffusion/installation).
2. Copy `python/sglang/multimodal_gen/apps/ComfyUI_SGLDiffusion` into ComfyUI's `custom_nodes/` directory.
3. Restart ComfyUI.

Example workflows live next to the plugin under `workflows/`.

## Integrated mode

1. Load the DiT with `SGLDiffusion UNET Loader`.
2. Set `num_gpus`, `tp_size`, `model_type`, or compile flags with `SGLDiffusion Options`.
3. Connect the loaded model to a standard ComfyUI sampler.

Supported integrated-mode families: Flux, Qwen-Image, Z-Image. Qwen-Image edit is experimental.

## How a sampler step reaches the worker

The plugin README has the [architecture diagram](https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/apps/ComfyUI_SGLDiffusion#architecture). The hop is:

1. A model adapter packs ComfyUI tensors into an SGLang `Req`.
2. Local ZMQ replaces CUDA tensors with IPC handles so latents stay on GPU.
3. Rank 0 materializes the handles. Multi-rank `--comfyui-mode` then detaches CUDA tensors and broadcasts them over NCCL. The general SP / CFG / TP path is still the original `broadcast_pyobj`.
4. The worker pipeline keeps `transformer` plus a pass-through scheduler. After the first step, conditioning stays in a worker session; later steps send latents and the timestep.

Integrated mode currently supports:

| Family | ComfyUI `model_type` | Native pipeline | Example checkpoints |
| --- | --- | --- | --- |
| Flux | `flux` | `FluxPipeline` | `FLUX.1-dev` |
| Z-Image | `lumina2` | `ZImagePipeline` | `Z-Image-Turbo` |
| Qwen-Image | `qwen_image` | `QwenImagePipeline` | `Qwen-Image`, `Qwen-Image-2512` |
| Qwen-Image edit | `qwen_image_edit` | `QwenImageEditPlusPipeline` | `Qwen-Image-Edit-2511` (experimental) |

To add a model, register a checkpoint spec under `runtime/loader/comfyui_checkpoints/` and a `ComfyUIModelAdapter`. Do not add another ComfyUI pipeline class.
1 change: 1 addition & 0 deletions docs/docs/sglang-diffusion/index.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ sglang serve --model-path Qwen/Qwen-Image --port 30010
- [Supported Models](/docs/sglang-diffusion/compatibility_matrix): browse supported model families, tasks, and public checkpoints
- [CLI](/docs/sglang-diffusion/api/cli): run one-off generation jobs or launch a persistent server
- [OpenAI-Compatible API](/docs/sglang-diffusion/api/openai_api): send image and video requests to the HTTP server
- [ComfyUI plugin](/docs/sglang-diffusion/comfyui): use SGLang from ComfyUI in server mode or as a per-step DiT backend
- [Performance Overview](/docs/sglang-diffusion/performance-optimization): choose speed, memory, parallelism, caching, and quality-tradeoff levers
- [Caching Acceleration](/docs/sglang-diffusion/caching-acceleration): use Cache-DiT, TeaCache, or Spectrum to reduce denoising cost
- [Quantization](/docs/sglang-diffusion/quantization): configure component checkpoint and causal KV-cache quantization
Expand Down
2 changes: 1 addition & 1 deletion python/sglang/multimodal_gen/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ SGLang diffusion features an end-to-end unified pipeline for accelerating diffus
SGLang Diffusion has the following features:
- Broad model support: Wan, FastWan, FLUX, Qwen-Image / Qwen-Image 2.1, LongCat-Image, Z-Image, Ideogram 4, Krea-2, Cosmos3, LTX-2/LTX-2.3/LTX-2.5, MiniMax-H3, FastH3, VDN-H3, LingBot Video MoE, LingBot World, SANA-Video/SANA-WM, JoyEcho, MOVA, GLM-Image, ERNIE-Image, Hunyuan3D, and more
- Fast inference speed: empowered by optimized `sgl-kernel` kernels, scheduler/runtime improvements, caching acceleration, and native diffusion hot-path optimizations
- Ease of use: OpenAI-compatible api, CLI, and python sdk support
- Ease of use: OpenAI-compatible api, CLI, python sdk, and a [ComfyUI plugin](apps/ComfyUI_SGLDiffusion/README.md)
- Multi-platform support:
- NVIDIA GPUs (H100, H200, A100, B200, 4090, 5090)
- AMD GPUs (MI300X, MI325X, MI355X)
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,13 @@
# ComfyUI SGLDiffusion Plugin

A ComfyUI plugin for integrating with SGLang Diffusion server, supporting image and video generation capabilities.
A ComfyUI plugin for SGLang Diffusion. Server mode talks to a standalone HTTP
server. Integrated mode keeps ComfyUI's CLIP / VAE / sampler loop and uses
SGLang only as a per-step DiT forward.

Integrated mode no longer ships dedicated `comfyui_*` pipelines. It starts the
native Flux / Qwen-Image / Z-Image pipeline under `--comfyui-mode`, loads a
single-file ComfyUI `.safetensors` through a checkpoint spec, and translates
each `apply_model` call through a small per-model adapter.

## Installation

Expand Down Expand Up @@ -29,11 +36,13 @@ Connect to a standalone SGLang Diffusion server.
4. **LoRA Support**: Use `SGLDiffusion Server Set LoRA` and `SGLDiffusion Server Unset LoRA`.

### Mode 2: Integrated Mode (Tight Integration)
Leverage SGLang's high-performance sampling directly within ComfyUI while using ComfyUI's front-end nodes (CLIP, VAE, etc.).
ComfyUI keeps CLIP, VAE, and the sampler loop. SGLang loads only the DiT and
runs one forward per sampler step (pass-through scheduler, no text encode /
decode on the worker).

1. **Load Model**: Use the `SGLDiffusion UNET Loader` node to load your diffusion model.
2. **Configure Options**: Use the `SGLDiffusion Options` node to set runtime parameters like `num_gpus`, `tp_size`, `model_type`, or `enable_torch_compile`.
3. **Sample**: Connect the loaded model to standard ComfyUI samplers. SGLang will handle the sampling process efficiently.
3. **Sample**: Connect the loaded model to standard ComfyUI samplers. Each step is packed by a model adapter and sent to the SGLang scheduler.
4. **LoRA Support**: Use the `SGLDiffusion LoRA Loader` for native LoRA integration.

## Adding a Model
Expand Down Expand Up @@ -85,4 +94,45 @@ To use these workflows:

## Current Implementation

This plugin provides a high-performance backend for diffusion models in ComfyUI. By leveraging SGLang's optimized kernels and parallelization techniques (Tensor Parallelism, TeaCache, etc.), it significantly accelerates the sampling process, especially for large models like FLUX.
SGLang's optimized kernels and parallelism (TP / SP, compile, cache) run on the
DiT only. Text encoding and VAE stay in ComfyUI.

## Architecture

```mermaid
flowchart LR
subgraph comfy [ComfyUI process]
CLIP[CLIP / text encode]
VAE[VAE]
SAMPLER[Sampler loop]
EXEC["SGLDiffusionExecutor"]
ADAPT["Model adapter<br/>pack / unpack"]
CLIP --> SAMPLER
SAMPLER --> EXEC --> ADAPT
end

ADAPT -->|CUDA IPC spill| R0

subgraph sgl [SGLang]
R0[Rank-0 scheduler]
R0 -->|comfyui_mode and multi-rank| NCCL["Detach CUDA tensors<br/>NCCL broadcast"]
R0 -->|otherwise| PYO["Original SP / CFG / TP<br/>broadcast_pyobj"]
NCCL --> PIPE
PYO --> PIPE
PIPE["Native pipeline<br/>--comfyui-mode"]
SPEC["Checkpoint spec<br/>single .safetensors"] --> PIPE
PIPE --> STAGE["Latent prep + session cache<br/>+ DenoisingStage"]
end

STAGE -->|noise_pred IPC| EXEC
STAGE --> VAE
```

Per sampler step:

1. The adapter turns ComfyUI `apply_model` tensors into an SGLang `Req`.
2. Local ZMQ pickle replaces CUDA tensors with IPC handles so latents stay on GPU.
3. Rank 0 materializes the handles. Multi-rank `--comfyui-mode` then detaches CUDA tensors and broadcasts them over NCCL; the general SP / CFG / TP path is still the original whole-list `broadcast_pyobj`.
4. The worker pipeline is the native model class with modules trimmed to `transformer` + pass-through scheduler. A single-file checkpoint goes through `comfyui_checkpoints`. After the first step, conditioning stays in a worker session; later steps send latents and the timestep.

Adding a model means a checkpoint spec plus a `ComfyUIModelAdapter`. There is no extra ComfyUI pipeline class.
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,13 @@
)
from .model_patcher import SGLDModelPatcher

_EXECUTOR_CLASSES = (
FluxExecutor,
ZImageExecutor,
QwenImageExecutor,
QwenImageEditExecutor,
)


class SGLDiffusionGenerator:
"""Generator for SGLang Diffusion models in ComfyUI."""
Expand All @@ -41,18 +48,15 @@ def __init__(self):
self.executor = None
self.last_options = None

self.pipeline_class_dict = {
"flux": "ComfyUIFluxPipeline",
"lumina2": "ComfyUIZImagePipeline", # zimage
"qwen_image": "ComfyUIQwenImagePipeline",
"qwen_image_edit": "ComfyUIQwenImageEditPipeline",
}
self.executor_class_dict = {
"flux": FluxExecutor,
"lumina2": ZImageExecutor,
"qwen_image": QwenImageExecutor,
"qwen_image_edit": QwenImageEditExecutor,
}
# Native pipelines, run under comfyui_mode as a DiT-only forward service.
self.pipeline_class_dict = {}
self.executor_class_dict = {}
for executor_cls in _EXECUTOR_CLASSES:
for model_type in executor_cls.adapter_cls.model_types:
self.executor_class_dict[model_type] = executor_cls
self.pipeline_class_dict[model_type] = (
executor_cls.adapter_cls.pipeline_class_name
)

def __del__(self):
self.close_generator()
Expand All @@ -66,14 +70,37 @@ def init_generator(
if kwargs is None:
kwargs = {}
# Set comfyui_mode for ComfyUI integration
kwargs = dict(kwargs)
kwargs["comfyui_mode"] = True
# ComfyUI already keeps CLIP/VAE in the parent process. Auto image
# policy otherwise sets dit_cpu_offload=True and every sampler step
# reloads the DiT from CPU.
kwargs.setdefault("dit_cpu_offload", False)
kwargs = self._server_args_kwargs(kwargs)
self.generator = DiffGenerator.from_pretrained(
model_path=model_path,
pipeline_class_name=pipeline_class_name,
**kwargs,
)
return self.generator

@staticmethod
def _server_args_kwargs(kwargs: dict) -> dict:
"""Drop plugin-only / stale flags that ServerArgs no longer accepts."""
import dataclasses

from sglang.multimodal_gen.runtime.server_args import ServerArgs

valid = {f.name for f in dataclasses.fields(ServerArgs)}
aliases = {"dp_degree": "dp_size", "cache_strategy": None, "model_type": None}
cleaned = {}
for key, value in kwargs.items():
dest = aliases.get(key, key)
if dest is None or dest not in valid:
continue
cleaned[dest] = value
return cleaned

def kill_generator(self):
"""Kill worker processes manually because generator shutdown cannot terminate them."""
current_pid = os.getpid()
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -3,15 +3,47 @@
Provides executor classes for different model types.
"""

from .adapter import ComfyUIModelAdapter, PackedForward, get_adapter_class
from .base import SGLDiffusionExecutor
from .flux import FluxExecutor
from .qwen_image import QwenImageEditExecutor, QwenImageExecutor
from .zimage import ZImageExecutor
from .flux import FluxAdapter, FluxExecutor
from .zimage import ZImageAdapter, ZImageExecutor

# Qwen adapters import ComfyUI (`comfy.ldm.common_dit`). Keep that optional so
# unit tests and SGLD-only paths can load Flux / Z-Image without ComfyUI.

__all__ = [
"ComfyUIModelAdapter",
"PackedForward",
"SGLDiffusionExecutor",
"FluxAdapter",
"FluxExecutor",
"ZImageAdapter",
"ZImageExecutor",
"QwenImageExecutor",
"QwenImageEditExecutor",
"get_adapter_class",
]


def __getattr__(name):
if name in {
"QwenImageExecutor",
"QwenImageEditExecutor",
"QwenImageAdapter",
"QwenImageEditAdapter",
}:
from .qwen_image import (
QwenImageAdapter,
QwenImageEditAdapter,
QwenImageEditExecutor,
QwenImageExecutor,
)

mapping = {
"QwenImageAdapter": QwenImageAdapter,
"QwenImageEditAdapter": QwenImageEditAdapter,
"QwenImageExecutor": QwenImageExecutor,
"QwenImageEditExecutor": QwenImageEditExecutor,
}
return mapping[name]
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
# SPDX-License-Identifier: Apache-2.0
"""Per-model adapters for the ComfyUI DiT-forward contract.

The shared executor owns request construction and the ZMQ round-trip.
Each adapter only translates between ComfyUI's ``apply_model`` tensors and
the fields SGLang's ``Req`` expects. Design the interface around H3 (nested
latents, structured payload), not Flux's three-tensor case.
"""

from __future__ import annotations

from dataclasses import dataclass, field
from typing import Any

import torch

_ADAPTERS: dict[str, type[ComfyUIModelAdapter]] = {}


@dataclass
class PackedForward:
"""One ComfyUI sampler step, already translated into SGLang tensors."""

latents: torch.Tensor
timesteps: torch.Tensor
prompt_embeds: list[torch.Tensor]
height: int
width: int
guidance_scale: float = 1.0
prompt_seq_lens: list[list[int]] | None = None
pooled_embeds: list[torch.Tensor] | None = None
extra_req: dict[str, Any] = field(default_factory=dict)
unpack_ctx: dict[str, Any] = field(default_factory=dict)


class ComfyUIModelAdapter:
"""Model-specific pack / unpack / fill_req for one ComfyUI DiT family."""

model_types: tuple[str, ...] = ()
pipeline_class_name: str = ""

def __init_subclass__(cls, **kwargs):
super().__init_subclass__(**kwargs)
for model_type in cls.model_types:
_ADAPTERS[model_type] = cls

def pack(
self, x: torch.Tensor, timestep: torch.Tensor, context, **kwargs
) -> PackedForward:
raise NotImplementedError

def unpack(
self, noise_pred: torch.Tensor, packed: PackedForward, x: torch.Tensor
) -> torch.Tensor:
return noise_pred.to(x.device)

def fill_req(self, req, packed: PackedForward) -> None:
req.latents = packed.latents
req.timesteps = packed.timesteps
req.prompt_embeds = packed.prompt_embeds
req.raw_latent_shape = torch.tensor(packed.latents.shape, dtype=torch.long)
req.do_classifier_free_guidance = False
if packed.prompt_seq_lens is not None:
req.prompt_seq_lens = packed.prompt_seq_lens
if packed.pooled_embeds is not None:
req.pooled_embeds = packed.pooled_embeds
for key, value in packed.extra_req.items():
setattr(req, key, value)


def get_adapter_class(model_type: str) -> type[ComfyUIModelAdapter]:
if model_type not in _ADAPTERS:
raise ValueError(
f"Unsupported ComfyUI model type {model_type!r}. "
f"Registered: {sorted(_ADAPTERS)}"
)
return _ADAPTERS[model_type]


def registered_model_types() -> list[str]:
return sorted(_ADAPTERS)
Loading
Loading