Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
139 changes: 139 additions & 0 deletions docs/cookbook/diffusion/CircleStone/Anima.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,139 @@
---
title: Anima
description: "Deploy Anima Base v1.0 with SGLang Diffusion for anime and illustration generation, using native components and single- or multi-GPU execution."
---

import { DiffusionModelTags } from '/src/snippets/diffusion/model-tags.jsx';
import { Deployment } from '/src/snippets/_deployment.jsx';
import { config } from '/src/snippets/configs/CircleStone/anima.jsx';

<DiffusionModelTags tags={["image", "text-to-image", "anime + illustration", "2B transformer"]} />

## 1. Quick start

Install the diffusion dependencies on Linux with NVIDIA CUDA, then install this
integration from its source checkout:

```bash Installation
uv pip install "sglang[diffusion]" --prerelease=allow
uv pip install -e "python[diffusion]"
```

Use the official `circlestone-labs/Anima-Base-v1.0-Diffusers` checkpoint. SGLang reads
its `modular_model_index.json` directly; no checkpoint conversion or newer Diffusers
runtime is required.

Start with **Resident + Compiled** for repeated generation on RTX 5090, DGX Spark,
or H200. Select **Eager** for shorter startup or frequent shape changes. Spark's
128 GB is unified memory shared with the CPU, not dedicated VRAM. See the
[measurements and tradeoffs](#5-measured-tuning) below.

<Deployment config={config} />

## 2. Model capabilities

[Anima](https://huggingface.co/circlestone-labs/Anima) is CircleStone Labs and Comfy
Org's text-to-image model for anime and illustration. It combines a Cosmos Predict2
transformer with Qwen3 text encoding, a learned T5-token conditioner, and a Qwen-Image
VAE. Prompts can combine tags with natural-language descriptions.

This integration targets **Base v1.0 text-to-image**. It does not claim support for
the separate Aesthetic or Turbo checkpoints, single-file ComfyUI weights, or image
editing. The checkpoint uses the CircleStone Labs Non-Commercial License; review
the [official license](https://huggingface.co/circlestone-labs/Anima/blob/main/LICENSE.md)
before deployment.

## 3. Sampling

Defaults are 1024 x 1024, 30 steps, guidance scale 4, and an empty negative prompt.
Width and height should be divisible by 16. `max_sequence_length` defaults to 512
and accepts 1 through 4096. The text conditioner pads short sequences to 512 tokens;
these padding positions remain part of the transformer's cross-attention, matching
the official implementation.

For offline generation:

```bash Generate
sglang generate \
--model-path circlestone-labs/Anima-Base-v1.0-Diffusers \
--prompt "masterpiece, best quality, safe, watercolor landscape, a quiet seaside village at sunset" \
--seed 42 \
--save-output
```

Use the same seed, generator device, dimensions, scheduler settings, and precision
when comparing runtimes. Latents and scheduler updates remain FP32; the transformer,
text components, and VAE default to BF16. Different attention kernels can produce
small floating-point differences that accumulate during denoising.

## 4. Runtime features

The pipeline reuses SGLang's native Qwen3 encoder, Qwen-Image VAE, component loaders,
denoising loop, and residency management. Anima's transformer and text conditioner
are native modules, not wrappers around Diffusers models.

For multi-GPU execution, choose TP to shard transformer weights or Ulysses/Ring to
shard image tokens. The transformer has 16 heads, so `tp_size * ulysses_degree`
must divide 16. CFG parallelism additionally splits conditional and unconditional
denoising and requires guidance greater than 1. The total GPU count must match the
selected parallel topology.

CPU and layerwise offload trade memory for transfers. The additional
`text_conditioner` component accepts the same residency controls as other native
components. Cache-DiT and quantized attention are approximate optimizations; assess
image quality for your prompts before enabling them. See
[performance optimization](/docs/sglang-diffusion/performance-optimization) for the
shared controls. H200 functional checks cover TP, Ulysses, Ring, CFG parallelism,
encoder folding, parallel tiled VAE decode, layerwise offload, FlashAttention,
Torch SDPA, SageAttention, Cache-DiT, and breakable CUDA graphs. These checks are
not a quality guarantee for approximate optimizations. Combined TP and SP,
multi-node execution, non-NVIDIA devices, and third-party LoRAs remain unverified.

Breakable CUDA graphs reuse the conditioner's actual sequence length, without
additional text-bucket padding. Requests with uncaptured shapes fall back to eager
execution. Standard short prompts use the same 512-token conditioning shape.

## 5. Measured tuning

For repeated requests, compilation was faster than eager execution on all three
tested platforms. The picker keeps eager available because initial compilation
and recompilation for new shapes can take minutes and use additional CPU memory.
No CPU offload, quantized attention, or Cache-DiT is enabled in the recommended
recipes.

The following are warm, sequential HTTP request medians: 1024 x 1024, 30 steps,
CFG 4, one image, CPU generator, and PNG/base64 output. Startup is excluded.

| GPU | Recommended recipe | Original eager | Recommended | Peak allocated |
| --- | --- | --- | --- | --- |
| RTX 5090, 32 GB | cuDNN SDPA + compile, tiled VAE | 6.99 s | 6.10 s | 5.7 GiB |
| DGX Spark, 128 GB unified | Torch SDPA + compile, tiled VAE | 26.03 s | 17.41 s | 5.6 GiB |
| H200, 141 GB | FlashAttention + compile, untiled VAE | 3.22 s | 2.22 s | 10.2 GiB |

Measurements used five timed requests after warmup, PyTorch
2.13.0+cu130, and the native Anima implementation at `214c9a48bb1` (Spark used
`7f3f903f039`, with identical runtime code). Results are workload-specific, not
throughput-at-saturation measurements. The original eager baseline uses Torch
SDPA on 5090/Spark and FlashAttention on H200, with VAE tiling enabled. Peak
allocated memory excludes the CUDA context, reserved pool, and compilation CPU
memory; it is not total device usage.

On **two NVLink-connected H200s**, Auto selects CFG parallelism instead of TP:
the compiled, untiled recipe measured **1.28 s**. Untiled decoding trades about
4-5 GiB more allocated memory for lower latency and is verified at 1024 x 1024,
one output. The tiled recipes additionally passed repeated 512 x 512 and
1536 x 1536 requests, plus two outputs at 1024 x 1024. Use **Tiled** for that
verified scope. At 512 x 512, two-GPU CFG eager was faster than compiled;
more GPUs or compilation are not universally better.

The recommended recipes change floating-point kernels and, on H200, VAE tiling.
They are **not pixel-identical to eager**: across three fixed prompt/seed pairs,
the single-GPU recipes above measured PSNR 21.95-35.44 dB and SSIM 0.851-0.969
against their same-device eager outputs.
These measure output differences, not perceptual quality guarantees. Use eager
with the same backend and tiling settings when reproducing an eager reference.

Breakable CUDA graphs were near parity at 1024 x 1024 while consuming about
1.2 GiB more allocated GPU memory for a single captured shape. They are not the
default. Cache-DiT can accelerate further, but changes the denoising computation;
validate it separately against your quality requirements.
6 changes: 6 additions & 0 deletions docs/docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -1521,6 +1521,12 @@
"cookbook/diffusion/Qwen-Image/Qwen-Image-Edit"
]
},
{
"group": "CircleStone",
"pages": [
"cookbook/diffusion/CircleStone/Anima"
]
},
{
"group": "SenseNova",
"tag": "NEW",
Expand Down
15 changes: 9 additions & 6 deletions docs/src/snippets/_deployment.jsx
Original file line number Diff line number Diff line change
Expand Up @@ -1865,16 +1865,19 @@ export const Deployment = ({ config, benchmarks }) => {
&& Number(sel.nodes) === recommendedRecipe.nodes
&& Number(sel.gpus_per_node) === recommendedRecipe.gpus_per_node
&& sel.topology_mode === "auto"
&& ["auto", recommendedRecipe.placement].includes(sel.placement)
&& sel.attention === "platform"
&& sel.precision === "native"
&& ["auto", recommendedRecipe.encoder].includes(sel.encoder)
&& sel.execution === "eager";
&& serveDims.every((dim) => {
const expected = recommendedRecipe[dim.id] ?? dim.default;
if (dim.id === "attention") return sel.attention === "platform";
return expected === undefined || sel[dim.id] === expected
|| (["placement", "encoder"].includes(dim.id) && sel[dim.id] === "auto");
});

const restoreRecommendedRecipe = () => {
if (!recommendedRecipe) return;
setSel((prev) => reseatHiddenPicks(normalizeBuilderSelection({
...prev,
...Object.fromEntries(serveDims
.map((dim) => [dim.id, recommendedRecipe[dim.id] ?? dim.default ?? dim.options?.[0]?.id])),
nodes: recommendedRecipe.nodes,
gpus_per_node: recommendedRecipe.gpus_per_node,
topology_mode: "auto",
Expand All @@ -1885,7 +1888,7 @@ export const Deployment = ({ config, benchmarks }) => {
attention: "platform",
precision: "native",
encoder: recommendedRecipe.encoder || "auto",
execution: "eager",
execution: recommendedRecipe.execution || "eager",
})));
};

Expand Down
141 changes: 141 additions & 0 deletions docs/src/snippets/configs/CircleStone/anima.jsx
Original file line number Diff line number Diff line change
@@ -0,0 +1,141 @@
export const config = {
modelName: "Anima Base v1.0",
supportedHardware: ["rtx5090", "dgx-spark", "h200"],
hardware: [
{ id: "rtx5090", label: "RTX 5090", vram: "32GB", vendor: "consumer" },
{ id: "dgx-spark", label: "DGX Spark", vram: "128GB unified", vendor: "blackwell" },
],
groupHardware: false,
matchDims: [],
overlayDims: [
{
id: "weights", title: "Checkpoint", scope: "base", default: "default",
description: "Official self-contained Diffusers checkpoint.",
options: [{ id: "default", label: "Base v1.0" }],
},
{
id: "mode", title: "Request mode", scope: "base", default: "text",
options: [{ id: "text", label: "Text to image" }],
},
{
id: "placement", title: "Placement", scope: "serve", default: "resident",
description: "Keep the pipeline resident when it fits. Use offload for a smaller memory budget.",
learnMore: "#4-runtime-features",
options: [
{ id: "resident", label: "Resident", flags: ["--performance-mode speed"], recommended: true },
{ id: "offload", label: "Layerwise offload", flags: ["--performance-mode memory", "--dit-layerwise-offload true"], soft: true, softReason: "This exact HTTP recipe is unverified." },
],
},
{
id: "execution", title: "Execution", scope: "serve", default: "compile",
description: "Compile for repeated requests. Eager avoids compilation at startup and for new shapes.",
learnMore: "#5-measured-tuning",
options: [
{ id: "compile", label: "Compiled", flags: ["--enable-torch-compile true"], recommended: true },
{ id: "eager", label: "Eager", flags: [] },
],
},
{
id: "attention", title: "Attention", scope: "serve", default: "platform",
description: "Exact attention backends can differ in floating-point reduction order.",
options: [
{ id: "platform", label: "Platform default", flags: (s) => [`--attention-backend ${s.hw === "h200" ? "fa" : s.hw === "rtx5090" ? "torch_cudnn_sdpa" : "torch_sdpa"}`], recommended: true },
{ id: "fa", label: "FlashAttention", flags: ["--attention-backend fa"], soft: (s) => s.hw !== "h200", softReason: "On RTX 5090 and Spark, this selector falls back to Torch SDPA. Select Torch SDPA explicitly." },
{ id: "sdpa", label: "Torch SDPA", flags: ["--attention-backend torch_sdpa"] },
{ id: "cudnn", label: "cuDNN SDPA", flags: ["--attention-backend torch_cudnn_sdpa"] },
],
},
{
id: "cfg", title: "CFG parallelism", scope: "serve", default: "auto",
description: "Auto uses CFG parallelism on two H200 GPUs. Manual TP/SP overrides disable automatic CFG splitting.",
learnMore: "#5-measured-tuning",
options: [
{ id: "auto", label: "Auto", flags: [] },
{ id: "off", label: "Off", flags: [] },
{ id: "on", label: "On", flags: [], disabled: (s) => Number(s.gpus_per_node) % 2 !== 0, disableReason: "CFG parallelism requires an even number of GPUs and guidance greater than 1." },
],
},
{
id: "vae", title: "VAE decoding", scope: "serve", default: "auto",
description: "Auto disables tiling for compiled H200 recipes. Untiled decoding uses about 4-5 GiB more memory at 1024 x 1024.",
learnMore: "#5-measured-tuning",
options: [
{ id: "auto", label: "Auto", flags: [] },
{ id: "tiled", label: "Tiled", flags: [] },
{ id: "full", label: "Untiled", flags: [] },
],
},
{
id: "outputs", title: "Outputs", scope: "request", kind: "number",
default: 1, min: 1, max: 4, options: [],
},
],
commandBuilder: {
defaultSelection: { hw: "rtx5090", nodes: 1, gpus_per_node: 1, topology_mode: "auto", tp_size: 1, ulysses_degree: 1, ring_degree: 1 },
resource: {
limits: { nodes: { min: 1, max: 1 }, gpus_per_node: { min: 1, max: 8 } },
verifiedRecipes: [
{ id: "rtx5090-1gpu-resident-cudnn", hw: "rtx5090", nodes: 1, gpus_per_node: 1, tp_size: 1, ulysses_degree: 1, ring_degree: 1, placement: "resident", attention: "cudnn", execution: "compile", cfg: "auto", vae: "auto" },
{ id: "dgx-spark-1gpu-resident-sdpa", hw: "dgx-spark", nodes: 1, gpus_per_node: 1, tp_size: 1, ulysses_degree: 1, ring_degree: 1, placement: "resident", attention: "sdpa", execution: "compile", cfg: "auto", vae: "auto" },
{ id: "h200-1gpu-resident-fa", hw: "h200", nodes: 1, gpus_per_node: 1, tp_size: 1, ulysses_degree: 1, ring_degree: 1, placement: "resident", attention: "fa", execution: "compile", cfg: "auto", vae: "auto" },
{ id: "h200-2gpu-resident-cfg", hw: "h200", nodes: 1, gpus_per_node: 2, tp_size: 1, ulysses_degree: 1, ring_degree: 1, placement: "resident", attention: "fa", execution: "compile", cfg: "auto", vae: "auto" },
],
cfgDegree: (s) => s.cfg === "on" || (s.cfg === "auto" && s.topology_mode !== "manual" && s.hw === "h200" && Number(s.gpus_per_node) === 2) ? 2 : 1,
vaeTiling: (s) => s.vae === "tiled" || (s.vae === "auto" && (s.hw !== "h200" || s.execution !== "compile")),
autoTopology: (s) => ({ tp_size: Number(s.gpus_per_node) / config.commandBuilder.resource.cfgDegree(s), ulysses_degree: 1, ring_degree: 1 }),
validateTopology: (s, t) => {
const errors = [];
if (Number(s.nodes) !== 1) errors.push("This picker covers single-node deployment only.");
if (s.hw === "dgx-spark" && Number(s.gpus_per_node) !== 1) errors.push("DGX Spark has one GPU per node.");
const cfg = config.commandBuilder.resource.cfgDegree(s);
if (!Number.isInteger(t.tp_size) || t.tp_size < 1 || Number(s.nodes) * Number(s.gpus_per_node) !== cfg * t.tp_size * t.ulysses_degree * t.ring_degree) errors.push("GPU count must equal CFG * TP * Ulysses * Ring.");
if (16 % (t.tp_size * t.ulysses_degree)) errors.push("TP * Ulysses must divide Anima's 16 attention heads.");
if (t.ring_degree > 1 && (s.hw !== "h200" || ["sdpa", "cudnn"].includes(s.attention))) errors.push("Ring requires FlashAttention.");
return errors;
},
},
resolveDeployment: (s) => {
const r = config.commandBuilder.resource;
const attention = s.attention === "platform" ? (s.hw === "h200" ? "fa" : s.hw === "rtx5090" ? "cudnn" : "sdpa") : s.attention;
const t = s.topology_mode === "manual"
? { tp_size: Number(s.tp_size), ulysses_degree: Number(s.ulysses_degree), ring_degree: Number(s.ring_degree) }
: r.autoTopology(s);
const errors = r.validateTopology(s, t);
const cfg = r.cfgDegree(s);
const tiled = r.vaeTiling(s);
const world = Number(s.nodes) * Number(s.gpus_per_node);
const recipe = r.verifiedRecipes.find((v) => v.hw === s.hw && v.gpus_per_node === Number(s.gpus_per_node) && v.tp_size === t.tp_size && v.ulysses_degree === t.ulysses_degree && v.ring_degree === t.ring_degree && v.placement === s.placement && r.cfgDegree(v) === cfg);
const attentions = world > 1 ? ["fa"] : s.execution === "eager"
? ["sdpa", "cudnn", ...(s.hw === "h200" ? ["fa"] : [])]
: s.hw === "rtx5090" ? ["sdpa", "cudnn"] : [s.hw === "h200" ? "fa" : "sdpa"];
const multiOutput = tiled && (s.execution === "compile" || attention === "fa" || (attention === "sdpa" && s.hw !== "h200"));
const vaeVerified = tiled || (s.execution === "compile" && attention === (s.hw === "h200" ? "fa" : "sdpa"));
const verified = !!recipe && attentions.includes(attention) && vaeVerified && Number(s.outputs) <= (multiOutput ? 2 : 1) && errors.length === 0;
const flags = ["--model-path {{MODEL_NAME}}"];
if (world > 1) flags.push(`--num-gpus ${world}`, `--tp-size ${t.tp_size}`, `--ulysses-degree ${t.ulysses_degree}`, `--ring-degree ${t.ring_degree}`);
if (cfg > 1) flags.push("--enable-cfg-parallel");
if (!tiled) flags.push("--vae-tiling false", "--vae-sp false");
flags.push("--host {{HOST_IP}}", "--port {{PORT}}");
return {
match: { hw: s.hw }, nnodes: Number(s.nodes), verified, flags,
builder: {
topology: t, topologySummary: `CFG ${cfg} / TP ${t.tp_size} / Ulysses ${t.ulysses_degree} / Ring ${t.ring_degree}`,
errors, warnings: verified ? [] : ["This exact HTTP recipe has not been verified."],
verification: { serve: verified ? "verified" : "unverified", request: verified ? "verified" : "unverified" },
resolvedSettings: { attention: s.attention === "platform" ? ({ fa: "FlashAttention", sdpa: "Torch SDPA", cudnn: "cuDNN SDPA" }[attention]) : undefined, cfg: cfg > 1 ? "2-way CFG" : "Off", vae: tiled ? "Tiled" : "Untiled" },
},
};
},
},
modelNames: { default: "circlestone-labs/Anima-Base-v1.0-Diffusers" },
placeholders: {
HOST_IP: { target: "command", label: "Bind host", default: "0.0.0.0" },
PORT: { target: "command", label: "Bind port", default: "30000" },
CURL_HOST: { target: "curl", label: "Server host", default: "localhost" },
CURL_PORT: { target: "curl", label: "Server port", default: "30000" },
},
curl: (s) => `curl -sS http://{{CURL_HOST}}:{{CURL_PORT}}/v1/images/generations \\
-H 'Content-Type: application/json' \\
-d '${JSON.stringify({ model: "{{MODEL_NAME}}", prompt: "masterpiece, best quality, safe, watercolor landscape, a quiet seaside village at sunset", size: "1024x1024", n: Number(s.outputs), seed: 42, generator_device: "cpu", output_format: "png", response_format: "b64_json" }, null, 2)}'`,
cells: [],
};
5 changes: 5 additions & 0 deletions docs/src/snippets/diffusion/model-catalog.jsx
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,11 @@
export const DiffusionModelCatalog = ({ category }) => {
const MODEL_CATALOG = {
image: [
{
name: "Anima",
modelIds: ["circlestone-labs/Anima-Base-v1.0-Diffusers"],
cookbook: "/cookbook/diffusion/CircleStone/Anima",
},
{
name: "Ming-Image",
modelIds: ["inclusionAI/Ming-Image-0.1-Design", "inclusionAI/Ming-Image-0.1-Design-Layer"],
Expand Down
16 changes: 7 additions & 9 deletions python/sglang/cli/utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -43,15 +43,13 @@ def _is_diffusion_model_from_registry(model_path: str) -> bool:


def _is_diffusers_model_dir(model_dir: str) -> bool:
"""Check if a local directory contains a valid diffusers model_index.json."""
config_path = os.path.join(model_dir, "model_index.json")
if not os.path.exists(config_path):
return False

with open(config_path) as f:
config = json.load(f)

return "_diffusers_version" in config
"""Check for a standard or modular Diffusers pipeline index."""
for filename in ("model_index.json", "modular_model_index.json"):
config_path = os.path.join(model_dir, filename)
if os.path.isfile(config_path):
with open(config_path) as f:
return "_diffusers_version" in json.load(f)
return False


def _is_gated_diffusion_repo(repo_id: str) -> bool:
Expand Down
Loading
Loading