Skip to content

[Diffusion] Optimize Qwen-Image TP collectives and attention - #36680

Merged
BBuf merged 19 commits into
sgl-project:mainfrom
BBuf:perf/diffusion-custom-allreduce-v2
Sep 1, 2026
Merged

BBuf merged 19 commits into
sgl-project:mainfrom
BBuf:perf/diffusion-custom-allreduce-v2

Conversation

@BBuf

@BBuf BBuf commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR consolidates the Qwen-Image TP kernel work previously split across
#36680, #36693, and #36694:

  • route CUDA diffusion TP collectives through the SRT custom-all-reduce
    dispatcher and admit 24 MiB Qwen-Image row-parallel outputs to
    CustomAllReduceV2
  • keep the three unquantized Qwen added-text Q/K/V weights packed, run
    three reference GEMMs for default quality=lossless, and mount the single
    column-parallel GEMM only for request-scoped quality=high; Nunchaku keeps
    its existing quantized packed path
  • pack masked text-prefix and image Q/K/V directly with a segmented Triton
    kernel instead of materializing three concatenated tensors
  • use the FA3 static persistent scheduler for batch-one dense packed sequences
    and scatter back to the padded/BCG bucket when needed
  • preserve the existing fallbacks for unsupported shapes, devices,
    quantization backends, and non-CUDA platforms
  • document the validated two-H200 recipe and update the diffusion performance
    skill/kernel documentation

Request quality semantics

Only the unquantized BF16 added-text QKV GEMM packing is non-bit-exact. It is
now a request-scoped quality site:

  • --quality=lossless (default): applies the packed Q/K/V weight slices as the
    original three independent projection GEMMs, preserving the reference
    reduction order and output;
  • --quality=high: mounts the single packed GEMM used for the performance
    numbers below;
  • switching back to lossless unmounts the site at the next batch boundary.

The checkpoint remains packed in memory in both modes, so this does not add a
second copy of the Q/K/V weights. The segmented pack, attention scheduler, and
collective changes keep their existing independently validated fallbacks.

Actual blast radius

This is broader than the original Qwen-only benchmark, but it does not
speed up every diffusion model or every Qwen checkpoint automatically.

Changed path Models/configurations that can execute it
CUDA diffusion TP custom-all-reduce dispatch + 32 MiB V2 workspace Any native CUDA diffusion run with TP > 1; checked-in affected presets include FLUX.1/2, Qwen-Image-2512, Qwen-Image-Edit-2511, Z-Image-Turbo, Ideogram 4 FP8, Cosmos3-Super T2V/I2V, MiniMax-H3, and the 4-GPU Cosmos TP variants
Generic masked-varlen FA3 dense/static fast path Native FA attention with a 2-D mask plus varlen metadata: ERNIE-Image, Krea-2, Ideogram 4, LingBot Video MoE, Z-Image, and Qwen-Image families
Quality-gated fused added-text QKV + segmented text/image packing The fused added-QKV GEMM is used only by quality=high on unquantized Qwen-Image architecture checkpoints: Qwen-Image, 2512, Edit, Edit-2509, Edit-2511, Layered, and FireRed Image Edit 1.0/1.1. Default lossless uses three projection GEMMs; segmented packing remains independently guarded.

LTX-2/2.3 and single-GPU FLUX.2 Klein do not execute these changed branches in
their tested presets; their measurements below are controls, not PR speedup
claims. Qwen/Ideogram NVFP4 are Blackwell-only and the quantized projection path
does not inherit the unquantized fused-QKV change.

Motivation and profiler evidence

The original Qwen-Image-2512 TP2 path left three independent gaps relative to
vLLM-Omni:

  1. Diffusion directly instantiated the legacy custom all-reduce
    implementation. The 24 MiB row-parallel outputs therefore fell back to NCCL
    instead of using the CUDA V2 dispatcher.
  2. Each attention block materialized joint text/image Q, K, and V tensors
    before packing. The trace contained 1,212 CatArray launches taking
    10.768 ms in the sampled window.
  3. SGLang selected FA3 VarlenDynamic at 159.403 us/call, while vLLM selected
    StaticPersistent at 147.402 us/call for this batch-one dense sequence.

The combined path changes those hot spots as follows:

  • 24 MiB collective microbenchmark: NCCL 104.196 us -> V2 graph one-shot pull
    81.754 us (21.5% faster)
  • materialized QKV pack -> segmented pack: 47.7293 us -> 21.0982 us at the
    production batch-one shape
  • FA3 attention: VarlenDynamic 159.403 us -> StaticPersistent 147.063 us,
    matching the vLLM kernel timing
  • under quality=high, fused added-text QKV removes 1,212 small split-K
    GEMMs/reductions and saves about 0.43 ms per denoise step

Cross-model H200 audit

Method

  • Base: a7e3f590cabb676a02181fb3328026587822f934
  • PR: 8e056cfbfe68ddc46c1f869b7853ec9fc7e94cfe
  • Same physical NVIDIA H200 pair (GPUs 6/7)
  • ABBA order: base-a, PR-a, PR-b, base-b; values below are two-run medians
  • Native SGLang backend, fixed prompt/seed, checked-in model preset. These
    measurements were collected before the request gate was added, when the packed
    GEMM ran unconditionally; they therefore describe the current quality=high
    path. The new default quality=lossless path retains the reference three-GEMM
    added-QKV computation. Eager with compile/BCG off unless explicitly shown
  • Every checkpoint used a task-owned cache. It was deleted and verified empty
    immediately after that model, including failed 403 attempts.

Full JSON summaries, output metrics, all contact sheets, and both cleanup
ledgers are in the
evidence branch.

Successfully executed affected checkpoints

Positive percentages mean lower PR latency. Bold rows cross the audit's 1.5%
signal threshold.

Preset / checkpoint Topology Base denoise (s) PR denoise (s) Change Output result
Qwen-Image-2512 TP2 + BCG 8.5406 7.8657 +8.58% repeats stable; Base/PR SSIM 0.9713
Qwen-Image 1 GPU + BCG 12.6528 12.3580 +2.39% bit-exact
Qwen-Image-Edit 1 GPU eager 32.0573 32.2576 -0.62% bit-exact
Qwen-Image-Edit-2509 1 GPU eager 22.3352 22.3464 -0.05% bit-exact
Qwen-Image-Edit-2511 TP2 eager 14.2704 14.2651 +0.04% bit-exact
Qwen-Image-Layered 1 GPU eager 32.9077 32.8971 +0.03% repeats stable; min SSIM 0.9891 over 4 layers
FireRed-Image-Edit-1.0 CFG2 eager 11.4178 11.4694 -0.45% bit-exact
FireRed-Image-Edit-1.1 CFG2 eager 11.4151 11.4170 -0.02% bit-exact
Z-Image-Turbo TP2 eager 0.5575 0.5307 +5.05% bit-exact
Z-Image 1 GPU eager 9.2629 9.3055 -0.46% bit-exact
Cosmos3-Super T2V TP2 eager 113.2269 113.0971 +0.11% MP4 and sampled frames bit-exact
Cosmos3-Super I2V TP2 eager 112.9026 112.9645 -0.05% MP4 and sampled frames bit-exact
ERNIE-Image 1 GPU eager 13.0868 12.8640 +1.73% prompt enhancement nondeterministic on both sides
ERNIE-Image-Turbo 1 GPU eager 12.9039 13.0717 -1.28% prompt enhancement nondeterministic on both sides
LingBot Video MoE 30B 1 GPU eager 3.9435 3.9126 +0.79% MP4 and sampled frames bit-exact

Qwen conclusion: no, the entire Qwen family does not get a measurable
speedup. The two T2I BCG lanes are clear wins; Edit, Edit-2509, Edit-2511,
Layered, and FireRed 1.0/1.1 are neutral or slightly slower at these fixed
shapes. The Qwen-specific code is shared, but its saved launches are too small
relative to the full denoiser on several variants, and only TP runs benefit
from the custom-all-reduce change.

ERNIE's internal prompt-enhancement generation is nondeterministic across fresh
processes on both base and PR. Its pictures are included for disclosure, but
their cross-side SSIM is not evidence of a PR correctness change. Denoise timing
is measured after the independently generated prompt.

Affected but blocked on this assignment

Preset Status
FLUX.1-dev TP2 Hugging Face 403 gated repo; isolated metadata cache cleaned
FLUX.2-dev TP2 Hugging Face 403 gated repo; isolated metadata cache cleaned
Ideogram 4 FP8 TP2 Hugging Face 403 gated repo; isolated metadata cache cleaned
Ideogram 4 Fast / Instant public wrappers resolve to gated Ideogram weights (403); caches cleaned
Krea-2 registry root Registered placeholder returns Hugging Face 404; metadata cache cleaned
Krea-2 Turbo / Raw Hugging Face 403 gated repos; caches cleaned
MiniMax-H3 TP2+Ulysses2 checked-in preset requires 4 GPUs; assignment had 2
Cosmos3-Super T2V CFG2+TP2 preset requires 4 GPUs; assignment had 2
Cosmos3-Super distilled T2I TP4 preset requires 4 GPUs; assignment had 2
Qwen-Image NVFP4 / Ideogram 4 NVFP4 Blackwell-only; cannot execute on H200

Negative controls (not attributed to this PR)

Preset Base (s) PR (s) Observed change Why it is a control
FLUX.2 Klein 4B 0.2390 0.2332 +2.52% 1 GPU, no masked-varlen metadata
FLUX.2 Klein Base 4B 6.0333 5.9889 +0.74% 1 GPU, no masked-varlen metadata
LTX-2 9.0477 9.4805 -4.57% CFG2, not TP; no varlen metadata
LTX-2.3 two-stage 13.5873 12.7691 +6.41% CFG2, not TP; no varlen metadata

These controls show why an isolated ABBA shift is not generalized unless the
changed branch is proven to execute.

Generated output: Base vs PR

Every successfully executed affected checkpoint has a Base/PR sheet below.
Video sheets show the first and middle frames; Layered shows all four layers.

Qwen family (8 checkpoints)

Qwen-Image-2512

Qwen-Image-2512 Base vs PR

Qwen-Image

Qwen-Image Base vs PR

Qwen-Image-Edit

Qwen-Image-Edit Base vs PR

Qwen-Image-Edit-2509

Qwen-Image-Edit-2509 Base vs PR

Qwen-Image-Edit-2511

Qwen-Image-Edit-2511 Base vs PR

Qwen-Image-Layered

Qwen-Image-Layered Base vs PR

FireRed-Image-Edit-1.0

FireRed-Image-Edit-1.0 Base vs PR

FireRed-Image-Edit-1.1

FireRed-Image-Edit-1.1 Base vs PR
Z-Image, Cosmos3, ERNIE, and LingBot (7 checkpoints)

Z-Image-Turbo

Z-Image-Turbo Base vs PR

Z-Image

Z-Image Base vs PR

Cosmos3-Super T2V

Cosmos3-Super T2V Base vs PR

Cosmos3-Super I2V

Cosmos3-Super I2V Base vs PR

ERNIE-Image (nondeterministic prompt enhancement on both sides)

ERNIE-Image Base vs PR

ERNIE-Image-Turbo (nondeterministic prompt enhancement on both sides)

ERNIE-Image-Turbo Base vs PR

LingBot Video MoE 30B

LingBot Video MoE Base vs PR

The control sheets are retained in the evidence branch but intentionally not
presented as affected-model results.

Separate vLLM-Omni comparison lane

This is the original fixed no-CFG benchmark and is separate from the cross-model
ABBA sweep above. It used Qwen-Image-2512, 1024x1024, 50 steps, guidance scale
1, seed 42, TP2, BCG enabled, compile off, five measured requests, and the same
adjacent full-NVLink H200 pair.

Runtime Median client latency Median denoise step
SGLang, combined candidate 4.101 s 76.9 ms
SGLang, CARv2-only baseline 4.256 s 78.6 ms
fastest observed vLLM-Omni regional-compile path 4.442 s -
vLLM-Omni forced eager 7.794 s -

The combined SGLang candidate is 7.7% faster than the fastest observed
vLLM-Omni result under that fixed workload. The fused added-QKV part of this
candidate now requires quality=high; default lossless keeps the reference
projection order.

Correctness and cleanup

  • Deterministic checkpoints were repeat-stable. Eleven affected lanes were
    bit-exact Base vs PR. Qwen-Image-2512 (SSIM 0.9713 / PSNR 30.15 dB) and
    Qwen-Image-Layered (minimum SSIM 0.9891 across four layers) were not bit-exact,
    but both were stable within each side.
  • The quality=high BF16 Qwen added-text projection fusion changes reduction
    association. Default quality=lossless now runs the reference three GEMMs. In
    the original no-CFG high-path lane, the fixed-seed image measured PSNR 40.55 dB and
    pixel MAE 1.393/255 against the unfused baseline.
  • ERNIE base and Turbo were nondeterministic within both base and PR because
    prompt enhancement is regenerated in each fresh process; the output sheets
    are disclosure rather than a correctness attribution.
  • The two cleanup ledgers contain 29 records (successful models, 403/404 attempts,
    and LTX-2.3's extra materialized checkpoint). Every record reports zero
    remaining bytes/weight files. The task cache itself is now empty except for
    its marker file.

Test plan

  • pre-commit run --from-ref origin/main --to-ref HEAD
  • Qwen-Image fused-QKV and BCG capture tests on H200: 12 passed, 3 subtests passed
  • request-scoped added-QKV quality gate on latest PR head: lossless/high/unmount switching, 14 import/Qwen tests passed on GB300
  • segmented pack/layout tests on H200: 16 passed
  • segmented pack benchmark on H200: 47.7293 us -> 21.0982 us (batch 1), 17.8346 us -> 7.7642 us (batch 2)
  • Qwen-Image-2512 TP2 BCG model-backed benchmark, five measured requests
  • same-pair vLLM-Omni comparison
  • affected-model H200 ABBA sweep, generated-media comparison, and per-model cleanup

The BF16 reduction-order tradeoff is now isolated behind quality=high and
documented explicitly. Supersedes #36693 and #36694.


CI States

Latest PR Test (Base): ⏳ Run #33459254730
Latest PR Test (Extra): ⏳ Run #33459254596
Latest PR Test (AMD ROCm 7.2): ⏳ Run #33459254715

@github-actions github-actions Bot added documentation Improvements or additions to documentation diffusion SGLang Diffusion labels Aug 27, 2026
@BBuf
BBuf force-pushed the perf/diffusion-custom-allreduce-v2 branch from d46d461 to 2a37870 Compare August 27, 2026 12:34
@BBuf BBuf changed the title perf(diffusion): use CustomAllReduceV2 for TP collectives [Diffusion] Optimize Qwen-Image TP collectives and attention Aug 27, 2026
@BBuf
BBuf marked this pull request as ready for review August 27, 2026 13:07
@BBuf BBuf added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) labels Aug 27, 2026
@BBuf
BBuf marked this pull request as draft August 27, 2026 18:27
@BBuf
BBuf marked this pull request as ready for review August 28, 2026 01:22
BBuf added 2 commits August 31, 2026 14:07
Preserve the PR's Qwen-Image NVFP4 fallback mapping while adopting main's transformer override quantization detection.
@BBuf

BBuf commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator Author

Updated the H100 consistency-data pin in 1ed919b73d to 883cf11, the merge commit of ci-data-diffusion#6.

This revision keeps the Qwen-Image fused-QKV GT from ci-data-diffusion#5, while restoring the three Sana H100 frames to the main-compatible blobs used before 88b17d4.

The Sana mismatch was unrelated to this PR's Qwen-Image runtime changes. 88b17d4 had captured output from the unmerged auto-residency experiment in #35335 / #36703, where post-warmup component placement changes were not numerically invariant for the two-stage Sana pipeline. That experiment later set SanaWMPipelineConfig.supports_auto_residency = False and moved its own ci-data pin back to the pre-88b17d4 revision, but the experimental images remained in the shared ci-data history and were picked up by this PR's repository-wide revision bump.

I also merged the latest main in b385ffe72a. The only content conflict was in transformer_load_utils.py; the resolution keeps this PR's quant_param_names_mapping for Qwen-Image NVFP4 fallback discovery and preserves main's new replacement-checkpoint quantization detection.

Validation performed locally:

  • the three Sana blob IDs in 883cf11 exactly match the pre-88b17d4 blobs;
  • the Qwen-Image fused-QKV GT remains unchanged;
  • git diff --check, Python compilation, and pre-commit on the touched/conflict-resolved files passed.

The updated head is b385ffe72a; CI should now compare both Qwen-Image and Sana against the correct references.

@BBuf

BBuf commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator Author

Applied the quality-gating follow-up in e5b6cdcbfd: the unquantized BF16 added-text QKV packed GEMM is now mounted only for request-scoped quality=high. Default quality=lossless applies the three packed weight slices as three independent reference GEMMs, without duplicating weights; switching back to lossless unmounts the site. On GB300, the lossless/high/unmount path test passed, the full quality-site suite passed (23 tests), and the latest-head import/Qwen selection passed (14 tests). The historical performance/SSIM numbers in the PR body are now explicitly labeled as the high path.

@BBuf

BBuf commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@BBuf

BBuf commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator Author

@BBuf
BBuf merged commit 71cee04 into sgl-project:main Sep 1, 2026
89 of 181 checks passed
StevenChenSE pushed a commit to StevenChenSE/sglang that referenced this pull request Sep 6, 2026
…ject#36680)

Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Rockdu added a commit to Rockdu/sglang that referenced this pull request Sep 22, 2026
…-wrapped layers

sgl-project#36680 added _split_unquantized_merged_linear, which reads
output_partition_sizes and weight straight off to_added_qkv. Once LoRA wraps
the layer (MergedColumnParallelLinearWithLoRA) the attribute is missing, and
splitting the base weight would drop the LoRA delta anyway. That broke
miles_diffusion test_qwenimage_pickscore_grpo_5xGPU on sglang-miles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one
MergedColumnParallelLinear (to_added_qkv). The parity patch routed every
ColumnParallelLinearWithLoRA through the 2D PEFT delta, which fails on
the merged layer's 3D stacked LoRA (size 8 vs 3072). Split the packed
weight into three reference GEMMs, each with its own PEFT-ordered delta,
which matches the separate linears PEFT trains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
…ity patch

sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one
MergedColumnParallelLinear (to_added_qkv). The parity patch forced every
LoRA layer through a 2D PEFT-ordered delta, which crashes on the merged
layer's 3D stacked LoRA (size 8 vs 3072). Let sgl-d's native LoRA
forwards run; the norm/RoPE/split_seqs parity patches stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
…ity patch

sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one
MergedColumnParallelLinear (to_added_qkv). The parity patch forced every
LoRA layer through a 2D PEFT-ordered delta, which crashes on the merged
layer's 3D stacked LoRA (size 8 vs 3072). Let sgl-d's native LoRA
forwards run; the norm/RoPE/split_seqs parity patches stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution, split by source:
- Step 0 (first rollout and the first train step): only this PR's
  7d83ef7, which drops the LoRA forward overrides from the qwen_image
  parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks)
  with this branch already gives this standard's step-0 values bit for
  bit.
- Everything after the first weight update: sgl-project/sglang#36680
  (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the
  only contributing commit in a CI-runner bisect from 5375babbac to
  5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into
  one to_added_qkv, so once LoRA is live the added-text Q/K/V run as
  one packed GEMM plus a stacked LoRA delta instead of three GEMMs.

Effect: the train<->rollout residual grows slightly
(log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series
barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
…ity patch

sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one
MergedColumnParallelLinear (to_added_qkv). The parity patch forced every
LoRA layer through a 2D PEFT-ordered delta, which crashes on the merged
layer's 3D stacked LoRA (size 8 vs 3072). Let sgl-d's native LoRA
forwards run; the norm/RoPE/split_seqs parity patches stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution, split by source:
- Step 0 (first rollout and the first train step): only this PR's
  7d83ef7, which drops the LoRA forward overrides from the qwen_image
  parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks)
  with this branch already gives this standard's step-0 values bit for
  bit.
- Everything after the first weight update: sgl-project/sglang#36680
  (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the
  only contributing commit in a CI-runner bisect from 5375babbac to
  5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into
  one to_added_qkv, so once LoRA is live the added-text Q/K/V run as
  one packed GEMM plus a stacked LoRA delta instead of three GEMMs.

Effect: the train<->rollout residual grows slightly
(log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series
barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to Rockdu/sglang that referenced this pull request Sep 24, 2026
…-wrapped layers

sgl-project#36680 added _split_unquantized_merged_linear, which reads
output_partition_sizes and weight straight off to_added_qkv. Once LoRA wraps
the layer (MergedColumnParallelLinearWithLoRA) the attribute is missing, and
splitting the base weight would drop the LoRA delta anyway. That broke
miles_diffusion test_qwenimage_pickscore_grpo_5xGPU on sglang-miles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 24, 2026
Attribution, split by source:
- Step 0 (first rollout and the first train step): only this PR's
  7d83ef7, which drops the LoRA forward overrides from the qwen_image
  parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks)
  with this branch already gives this standard's step-0 values bit for
  bit.
- Everything after the first weight update: sgl-project/sglang#36680
  (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the
  only contributing commit in a CI-runner bisect from 5375babbac to
  5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into
  one to_added_qkv, so once LoRA is live the added-text Q/K/V run as
  one packed GEMM plus a stacked LoRA delta instead of three GEMMs.

Effect: the train<->rollout residual grows slightly
(log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series
barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion documentation Improvements or additions to documentation jit-kernel run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants