Skip to content

[diffusion] Fix Z-Image accuracy - #29742

Merged
mickqian merged 18 commits into
sgl-project:mainfrom
qimcis:fix-zimage-dynamic-batch-qwen3-rope
Jul 8, 2026
Merged

mickqian merged 18 commits into
sgl-project:mainfrom
qimcis:fix-zimage-dynamic-batch-qwen3-rope

Conversation

@qimcis

@qimcis qimcis commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Fix #28502 , and general accuracy issues with Z-image-Turbo

Modifications

With --batching-mode dynamic, Tongyi-MAI/Z-Image-Turbo could produce severely degraded images when requests were batched.

Initially, i found part of this issue was that qwen3 text encoder default position_ids were shaped [1, seq] even when hidden_states were [batch, seq, dim], causing RoPE to be applied with the wrong flattened token layout for batched prompts,

But following this, I realize that our current z-image turbo implementation is severely degraded, even for singleton generations, compared to the native pytorch implementation:

Firstly, z-image sampling and normalization differed from the native implementation above; we use the scheduler default sigma path, outer autocast, and shared fp32-accumulating RMSNorm, changing the denoising trajectory and bf16 activations for this model

Secondly, we padded batched images/captions to shared batch len but did not preserve RoPE offsets for each request or mask extra batch-only padding in attention, which causes mixed-prompt dynamic batches to attend to invalid tokens and use wrong image positions

Accuracy Tests

Tested on 1xh100

python3 -m sglang.multimodal_gen.runtime.entrypoints.cli.main serve \
  --model-path Tongyi-MAI/Z-Image-Turbo \
  --host 0.0.0.0 \
  --port 30000 \
  --batching-mode dynamic \
  --batching-max-size 5 \
  --batching-delay-ms 1000 \
  --enable-batching-metrics
{
  "model": "Tongyi-MAI/Z-Image-Turbo",
  "prompt": "Please produce a black-and-white line art drawing.\n\nSpecifications: Render the artwork exclusively with black lines, with distinct, well-defined outlines. The piece shall consist solely of black linework on a white background, with no fills, no gradients, and no shading. Emphasize clarity of outlines and structural forms. All lines must be sharp, continuous, and with clean edges; the overall image shall be neat and uncluttered.\n\nThe subject to be depicted is: <PROMPT>",
  "n": 1,
  "size": "640x480",
  "response_format": "b64_json",
  "seed": 42,
  "num_inference_steps": 8,
  "guidance_scale": 0.0
}

Prompts tested:

  • Draw Donald Duck and Mickey Mouse fishing by the river.
  • Crayon Shin chan walks with a puppy on the street.
  • Naruto and Luffy sit side by side on a large boulder. On the left, Naruto has a smile on his face and raises one hand to make a victory gesture. On the right, Luffy wears a wide-brimmed straw hat, an unbuttoned short-sleeved jacket, shorts and flip-flops, with a long sword strapped to his back. He is also grinning happily. Behind them lie a wide expanse of water and the sky dotted with a few clouds, and patches of grass grow beside the boulder.
  • SpongeBob SquarePants and Patrick sat at the table having dinner together, and there was a TV in the middle of the house with an old phone next to it.
  • Mario is driving a go kart on the highway, surrounded by many low trees.
all_prompts_native_main_branch_singleton_batch5_no_prompt_column

Speed Tests and Profiling

1xh100, 1024x1024, 9 steps single generation

Loop Time
main 863.27 ms
current branch 851.57 ms

batch generation size 5:

Batch Wall Mean Client Denoising Avg Step
main 4907.5 ms 4893.5 ms 3749.0 ms 416.3 ms
current branch 4781.3 ms 4765.0 ms 3691.5 ms 409.9 ms

about 2% faster across single and batch generation

Checklist


CI States

Latest PR Test (Base): ❌ Run #28860860153
Latest PR Test (Extra): ❌ Run #28860859883

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for batched inference, attention masking, and custom RMS normalization (ZImageRMSNorm) to match the official Z-Image implementation. It also fixes a batch-dimension issue in the Qwen3 encoder's default position IDs and adds corresponding unit tests. The reviewer suggests caching the batched freqs_cis in the Z-Image model's forward pass to avoid redundant computations across denoising steps, which would improve serving throughput.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +1155 to +1163
if len(input_images) > 1 and get_sp_world_size() == 1:
freqs_cis = self._build_batched_freqs_cis(
input_images,
input_cap_feats,
patch_size,
f_patch_size,
image_target_len=x.shape[1],
cap_target_len=cap_feats.shape[1],
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

In the current implementation, _build_batched_freqs_cis is called on every single denoising step when batching is enabled (len(input_images) > 1). Since the shapes and devices of the input images and caption features do not change across denoising steps within a request, rebuilding the batched freqs_cis on every step introduces redundant overhead (such as coordinate grid creation, rotary embedding computation, padding, and stacking).

We can cache the batched freqs_cis based on the input shapes, devices, and patch parameters to completely avoid this redundant computation and improve serving throughput.

        if len(input_images) > 1 and get_sp_world_size() == 1:
            cache_key = (
                len(input_images),
                tuple(img.shape for img in input_images),
                tuple(cap.shape for cap in input_cap_feats),
                patch_size,
                f_patch_size,
                x.shape[1],
                cap_feats.shape[1],
                device,
            )
            if (
                getattr(self, "_cached_batched_freqs_cis_key", None) == cache_key
                and getattr(self, "_cached_batched_freqs_cis", None) is not None
            ):
                freqs_cis = self._cached_batched_freqs_cis
            else:
                freqs_cis = self._build_batched_freqs_cis(
                    input_images,
                    input_cap_feats,
                    patch_size,
                    f_patch_size,
                    image_target_len=x.shape[1],
                    cap_target_len=cap_feats.shape[1],
                )
                self._cached_batched_freqs_cis_key = cache_key
                self._cached_batched_freqs_cis = freqs_cis

@qimcis

qimcis commented Jun 30, 2026

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Jun 30, 2026

@mickqian mickqian left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

excellent. could you help confirm:

  1. this kind of problem only happens with Z-Image, and,
  2. does this ground truth needs to be updated, or we just need to tighten the consistency threshold? repro script is here


image_pos = torch.arange(image_target_len, device=device).unsqueeze(0)
cap_pos = torch.arange(cap_target_len, device=device).unsqueeze(0)
image_len = torch.tensor(image_lengths, device=device).unsqueeze(1)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: torch.as_tensor

@qimcis

qimcis commented Jul 1, 2026

Copy link
Copy Markdown
Contributor Author

/rerun-failed-ci

@qimcis

qimcis commented Jul 1, 2026

Copy link
Copy Markdown
Contributor Author

excellent. could you help confirm:

  1. this kind of problem only happens with Z-Image, and,
  2. does this ground truth needs to be updated, or we just need to tighten the consistency threshold? repro script is here

i tested across some other models, and didn't seem to see the same problem - gt also seems to be okay and i don't think it needs to be updated, let's see if ci passes

@qimcis
qimcis force-pushed the fix-zimage-dynamic-batch-qwen3-rope branch from 92f2b43 to 34cd67e Compare July 2, 2026 19:11
@qimcis

qimcis commented Jul 2, 2026

Copy link
Copy Markdown
Contributor Author

/rerun-failed-ci

@qimcis
qimcis force-pushed the fix-zimage-dynamic-batch-qwen3-rope branch from d7b91b6 to 15c3f12 Compare July 3, 2026 00:14
@mickqian

mickqian commented Jul 4, 2026

Copy link
Copy Markdown
Collaborator

could you resolve the conflict? cheers

@qimcis
qimcis force-pushed the fix-zimage-dynamic-batch-qwen3-rope branch from 11ecbc5 to e4d663c Compare July 6, 2026 07:41
@mickqian

mickqian commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@mickqian

mickqian commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@mickqian
mickqian merged commit fa185ed into sgl-project:main Jul 8, 2026
301 of 367 checks passed
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
Co-authored-by: Mick <mickjagger19@icloud.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion jit-kernel run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Dynamic batching causes severe image quality degradation, line art distorted and subjects missing with same seed & steps

2 participants