Skip to content

[qwen3_5] GDN/linear-attention performance pack: fused QKVZBA split, fused decode, PDL, MTP scatter (ports from sgl-project) - #752

Open
TobyMint wants to merge 51 commits into
bytedance-iaas:ep_mainfrom
TobyMint:ep-qwen38-gdn-perf
Open

TobyMint wants to merge 51 commits into
bytedance-iaas:ep_mainfrom
TobyMint:ep-qwen38-gdn-perf

Conversation

@TobyMint

@TobyMint TobyMint commented Sep 2, 2026 •

Copy link
Copy Markdown

What

Performance pack for the GDN / linear-attention path of the Qwen3.5/3.8 hybrids, ported from sgl-project (sgl-project#34859 / sgl-project#35758). Stacked on #751 (its 2 correctness commits are included below; after that merges this PR shows only the 4 perf commits). All validated on H20 (SM90) with Qwen3.8-27B bf16; each port keeps its upstream tests, all green locally.

1. Ratio-3 fused QKVZBA split (triton_gdn_fused_proj.py, qwen3_5.py)

tl.arange only accepts power-of-two extents, so the fused split/reshape/cat contiguous kernel could not serve the dense 27B hybrids' v/k head ratio 3 — the model fell back to fix_query_key_value_ordering (two .contiguous() copies plus a torch.cat of q/k/v) in every one of the 48 GDN layers. The kernel now walks non-power-of-two groups one HEAD_V-sized head at a time (bitwise-exact; power-of-two groups keep the wide vector access). Upstream test test_gdn_fused_split_head_ratios.py ported (ratio 1/2/3/4, 4/4).

2. Fused GDN decode QKVZ/BA unpack + indexed Conv1D (gdn_backend.py, default on)

Decode keeps the projection pair packed; one Triton kernel unpacks it, applies the indexed causal Conv1D state update, and materializes mixed_qkv/z/b/a — replacing the split kernel plus a second conv launch. Eligibility-checked with explicit fallback, env kill switch SGLANG_ENABLE_GDN_DECODE_FUSED_PROJ_CONV (registered per the Envs convention, names kept upstream-verbatim for future syncs). Upstream test ported: parity vs split+causal_conv1d_update, widths 2-4, dtypes, CUDA-graph replay, OOB-slot masking, fp8 fallback (5/5).

Known caveat (documented in code): the debug-only live parity check (SGLANG_GDN_DECODE_FUSION_VERIFY_REAL_TENSORS=1) fires during CUDA graph capture because capture dummy batches alias mamba slots across rows while the reference uses a private state copy; real batches have unique slots.

3. PDL on GDN recurrent and gated-norm kernels (fla/*, SM90+)

Chain fused_sigmoid_gating_delta_rule_update and layer_norm_gated_fwd behind their producer conv1d_update via programmatic dependent launch. Scheduling-only and bit-exact.

4. Single-launch fused conv-window scatter for the MTP verify commit (mamba_state_scatter_triton.py)

The post-verify conv-state commit looped fused_conv_window_scatter_with_mask once per conv type and again for the track set. fused_conv_window_scatter_multi covers up to 8 (dst, src) pairs and both index sets in one launch; falls back when ineligible. Also relaxes the scatter kernels' dst contiguity assert to accept envelope-strided unified-pool views. Upstream test extended with the multi path (4/4).

Measured-and-excluded: SM90 FlashInfer GDN prefill default

The same upstream PRs default FlashInfer GDN prefill on for SM90 (max_chunk=32768, fp32 state). We measured a regression on H20 and did not port the default — it stays reachable via --linear-attn-prefill-backend flashinfer:

  • chunk 8192, 16×8192: TTFT 15.1 s (Triton) vs 17.8 s (FlashInfer), input throughput 2896 vs 2643 tok/s
  • chunk 32768, identical flags: Triton wins both workloads by ~37% (16×8192: 4934 vs 3609 tok/s; 16×16384: 2403 vs 1757 tok/s)

Validation summary (H20 TP1, dense 27B bf16)

Item Result
Kernel tests ratio 4/4, decode-fusion 5/5, scatter 4/4 — all bitwise/unit-green
Decode TPOT @bs8 18.86–19.04 ms across the individual ports (neutral-positive each)
Decode TPOT @Bs64 21.80–22.06 ms; PDL +0.8% output throughput; gsm8k 0.980
gsm8k 5-shot 200q 0.965–0.980 across configs (bf16 baseline band)
MTP NEXTN 3/1/4 gsm8k 0.975, accept len 3.55–3.60
Prefill deployment note --chunked-prefill-size 32768 boosts prefill ~70% on H20 but needs manual --mem-fraction-static (~0.75) and a bounded --cuda-graph-bs-prefill list (auto planner goes negative on the mamba pool; full 106-size capture ≈35 min)

CI States

Latest PR Test (Base): ⚠️ Run #33582388625
Latest PR Test (Extra): ⚠️ Run #33582388460

Kangyan-Zhou and others added 30 commits August 5, 2026 00:26
…emantics (sgl-project#32588) (sgl-project#33668)

Signed-off-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
…pool's page granularity to its allocator (sgl-project#33348) (sgl-project#33762)

Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
…a graph for qwen3.5 nightly test (sgl-project#33772) (sgl-project#33786)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…ild hpc-ops with C++20 (sgl-project#33956) (sgl-project#34028)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…zation (sgl-project#33500) (sgl-project#34032)

Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
…edsl_mla (fold_sq) (sgl-project#33650) (sgl-project#34034)

Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
…eline stage, not per whole model (sgl-project#33666) (sgl-project#34035)

Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Keep the Onion runtime dependency independently reviewable and revertible.

Signed-off-by: Hank Han <hanhan7630@outlook.com>
Install the EIC SDK in the private runtime image and provide the v0.5.17-compatible deployment integration check as one feature.

Signed-off-by: Hank Han <hanhan7630@outlook.com>
Consolidate the Volcengine image build and sync paths, CUDA 13 variants, DeepSeek V4 nightly, reusable kernel build, ep_main PR suites, runner hardening, private schedule policy, gateway build metadata, bounded image provenance, and immutable-image runtime verification into one CI feature.

Signed-off-by: Hank Han <hanhan7630@outlook.com>
Add opt-in zstd (layer compression) and nydus (lazy-loading) image
formats to the SGLang private delivery build, on top of the existing
gzip OCI output. Formats are selected via a CSV `image_formats` input
(default `oci`, so production callers are byte-for-byte unchanged) and
tagged with `-zstd` / `-nydus` suffixes after the cuda suffix.

- get_volcengine_image_tag.py: extract pure build_tag()/validate_suffix()
  helpers and add `--format-suffix` (appended after the cuda suffix);
  covered by scripts/ci/test_get_volcengine_image_tag.py.
- _docker-build-and-publish.yml: add `image_formats` CSV input; build
  once then derive zstd via a cache-hit buildx re-export
  (compression=zstd,force-compression=true,oci-mediatypes=true) and
  nydus via `nydusify convert` from the pushed digest. Each format is
  gated by contains(inputs.image_formats, ...). base/zstd reuse the
  docker pull+run provenance check (falling back to a manifest
  media-type check when the daemon lacks zstd); nydus uses
  `nydusify check` + manifest media-type/annotation assertions (no run).
- release-docker-dev.yml: wire `private_debug_image_formats` into
  build-dev-debug-base so one bounded debug job exercises all three
  formats; production callers keep the default oci.

Also harden all framework_final egress fetches in docker/Dockerfile and
the nydus tooling/convert steps behind an outer retry() loop (github
tarballs, flashinfer pip index, nydus static tarball, just/oh-my-zsh
installers, git clones, and nydusify convert), so intermittent internal
proxy failures (curl 56 / 504 / closed pipe) are retried instead of
hard-failing. These robustness changes are orthogonal to the format
feature and also benefit the default oci production build.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Flip the reusable _docker-build-and-publish.yml image_formats default
from "oci" to "oci,zstd,nydus" so every private caller
(dev regular matrix, release-docker, deepseek-v4 nightly, and the
runtime target) produces and pushes all three formats by default
instead of only the debug base path. Callers can still pass a subset
to narrow the set.

The zstd/nydus steps already iterate every resolved tag, derive all
parameters from inputs (docker_target/cuda/extra_build_args), and reuse
the same provenance verification, so generalizing to the production
multi-tag / multi-variant callers is safe. The runtime target is fully
verifiable too: it ships oniond, the full /sgl-workspace source tree
(test/ + scripts/ci/verify_private_image_runtime.py), pytest (installed
in the framework stage), and identical provenance env/labels, so the
existing target-agnostic verify step applies without degradation.

Also add a private_debug_docker_target input to release-docker-dev.yml
(default framework_final) so a bounded debug dispatch can build the
runtime target with all three formats to empirically confirm the
runtime verify path.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
The zstd re-export step was recompiling the entire image from scratch
(~124min, longer than the base build) because BuildKit's local layer
cache was evicted between the base build and the zstd step by the large
image plus the verify `docker pull`, so every intermediate stage
(torch_deps `.[all]`, framework, deepep) re-ran with zero cache hits.

Export a full mode=max registry build cache from the base build (only
when a zstd format is requested) keyed to the resolved primary tag, and
import it in the zstd step via --cache-from. mode=max is required
because the heavy stages reach the final image via COPY --from and
mode=min/inline would not capture them. The zstd step now hits cache for
all layers and only redoes zstd compression.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
…kernel/

The bytedance/deepseek_v4 branch relocated its kernel build tree from
python/sglang/kernels/aot/ to sgl-kernel/. The nightly daily build
(release-docker-deepseek-v4-nightly.yml) checks out deepseek_v4 for the
kernel wheel but ran the reusable release-whl-kernel.yml from ep_main,
which still cd'd into the removed python/sglang/kernels/aot/ path,
failing build-kernel-wheel in <1s with "No such file or directory" and
cascading to skip build-nightly.

The workflow inputs were already described as sgl-kernel/build.sh, so the
job body was an incomplete migration. Align all 35 path references
(cd, artifact paths, version.py, working-directory) with sgl-kernel/, as
already done in deepseek_v4's own release-whl-kernel.yml.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
The nightly builds deepseek_v4 source but reused ep_main's
release-whl-kernel.yml, coupling that shared file to deepseek_v4's
directory layout. ep_main follows upstream (kernels under
python/sglang/kernels/aot/, per sgl-project#32648) while deepseek_v4 keeps them
under sgl-kernel/, so one shared file cannot serve both.

Split the two consumers:

- Nightly now calls the kernel workflow from the branch it builds:
  uses: .../release-whl-kernel.yml@bytedance/deepseek_v4 (which now
  exposes a workflow_call build-wheel job using sgl-kernel/).
- ep_main's release-whl-kernel.yml is reverted to python/sglang/kernels/aot,
  undoing the earlier sgl-kernel path change (31bf244). Its only local
  consumer, release-docker-dev.yml, builds ep_main source, which uses the
  upstream python/sglang/kernels/aot/ layout.

Each branch's kernel build now tracks its own source structure.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add a workflow_dispatch workflow on x64-docker-build-node that backfills
zstd and nydus image formats for existing OCI images in the Volcengine
serving registry. Uses the same buildx and nydusify commands as the
private delivery build pipeline.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
…er step

Co-authored-by: TRAE CLI <noreply@bytedance.com>
The shared _docker-build-and-publish.yml verify step runs ep_main-specific
runtime smoke tests inside the freshly built image:
`scripts/eic_integration_check.py` (py_compile) plus pytest for
`test_runtime_context.py::TestMoeFlagsGroup` and
`test_model_overrides.py::deepseek_spec_moe_resolution`. Those tests import
`sglang.srt.runtime_context` and `sglang.srt.arg_groups.arg_utils`, which
exist only on ep_main; the deepseek_v4 source tree has neither the modules
nor the test files. The nightly builds v4 source through ep_main's reusable
workflow, so the verify step failed after the image built and pushed
successfully, skipping the zstd/nydus stages.

Add a `verify_image` boolean input (default true, preserving the ep_main
dev/runtime/release callers) that gates all three verify steps, and pass
`verify_image: false` from the deepseek_v4 nightly. The provenance
labels/env are still baked into every image regardless; only the
in-container smoke tests are skipped for v4.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
skopeo does not share docker's credential store, so `skopeo inspect`
against a private registry fails even after `docker login`. Switch to
`docker manifest inspect` which uses the docker daemon's auth.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add source_image_ref input to allow renaming old-format images to standard
daily-build tags before backfilling zstd/nydus formats.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
luoroger37 and others added 21 commits August 24, 2026 11:29
…#729)

Co-authored-by: luoroger37 <luowenjie.roger7@bytedance.com>
Add source_image_is_customer input and customer registry login for
converting customer zstd images back to gzip format in dev registry.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
…r sync

Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
- qwen3_5_mtp: captured prefill pads embeddings while target hidden
  states keep real height; graft the real rows into an equal-height
  slot before cat+fc instead of letting cat broadcast-fail or misalign
  (upstream qwen3_5_mtp forward padding fix).
- gdn_backend: honor --linear-attn-verify-backend. The dispatcher
  re-derived the verify kernel with the auto rule and ignored the
  stored choice, so an explicit triton selection (required when
  --mamba-ssm-dtype bfloat16 meets FlashInfer SM90 verify's fp32-state
  requirement, upstream sgl-project#36611) had no effect.
- linear/utils: raise when --enable-deterministic-inference combines
  with a FlashInfer GDN prefill (upstream _validate_gdn_linear_attn_backends).

Validated on H20 TP1 with NEXTN (steps 3 / topk 1 / draft 4): boots,
gsm8k 200q = 0.975 (bf16 baseline 0.980), accept length 3.55-3.60.
…gl-project#34446)

The fused kernel loaded a single position per token, so with image
inputs ([3, T] temporal/height/width positions) every rotary lane
silently read the temporal row — wrong RoPE on image tokens in all
full-attention layers of Qwen3.5/3.8 hybrids. Text was unaffected
(the three rows coincide), so this never shows in text-only smoke.

Port: the kernel takes an mrope_axis_map ([rotary_dim//2] lane->axis)
and reads positions[axis[lane], t] when positions is 2-D;
MRotaryEmbedding now builds the axis map for every mrope_section style
(contiguous, interleaved, GLM round-robin) instead of GLM only, while
the legacy sgl_kernel call sites keep the GLM-only map via
_legacy_axis_map. Unit test mirrors the kernel math bitwise for 1-D
and mrope positions and checks both axis-map styles.
Port of the V_POW2 walk from sgl-project sgl-project#34859: tl.arange only accepts
power-of-two extents, so the fused split/reshape/cat contiguous kernel
could not serve the v/k head ratio 3 of the dense 27B hybrids and the
model fell back to fix_query_key_value_ordering (two .contiguous()
copies plus a torch.cat of q/k/v) in every GDN layer. The kernel now
walks non-power-of-two groups one HEAD_V-sized head at a time
(bitwise-exact; power-of-two groups keep the wide vector access), the
model whitelist admits ratio 3 on CUDA, and the ported ratio test
covers 1/2/3/4.

Measured on H20 SM90 (16x8192 prefill / 64x512 decode): neutral to
slightly positive; adopted over the FlashInfer GDN prefill default
from the same PR, which regressed prefill TTFT ~9-18% on this shape
and stays opt-in via --linear-attn-prefill-backend flashinfer.
sgl-project#35758)

Decode fusion: the projected QKVZ/BA pair stays packed and one Triton
kernel unpacks it, applies the indexed causal Conv1D state update, and
materializes mixed_qkv/z/b/a, replacing the split kernel plus a second
conv launch. Eligibility-checked with an explicit fallback, env-gated
(SGLANG_ENABLE_GDN_DECODE_FUSED_PROJ_CONV, default on; names kept
upstream-verbatim for future syncs), with an opt-in live parity check
(SGLANG_GDN_DECODE_FUSION_VERIFY_REAL_TENSORS) and per-layer hit logs.
Env vars registered per the Envs convention.

Ported test covers parity vs split+causal_conv1d_update, widths 2-4,
dtypes, CUDA-graph replay, OOB-slot masking, and the fp8 fallback
(5/5 on H20).

Measured on H20 TP1, dense 27B: neutral-to-slightly-positive decode
(bs8 TPOT 18.91 vs 19.04 ms; bs64 21.91 vs 22.02 ms); gsm8k 200q
0.965 vs 0.975 (within run noise). Caveat: the live parity check
fires during CUDA graph capture because capture dummy batches alias
mamba slots across rows (reference uses a private state copy); it is
a debug-only mode and real batches have unique slots.
After a speculative verify, the conv-state commit looped
fused_conv_window_scatter_with_mask once per conv type and again for
the track set — with the 27B hybrid's 48 GDN layers that is one kernel
per type per index set. Add fused_conv_window_scatter_multi: one
launch covers up to 8 (dst, src) conv-type pairs and both index sets
(accept commit + interval-crossing track) via an int64 metadata
tensor; falls back to the per-type loop when ineligible (non-bf16,
rank/shape mismatch, > 8 types). Also relax the scatter kernels' dst
contiguity assert to _require_entry_contiguous_dst (envelope-strided
unified-pool views keep arbitrary layer/slot strides; only the
trailing entry dims must be contiguous). The eager
_verify_commit_step_indices index math stays as is — upstream's
fused_commit_track_indices has no callers.

Upstream scatter test extended with the multi path: 4/4 on H20.
E2E (TP1, NEXTN 3/1/4): gsm8k 0.975, accept len 3.60, 513 tok/s.
Chain fused_sigmoid_gating_delta_rule_update and layer_norm_gated_fwd
behind their producer conv1d_update via programmatic dependent launch
(gdc_wait/gdc_launch_dependents + launch_pdl, gated on
is_arch_support_pdl). Scheduling-only and bit-exact: all global loads
sit behind the wait, consumers still fence on full completion.

H20 TP1 dense 27B: decode TPOT 21.80 vs 21.91 ms @ bs64 (+0.8% output
throughput), 18.86 vs 18.91 ms @ bs8; gsm8k 200q 0.980.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants