[qwen3_5] GDN/linear-attention performance pack: fused QKVZBA split, fused decode, PDL, MTP scatter (ports from sgl-project) - #752
Open
TobyMint wants to merge 51 commits into
Conversation
…emantics (sgl-project#32588) (sgl-project#33668) Signed-off-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
…pool's page granularity to its allocator (sgl-project#33348) (sgl-project#33762) Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
sgl-project#33779) Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
…a graph for qwen3.5 nightly test (sgl-project#33772) (sgl-project#33786) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…ild hpc-ops with C++20 (sgl-project#33956) (sgl-project#34028) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…zation (sgl-project#33500) (sgl-project#34032) Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
…regation (sgl-project#33599) (sgl-project#34033) Co-authored-by: Xinyi Song <xinyisong0111@gmail.com>
…edsl_mla (fold_sq) (sgl-project#33650) (sgl-project#34034) Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
…eline stage, not per whole model (sgl-project#33666) (sgl-project#34035) Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Keep the Onion runtime dependency independently reviewable and revertible. Signed-off-by: Hank Han <hanhan7630@outlook.com>
Install the EIC SDK in the private runtime image and provide the v0.5.17-compatible deployment integration check as one feature. Signed-off-by: Hank Han <hanhan7630@outlook.com>
Consolidate the Volcengine image build and sync paths, CUDA 13 variants, DeepSeek V4 nightly, reusable kernel build, ep_main PR suites, runner hardening, private schedule policy, gateway build metadata, bounded image provenance, and immutable-image runtime verification into one CI feature. Signed-off-by: Hank Han <hanhan7630@outlook.com>
Add opt-in zstd (layer compression) and nydus (lazy-loading) image formats to the SGLang private delivery build, on top of the existing gzip OCI output. Formats are selected via a CSV `image_formats` input (default `oci`, so production callers are byte-for-byte unchanged) and tagged with `-zstd` / `-nydus` suffixes after the cuda suffix. - get_volcengine_image_tag.py: extract pure build_tag()/validate_suffix() helpers and add `--format-suffix` (appended after the cuda suffix); covered by scripts/ci/test_get_volcengine_image_tag.py. - _docker-build-and-publish.yml: add `image_formats` CSV input; build once then derive zstd via a cache-hit buildx re-export (compression=zstd,force-compression=true,oci-mediatypes=true) and nydus via `nydusify convert` from the pushed digest. Each format is gated by contains(inputs.image_formats, ...). base/zstd reuse the docker pull+run provenance check (falling back to a manifest media-type check when the daemon lacks zstd); nydus uses `nydusify check` + manifest media-type/annotation assertions (no run). - release-docker-dev.yml: wire `private_debug_image_formats` into build-dev-debug-base so one bounded debug job exercises all three formats; production callers keep the default oci. Also harden all framework_final egress fetches in docker/Dockerfile and the nydus tooling/convert steps behind an outer retry() loop (github tarballs, flashinfer pip index, nydus static tarball, just/oh-my-zsh installers, git clones, and nydusify convert), so intermittent internal proxy failures (curl 56 / 504 / closed pipe) are retried instead of hard-failing. These robustness changes are orthogonal to the format feature and also benefit the default oci production build. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Flip the reusable _docker-build-and-publish.yml image_formats default from "oci" to "oci,zstd,nydus" so every private caller (dev regular matrix, release-docker, deepseek-v4 nightly, and the runtime target) produces and pushes all three formats by default instead of only the debug base path. Callers can still pass a subset to narrow the set. The zstd/nydus steps already iterate every resolved tag, derive all parameters from inputs (docker_target/cuda/extra_build_args), and reuse the same provenance verification, so generalizing to the production multi-tag / multi-variant callers is safe. The runtime target is fully verifiable too: it ships oniond, the full /sgl-workspace source tree (test/ + scripts/ci/verify_private_image_runtime.py), pytest (installed in the framework stage), and identical provenance env/labels, so the existing target-agnostic verify step applies without degradation. Also add a private_debug_docker_target input to release-docker-dev.yml (default framework_final) so a bounded debug dispatch can build the runtime target with all three formats to empirically confirm the runtime verify path. Co-authored-by: TRAE CLI <noreply@bytedance.com>
The zstd re-export step was recompiling the entire image from scratch (~124min, longer than the base build) because BuildKit's local layer cache was evicted between the base build and the zstd step by the large image plus the verify `docker pull`, so every intermediate stage (torch_deps `.[all]`, framework, deepep) re-ran with zero cache hits. Export a full mode=max registry build cache from the base build (only when a zstd format is requested) keyed to the resolved primary tag, and import it in the zstd step via --cache-from. mode=max is required because the heavy stages reach the final image via COPY --from and mode=min/inline would not capture them. The zstd step now hits cache for all layers and only redoes zstd compression. Co-authored-by: TRAE CLI <noreply@bytedance.com>
…kernel/ The bytedance/deepseek_v4 branch relocated its kernel build tree from python/sglang/kernels/aot/ to sgl-kernel/. The nightly daily build (release-docker-deepseek-v4-nightly.yml) checks out deepseek_v4 for the kernel wheel but ran the reusable release-whl-kernel.yml from ep_main, which still cd'd into the removed python/sglang/kernels/aot/ path, failing build-kernel-wheel in <1s with "No such file or directory" and cascading to skip build-nightly. The workflow inputs were already described as sgl-kernel/build.sh, so the job body was an incomplete migration. Align all 35 path references (cd, artifact paths, version.py, working-directory) with sgl-kernel/, as already done in deepseek_v4's own release-whl-kernel.yml. Co-authored-by: TRAE CLI <noreply@bytedance.com>
The nightly builds deepseek_v4 source but reused ep_main's release-whl-kernel.yml, coupling that shared file to deepseek_v4's directory layout. ep_main follows upstream (kernels under python/sglang/kernels/aot/, per sgl-project#32648) while deepseek_v4 keeps them under sgl-kernel/, so one shared file cannot serve both. Split the two consumers: - Nightly now calls the kernel workflow from the branch it builds: uses: .../release-whl-kernel.yml@bytedance/deepseek_v4 (which now exposes a workflow_call build-wheel job using sgl-kernel/). - ep_main's release-whl-kernel.yml is reverted to python/sglang/kernels/aot, undoing the earlier sgl-kernel path change (31bf244). Its only local consumer, release-docker-dev.yml, builds ep_main source, which uses the upstream python/sglang/kernels/aot/ layout. Each branch's kernel build now tracks its own source structure. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add a workflow_dispatch workflow on x64-docker-build-node that backfills zstd and nydus image formats for existing OCI images in the Volcengine serving registry. Uses the same buildx and nydusify commands as the private delivery build pipeline. Co-authored-by: TRAE CLI <noreply@bytedance.com>
…er step Co-authored-by: TRAE CLI <noreply@bytedance.com>
The shared _docker-build-and-publish.yml verify step runs ep_main-specific runtime smoke tests inside the freshly built image: `scripts/eic_integration_check.py` (py_compile) plus pytest for `test_runtime_context.py::TestMoeFlagsGroup` and `test_model_overrides.py::deepseek_spec_moe_resolution`. Those tests import `sglang.srt.runtime_context` and `sglang.srt.arg_groups.arg_utils`, which exist only on ep_main; the deepseek_v4 source tree has neither the modules nor the test files. The nightly builds v4 source through ep_main's reusable workflow, so the verify step failed after the image built and pushed successfully, skipping the zstd/nydus stages. Add a `verify_image` boolean input (default true, preserving the ep_main dev/runtime/release callers) that gates all three verify steps, and pass `verify_image: false` from the deepseek_v4 nightly. The provenance labels/env are still baked into every image regardless; only the in-container smoke tests are skipped for v4. Co-authored-by: TRAE CLI <noreply@bytedance.com>
skopeo does not share docker's credential store, so `skopeo inspect` against a private registry fails even after `docker login`. Switch to `docker manifest inspect` which uses the docker daemon's auth. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add source_image_ref input to allow renaming old-format images to standard daily-build tags before backfilling zstd/nydus formats. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
…#729) Co-authored-by: luoroger37 <luowenjie.roger7@bytedance.com>
Add source_image_is_customer input and customer registry login for converting customer zstd images back to gzip format in dev registry. Co-authored-by: TRAE CLI <noreply@bytedance.com> Co-authored-by: TRAE CLI <traecli@bytedance.com>
…r sync Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
- qwen3_5_mtp: captured prefill pads embeddings while target hidden states keep real height; graft the real rows into an equal-height slot before cat+fc instead of letting cat broadcast-fail or misalign (upstream qwen3_5_mtp forward padding fix). - gdn_backend: honor --linear-attn-verify-backend. The dispatcher re-derived the verify kernel with the auto rule and ignored the stored choice, so an explicit triton selection (required when --mamba-ssm-dtype bfloat16 meets FlashInfer SM90 verify's fp32-state requirement, upstream sgl-project#36611) had no effect. - linear/utils: raise when --enable-deterministic-inference combines with a FlashInfer GDN prefill (upstream _validate_gdn_linear_attn_backends). Validated on H20 TP1 with NEXTN (steps 3 / topk 1 / draft 4): boots, gsm8k 200q = 0.975 (bf16 baseline 0.980), accept length 3.55-3.60.
…gl-project#34446) The fused kernel loaded a single position per token, so with image inputs ([3, T] temporal/height/width positions) every rotary lane silently read the temporal row — wrong RoPE on image tokens in all full-attention layers of Qwen3.5/3.8 hybrids. Text was unaffected (the three rows coincide), so this never shows in text-only smoke. Port: the kernel takes an mrope_axis_map ([rotary_dim//2] lane->axis) and reads positions[axis[lane], t] when positions is 2-D; MRotaryEmbedding now builds the axis map for every mrope_section style (contiguous, interleaved, GLM round-robin) instead of GLM only, while the legacy sgl_kernel call sites keep the GLM-only map via _legacy_axis_map. Unit test mirrors the kernel math bitwise for 1-D and mrope positions and checks both axis-map styles.
Port of the V_POW2 walk from sgl-project sgl-project#34859: tl.arange only accepts power-of-two extents, so the fused split/reshape/cat contiguous kernel could not serve the v/k head ratio 3 of the dense 27B hybrids and the model fell back to fix_query_key_value_ordering (two .contiguous() copies plus a torch.cat of q/k/v) in every GDN layer. The kernel now walks non-power-of-two groups one HEAD_V-sized head at a time (bitwise-exact; power-of-two groups keep the wide vector access), the model whitelist admits ratio 3 on CUDA, and the ported ratio test covers 1/2/3/4. Measured on H20 SM90 (16x8192 prefill / 64x512 decode): neutral to slightly positive; adopted over the FlashInfer GDN prefill default from the same PR, which regressed prefill TTFT ~9-18% on this shape and stays opt-in via --linear-attn-prefill-backend flashinfer.
sgl-project#35758) Decode fusion: the projected QKVZ/BA pair stays packed and one Triton kernel unpacks it, applies the indexed causal Conv1D state update, and materializes mixed_qkv/z/b/a, replacing the split kernel plus a second conv launch. Eligibility-checked with an explicit fallback, env-gated (SGLANG_ENABLE_GDN_DECODE_FUSED_PROJ_CONV, default on; names kept upstream-verbatim for future syncs), with an opt-in live parity check (SGLANG_GDN_DECODE_FUSION_VERIFY_REAL_TENSORS) and per-layer hit logs. Env vars registered per the Envs convention. Ported test covers parity vs split+causal_conv1d_update, widths 2-4, dtypes, CUDA-graph replay, OOB-slot masking, and the fp8 fallback (5/5 on H20). Measured on H20 TP1, dense 27B: neutral-to-slightly-positive decode (bs8 TPOT 18.91 vs 19.04 ms; bs64 21.91 vs 22.02 ms); gsm8k 200q 0.965 vs 0.975 (within run noise). Caveat: the live parity check fires during CUDA graph capture because capture dummy batches alias mamba slots across rows (reference uses a private state copy); it is a debug-only mode and real batches have unique slots.
After a speculative verify, the conv-state commit looped fused_conv_window_scatter_with_mask once per conv type and again for the track set — with the 27B hybrid's 48 GDN layers that is one kernel per type per index set. Add fused_conv_window_scatter_multi: one launch covers up to 8 (dst, src) conv-type pairs and both index sets (accept commit + interval-crossing track) via an int64 metadata tensor; falls back to the per-type loop when ineligible (non-bf16, rank/shape mismatch, > 8 types). Also relax the scatter kernels' dst contiguity assert to _require_entry_contiguous_dst (envelope-strided unified-pool views keep arbitrary layer/slot strides; only the trailing entry dims must be contiguous). The eager _verify_commit_step_indices index math stays as is — upstream's fused_commit_track_indices has no callers. Upstream scatter test extended with the multi path: 4/4 on H20. E2E (TP1, NEXTN 3/1/4): gsm8k 0.975, accept len 3.60, 513 tok/s.
Chain fused_sigmoid_gating_delta_rule_update and layer_norm_gated_fwd behind their producer conv1d_update via programmatic dependent launch (gdc_wait/gdc_launch_dependents + launch_pdl, gated on is_arch_support_pdl). Scheduling-only and bit-exact: all global loads sit behind the wait, consumers still fence on full completion. H20 TP1 dense 27B: decode TPOT 21.80 vs 21.91 ms @ bs64 (+0.8% output throughput), 18.86 vs 18.91 ms @ bs8; gsm8k 200q 0.980.
HanHan009527
force-pushed
the
ep_main
branch
from
September 15, 2026 04:36
8ab9652 to
f4c61f3
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Performance pack for the GDN / linear-attention path of the Qwen3.5/3.8 hybrids, ported from sgl-project (sgl-project#34859 / sgl-project#35758). Stacked on #751 (its 2 correctness commits are included below; after that merges this PR shows only the 4 perf commits). All validated on H20 (SM90) with Qwen3.8-27B bf16; each port keeps its upstream tests, all green locally.
1. Ratio-3 fused QKVZBA split (
triton_gdn_fused_proj.py,qwen3_5.py)tl.arangeonly accepts power-of-two extents, so the fused split/reshape/cat contiguous kernel could not serve the dense 27B hybrids' v/k head ratio 3 — the model fell back tofix_query_key_value_ordering(two.contiguous()copies plus atorch.catof q/k/v) in every one of the 48 GDN layers. The kernel now walks non-power-of-two groups one HEAD_V-sized head at a time (bitwise-exact; power-of-two groups keep the wide vector access). Upstream testtest_gdn_fused_split_head_ratios.pyported (ratio 1/2/3/4, 4/4).2. Fused GDN decode QKVZ/BA unpack + indexed Conv1D (
gdn_backend.py, default on)Decode keeps the projection pair packed; one Triton kernel unpacks it, applies the indexed causal Conv1D state update, and materializes
mixed_qkv/z/b/a— replacing the split kernel plus a second conv launch. Eligibility-checked with explicit fallback, env kill switchSGLANG_ENABLE_GDN_DECODE_FUSED_PROJ_CONV(registered per the Envs convention, names kept upstream-verbatim for future syncs). Upstream test ported: parity vs split+causal_conv1d_update, widths 2-4, dtypes, CUDA-graph replay, OOB-slot masking, fp8 fallback (5/5).Known caveat (documented in code): the debug-only live parity check (
SGLANG_GDN_DECODE_FUSION_VERIFY_REAL_TENSORS=1) fires during CUDA graph capture because capture dummy batches alias mamba slots across rows while the reference uses a private state copy; real batches have unique slots.3. PDL on GDN recurrent and gated-norm kernels (
fla/*, SM90+)Chain
fused_sigmoid_gating_delta_rule_updateandlayer_norm_gated_fwdbehind their producer conv1d_update via programmatic dependent launch. Scheduling-only and bit-exact.4. Single-launch fused conv-window scatter for the MTP verify commit (
mamba_state_scatter_triton.py)The post-verify conv-state commit looped
fused_conv_window_scatter_with_maskonce per conv type and again for the track set.fused_conv_window_scatter_multicovers up to 8 (dst, src) pairs and both index sets in one launch; falls back when ineligible. Also relaxes the scatter kernels' dst contiguity assert to accept envelope-strided unified-pool views. Upstream test extended with the multi path (4/4).Measured-and-excluded: SM90 FlashInfer GDN prefill default
The same upstream PRs default FlashInfer GDN prefill on for SM90 (
max_chunk=32768, fp32 state). We measured a regression on H20 and did not port the default — it stays reachable via--linear-attn-prefill-backend flashinfer:Validation summary (H20 TP1, dense 27B bf16)
--chunked-prefill-size 32768boosts prefill ~70% on H20 but needs manual--mem-fraction-static(~0.75) and a bounded--cuda-graph-bs-prefilllist (auto planner goes negative on the mamba pool; full 106-size capture ≈35 min)CI States
Latest PR Test (Base):⚠️ Run #33582388625⚠️ Run #33582388460
Latest PR Test (Extra):