Skip to content

[qwen3_5] Correctness fixes: mrope axis handling in fused QK kernel + MTP verify paths (ports from sgl-project) - #751

Open
TobyMint wants to merge 47 commits into
bytedance-iaas:ep_mainfrom
TobyMint:ep-qwen38-correctness
Open

TobyMint wants to merge 47 commits into
bytedance-iaas:ep_mainfrom
TobyMint:ep-qwen38-correctness

Conversation

@TobyMint

@TobyMint TobyMint commented Sep 2, 2026 •

Copy link
Copy Markdown

What

Two correctness fixes for the Qwen3.5/3.8 family, ported from sgl-project where they landed after our sync point. Validated against Qwen3.8-27B (bf16) on 8xH20. No behavior change for paths we already exercise — these close gaps that surface only under image inputs or speculative decoding.

1. mrope axis handling in the fused QK-norm+RoPE+gate kernel (sgl-project sgl-project#34446)

The fused kernel loaded a single position per token, so with image inputs ([3, T] temporal/height/width positions) every rotary lane silently read the temporal row → wrong RoPE on image tokens in all 16 full-attention layers. Text is unaffected (the three rows coincide), which is why text-only smoke can never catch this.

  • The kernel now takes mrope_axis_map ([rotary_dim//2] lane→axis) and reads positions[axis[lane], t] when positions are 2-D.
  • MRotaryEmbedding builds the axis map for every mrope_section style (contiguous / interleaved / GLM round-robin) instead of GLM-only; legacy sgl_kernel call sites keep the GLM-only map via _legacy_axis_map.
  • Adds test/registered/unit/layers/attention/test_fused_qk_rmsnorm_rope_gate_mrope.py — bitwise parity vs a torch reference for 1-D and mrope positions, plus both axis-map styles. Upstream has no dedicated test for this kernel branch; this fills that hole.

2. MTP / linear-attention verify correctness

  • qwen3_5_mtp: captured prefill pads embeddings while target hidden states keep real height; graft the real rows into an equal-height slot before cat+fc instead of misaligning under CUDA graph capture.
  • gdn_backend: honor --linear-attn-verify-backend. The dispatcher re-derived the verify kernel with the auto rule and ignored the stored choice, so an explicit triton selection (required when --mamba-ssm-dtype bfloat16 meets FlashInfer SM90 verify's fp32-state requirement — same class of fix as sgl-project docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state sgl-project/sglang#36611 for H200) had no effect.
  • linear/utils: raise when --enable-deterministic-inference is combined with a FlashInfer GDN prefill.

Validation (H20 x1, Qwen3.8-27B bf16, local)

Path Result
Text TP1 gsm8k 5-shot 200q = 0.975–0.980 (unchanged vs baseline); TP1/2/4/8 smoke correct
MTP (NEXTN steps=3 topk=1 draft=4) boots, gsm8k 0.975, accept len 3.55–3.60 / 4
Vision shape / color / spatial position / count QA all correct post-fix (position QA is exactly the capability the T/H/W axis separation controls)
Unit tests mrope 3/3 bitwise

Scope note

Ports of upstream changes only; no new features. Follow-ups come as separate stacked PRs: GDN kernel perf pack (fused QKVZBA unpack / fused decode / PDL / MTP scatter), SM90 bf16 GEMV (opt-in), compressed-tensors KV scales, and a PD×GDN validation harness.


CI States

Latest PR Test (Base): ❌ Run #33581930370
Latest PR Test (Extra): ❌ Run #33581930143
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.

Kangyan-Zhou and others added 30 commits August 5, 2026 00:26
…emantics (sgl-project#32588) (sgl-project#33668)

Signed-off-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
…pool's page granularity to its allocator (sgl-project#33348) (sgl-project#33762)

Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
…a graph for qwen3.5 nightly test (sgl-project#33772) (sgl-project#33786)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…ild hpc-ops with C++20 (sgl-project#33956) (sgl-project#34028)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…zation (sgl-project#33500) (sgl-project#34032)

Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
…edsl_mla (fold_sq) (sgl-project#33650) (sgl-project#34034)

Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
…eline stage, not per whole model (sgl-project#33666) (sgl-project#34035)

Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Keep the Onion runtime dependency independently reviewable and revertible.

Signed-off-by: Hank Han <hanhan7630@outlook.com>
Install the EIC SDK in the private runtime image and provide the v0.5.17-compatible deployment integration check as one feature.

Signed-off-by: Hank Han <hanhan7630@outlook.com>
Consolidate the Volcengine image build and sync paths, CUDA 13 variants, DeepSeek V4 nightly, reusable kernel build, ep_main PR suites, runner hardening, private schedule policy, gateway build metadata, bounded image provenance, and immutable-image runtime verification into one CI feature.

Signed-off-by: Hank Han <hanhan7630@outlook.com>
Add opt-in zstd (layer compression) and nydus (lazy-loading) image
formats to the SGLang private delivery build, on top of the existing
gzip OCI output. Formats are selected via a CSV `image_formats` input
(default `oci`, so production callers are byte-for-byte unchanged) and
tagged with `-zstd` / `-nydus` suffixes after the cuda suffix.

- get_volcengine_image_tag.py: extract pure build_tag()/validate_suffix()
  helpers and add `--format-suffix` (appended after the cuda suffix);
  covered by scripts/ci/test_get_volcengine_image_tag.py.
- _docker-build-and-publish.yml: add `image_formats` CSV input; build
  once then derive zstd via a cache-hit buildx re-export
  (compression=zstd,force-compression=true,oci-mediatypes=true) and
  nydus via `nydusify convert` from the pushed digest. Each format is
  gated by contains(inputs.image_formats, ...). base/zstd reuse the
  docker pull+run provenance check (falling back to a manifest
  media-type check when the daemon lacks zstd); nydus uses
  `nydusify check` + manifest media-type/annotation assertions (no run).
- release-docker-dev.yml: wire `private_debug_image_formats` into
  build-dev-debug-base so one bounded debug job exercises all three
  formats; production callers keep the default oci.

Also harden all framework_final egress fetches in docker/Dockerfile and
the nydus tooling/convert steps behind an outer retry() loop (github
tarballs, flashinfer pip index, nydus static tarball, just/oh-my-zsh
installers, git clones, and nydusify convert), so intermittent internal
proxy failures (curl 56 / 504 / closed pipe) are retried instead of
hard-failing. These robustness changes are orthogonal to the format
feature and also benefit the default oci production build.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Flip the reusable _docker-build-and-publish.yml image_formats default
from "oci" to "oci,zstd,nydus" so every private caller
(dev regular matrix, release-docker, deepseek-v4 nightly, and the
runtime target) produces and pushes all three formats by default
instead of only the debug base path. Callers can still pass a subset
to narrow the set.

The zstd/nydus steps already iterate every resolved tag, derive all
parameters from inputs (docker_target/cuda/extra_build_args), and reuse
the same provenance verification, so generalizing to the production
multi-tag / multi-variant callers is safe. The runtime target is fully
verifiable too: it ships oniond, the full /sgl-workspace source tree
(test/ + scripts/ci/verify_private_image_runtime.py), pytest (installed
in the framework stage), and identical provenance env/labels, so the
existing target-agnostic verify step applies without degradation.

Also add a private_debug_docker_target input to release-docker-dev.yml
(default framework_final) so a bounded debug dispatch can build the
runtime target with all three formats to empirically confirm the
runtime verify path.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
The zstd re-export step was recompiling the entire image from scratch
(~124min, longer than the base build) because BuildKit's local layer
cache was evicted between the base build and the zstd step by the large
image plus the verify `docker pull`, so every intermediate stage
(torch_deps `.[all]`, framework, deepep) re-ran with zero cache hits.

Export a full mode=max registry build cache from the base build (only
when a zstd format is requested) keyed to the resolved primary tag, and
import it in the zstd step via --cache-from. mode=max is required
because the heavy stages reach the final image via COPY --from and
mode=min/inline would not capture them. The zstd step now hits cache for
all layers and only redoes zstd compression.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
…kernel/

The bytedance/deepseek_v4 branch relocated its kernel build tree from
python/sglang/kernels/aot/ to sgl-kernel/. The nightly daily build
(release-docker-deepseek-v4-nightly.yml) checks out deepseek_v4 for the
kernel wheel but ran the reusable release-whl-kernel.yml from ep_main,
which still cd'd into the removed python/sglang/kernels/aot/ path,
failing build-kernel-wheel in <1s with "No such file or directory" and
cascading to skip build-nightly.

The workflow inputs were already described as sgl-kernel/build.sh, so the
job body was an incomplete migration. Align all 35 path references
(cd, artifact paths, version.py, working-directory) with sgl-kernel/, as
already done in deepseek_v4's own release-whl-kernel.yml.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
The nightly builds deepseek_v4 source but reused ep_main's
release-whl-kernel.yml, coupling that shared file to deepseek_v4's
directory layout. ep_main follows upstream (kernels under
python/sglang/kernels/aot/, per sgl-project#32648) while deepseek_v4 keeps them
under sgl-kernel/, so one shared file cannot serve both.

Split the two consumers:

- Nightly now calls the kernel workflow from the branch it builds:
  uses: .../release-whl-kernel.yml@bytedance/deepseek_v4 (which now
  exposes a workflow_call build-wheel job using sgl-kernel/).
- ep_main's release-whl-kernel.yml is reverted to python/sglang/kernels/aot,
  undoing the earlier sgl-kernel path change (31bf244). Its only local
  consumer, release-docker-dev.yml, builds ep_main source, which uses the
  upstream python/sglang/kernels/aot/ layout.

Each branch's kernel build now tracks its own source structure.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add a workflow_dispatch workflow on x64-docker-build-node that backfills
zstd and nydus image formats for existing OCI images in the Volcengine
serving registry. Uses the same buildx and nydusify commands as the
private delivery build pipeline.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
…er step

Co-authored-by: TRAE CLI <noreply@bytedance.com>
The shared _docker-build-and-publish.yml verify step runs ep_main-specific
runtime smoke tests inside the freshly built image:
`scripts/eic_integration_check.py` (py_compile) plus pytest for
`test_runtime_context.py::TestMoeFlagsGroup` and
`test_model_overrides.py::deepseek_spec_moe_resolution`. Those tests import
`sglang.srt.runtime_context` and `sglang.srt.arg_groups.arg_utils`, which
exist only on ep_main; the deepseek_v4 source tree has neither the modules
nor the test files. The nightly builds v4 source through ep_main's reusable
workflow, so the verify step failed after the image built and pushed
successfully, skipping the zstd/nydus stages.

Add a `verify_image` boolean input (default true, preserving the ep_main
dev/runtime/release callers) that gates all three verify steps, and pass
`verify_image: false` from the deepseek_v4 nightly. The provenance
labels/env are still baked into every image regardless; only the
in-container smoke tests are skipped for v4.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
skopeo does not share docker's credential store, so `skopeo inspect`
against a private registry fails even after `docker login`. Switch to
`docker manifest inspect` which uses the docker daemon's auth.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add source_image_ref input to allow renaming old-format images to standard
daily-build tags before backfilling zstd/nydus formats.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
luoroger37 and others added 17 commits August 24, 2026 11:29
…#729)

Co-authored-by: luoroger37 <luowenjie.roger7@bytedance.com>
Add source_image_is_customer input and customer registry login for
converting customer zstd images back to gzip format in dev registry.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
…r sync

Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
- qwen3_5_mtp: captured prefill pads embeddings while target hidden
  states keep real height; graft the real rows into an equal-height
  slot before cat+fc instead of letting cat broadcast-fail or misalign
  (upstream qwen3_5_mtp forward padding fix).
- gdn_backend: honor --linear-attn-verify-backend. The dispatcher
  re-derived the verify kernel with the auto rule and ignored the
  stored choice, so an explicit triton selection (required when
  --mamba-ssm-dtype bfloat16 meets FlashInfer SM90 verify's fp32-state
  requirement, upstream sgl-project#36611) had no effect.
- linear/utils: raise when --enable-deterministic-inference combines
  with a FlashInfer GDN prefill (upstream _validate_gdn_linear_attn_backends).

Validated on H20 TP1 with NEXTN (steps 3 / topk 1 / draft 4): boots,
gsm8k 200q = 0.975 (bf16 baseline 0.980), accept length 3.55-3.60.
…gl-project#34446)

The fused kernel loaded a single position per token, so with image
inputs ([3, T] temporal/height/width positions) every rotary lane
silently read the temporal row — wrong RoPE on image tokens in all
full-attention layers of Qwen3.5/3.8 hybrids. Text was unaffected
(the three rows coincide), so this never shows in text-only smoke.

Port: the kernel takes an mrope_axis_map ([rotary_dim//2] lane->axis)
and reads positions[axis[lane], t] when positions is 2-D;
MRotaryEmbedding now builds the axis map for every mrope_section style
(contiguous, interleaved, GLM round-robin) instead of GLM only, while
the legacy sgl_kernel call sites keep the GLM-only map via
_legacy_axis_map. Unit test mirrors the kernel math bitwise for 1-D
and mrope positions and checks both axis-map styles.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants