Conversation
…emantics (sgl-project#32588) (sgl-project#33668) Signed-off-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
…pool's page granularity to its allocator (sgl-project#33348) (sgl-project#33762) Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
sgl-project#33779) Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
…a graph for qwen3.5 nightly test (sgl-project#33772) (sgl-project#33786) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…ild hpc-ops with C++20 (sgl-project#33956) (sgl-project#34028) Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
…zation (sgl-project#33500) (sgl-project#34032) Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
…regation (sgl-project#33599) (sgl-project#34033) Co-authored-by: Xinyi Song <xinyisong0111@gmail.com>
…edsl_mla (fold_sq) (sgl-project#33650) (sgl-project#34034) Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
…eline stage, not per whole model (sgl-project#33666) (sgl-project#34035) Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
Keep the Onion runtime dependency independently reviewable and revertible. Signed-off-by: Hank Han <hanhan7630@outlook.com>
Install the EIC SDK in the private runtime image and provide the v0.5.17-compatible deployment integration check as one feature. Signed-off-by: Hank Han <hanhan7630@outlook.com>
Consolidate the Volcengine image build and sync paths, CUDA 13 variants, DeepSeek V4 nightly, reusable kernel build, ep_main PR suites, runner hardening, private schedule policy, gateway build metadata, bounded image provenance, and immutable-image runtime verification into one CI feature. Signed-off-by: Hank Han <hanhan7630@outlook.com>
Add opt-in zstd (layer compression) and nydus (lazy-loading) image formats to the SGLang private delivery build, on top of the existing gzip OCI output. Formats are selected via a CSV `image_formats` input (default `oci`, so production callers are byte-for-byte unchanged) and tagged with `-zstd` / `-nydus` suffixes after the cuda suffix. - get_volcengine_image_tag.py: extract pure build_tag()/validate_suffix() helpers and add `--format-suffix` (appended after the cuda suffix); covered by scripts/ci/test_get_volcengine_image_tag.py. - _docker-build-and-publish.yml: add `image_formats` CSV input; build once then derive zstd via a cache-hit buildx re-export (compression=zstd,force-compression=true,oci-mediatypes=true) and nydus via `nydusify convert` from the pushed digest. Each format is gated by contains(inputs.image_formats, ...). base/zstd reuse the docker pull+run provenance check (falling back to a manifest media-type check when the daemon lacks zstd); nydus uses `nydusify check` + manifest media-type/annotation assertions (no run). - release-docker-dev.yml: wire `private_debug_image_formats` into build-dev-debug-base so one bounded debug job exercises all three formats; production callers keep the default oci. Also harden all framework_final egress fetches in docker/Dockerfile and the nydus tooling/convert steps behind an outer retry() loop (github tarballs, flashinfer pip index, nydus static tarball, just/oh-my-zsh installers, git clones, and nydusify convert), so intermittent internal proxy failures (curl 56 / 504 / closed pipe) are retried instead of hard-failing. These robustness changes are orthogonal to the format feature and also benefit the default oci production build. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Flip the reusable _docker-build-and-publish.yml image_formats default from "oci" to "oci,zstd,nydus" so every private caller (dev regular matrix, release-docker, deepseek-v4 nightly, and the runtime target) produces and pushes all three formats by default instead of only the debug base path. Callers can still pass a subset to narrow the set. The zstd/nydus steps already iterate every resolved tag, derive all parameters from inputs (docker_target/cuda/extra_build_args), and reuse the same provenance verification, so generalizing to the production multi-tag / multi-variant callers is safe. The runtime target is fully verifiable too: it ships oniond, the full /sgl-workspace source tree (test/ + scripts/ci/verify_private_image_runtime.py), pytest (installed in the framework stage), and identical provenance env/labels, so the existing target-agnostic verify step applies without degradation. Also add a private_debug_docker_target input to release-docker-dev.yml (default framework_final) so a bounded debug dispatch can build the runtime target with all three formats to empirically confirm the runtime verify path. Co-authored-by: TRAE CLI <noreply@bytedance.com>
The zstd re-export step was recompiling the entire image from scratch (~124min, longer than the base build) because BuildKit's local layer cache was evicted between the base build and the zstd step by the large image plus the verify `docker pull`, so every intermediate stage (torch_deps `.[all]`, framework, deepep) re-ran with zero cache hits. Export a full mode=max registry build cache from the base build (only when a zstd format is requested) keyed to the resolved primary tag, and import it in the zstd step via --cache-from. mode=max is required because the heavy stages reach the final image via COPY --from and mode=min/inline would not capture them. The zstd step now hits cache for all layers and only redoes zstd compression. Co-authored-by: TRAE CLI <noreply@bytedance.com>
…kernel/ The bytedance/deepseek_v4 branch relocated its kernel build tree from python/sglang/kernels/aot/ to sgl-kernel/. The nightly daily build (release-docker-deepseek-v4-nightly.yml) checks out deepseek_v4 for the kernel wheel but ran the reusable release-whl-kernel.yml from ep_main, which still cd'd into the removed python/sglang/kernels/aot/ path, failing build-kernel-wheel in <1s with "No such file or directory" and cascading to skip build-nightly. The workflow inputs were already described as sgl-kernel/build.sh, so the job body was an incomplete migration. Align all 35 path references (cd, artifact paths, version.py, working-directory) with sgl-kernel/, as already done in deepseek_v4's own release-whl-kernel.yml. Co-authored-by: TRAE CLI <noreply@bytedance.com>
The nightly builds deepseek_v4 source but reused ep_main's release-whl-kernel.yml, coupling that shared file to deepseek_v4's directory layout. ep_main follows upstream (kernels under python/sglang/kernels/aot/, per sgl-project#32648) while deepseek_v4 keeps them under sgl-kernel/, so one shared file cannot serve both. Split the two consumers: - Nightly now calls the kernel workflow from the branch it builds: uses: .../release-whl-kernel.yml@bytedance/deepseek_v4 (which now exposes a workflow_call build-wheel job using sgl-kernel/). - ep_main's release-whl-kernel.yml is reverted to python/sglang/kernels/aot, undoing the earlier sgl-kernel path change (31bf244). Its only local consumer, release-docker-dev.yml, builds ep_main source, which uses the upstream python/sglang/kernels/aot/ layout. Each branch's kernel build now tracks its own source structure. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add a workflow_dispatch workflow on x64-docker-build-node that backfills zstd and nydus image formats for existing OCI images in the Volcengine serving registry. Uses the same buildx and nydusify commands as the private delivery build pipeline. Co-authored-by: TRAE CLI <noreply@bytedance.com>
…er step Co-authored-by: TRAE CLI <noreply@bytedance.com>
The shared _docker-build-and-publish.yml verify step runs ep_main-specific runtime smoke tests inside the freshly built image: `scripts/eic_integration_check.py` (py_compile) plus pytest for `test_runtime_context.py::TestMoeFlagsGroup` and `test_model_overrides.py::deepseek_spec_moe_resolution`. Those tests import `sglang.srt.runtime_context` and `sglang.srt.arg_groups.arg_utils`, which exist only on ep_main; the deepseek_v4 source tree has neither the modules nor the test files. The nightly builds v4 source through ep_main's reusable workflow, so the verify step failed after the image built and pushed successfully, skipping the zstd/nydus stages. Add a `verify_image` boolean input (default true, preserving the ep_main dev/runtime/release callers) that gates all three verify steps, and pass `verify_image: false` from the deepseek_v4 nightly. The provenance labels/env are still baked into every image regardless; only the in-container smoke tests are skipped for v4. Co-authored-by: TRAE CLI <noreply@bytedance.com>
skopeo does not share docker's credential store, so `skopeo inspect` against a private registry fails even after `docker login`. Switch to `docker manifest inspect` which uses the docker daemon's auth. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Add source_image_ref input to allow renaming old-format images to standard daily-build tags before backfilling zstd/nydus formats. Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
…#729) Co-authored-by: luoroger37 <luowenjie.roger7@bytedance.com>
Add source_image_is_customer input and customer registry login for converting customer zstd images back to gzip format in dev registry. Co-authored-by: TRAE CLI <noreply@bytedance.com> Co-authored-by: TRAE CLI <traecli@bytedance.com>
…r sync Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
- qwen3_5_mtp: captured prefill pads embeddings while target hidden states keep real height; graft the real rows into an equal-height slot before cat+fc instead of letting cat broadcast-fail or misalign (upstream qwen3_5_mtp forward padding fix). - gdn_backend: honor --linear-attn-verify-backend. The dispatcher re-derived the verify kernel with the auto rule and ignored the stored choice, so an explicit triton selection (required when --mamba-ssm-dtype bfloat16 meets FlashInfer SM90 verify's fp32-state requirement, upstream sgl-project#36611) had no effect. - linear/utils: raise when --enable-deterministic-inference combines with a FlashInfer GDN prefill (upstream _validate_gdn_linear_attn_backends). Validated on H20 TP1 with NEXTN (steps 3 / topk 1 / draft 4): boots, gsm8k 200q = 0.975 (bf16 baseline 0.980), accept length 3.55-3.60.
…gl-project#34446) The fused kernel loaded a single position per token, so with image inputs ([3, T] temporal/height/width positions) every rotary lane silently read the temporal row — wrong RoPE on image tokens in all full-attention layers of Qwen3.5/3.8 hybrids. Text was unaffected (the three rows coincide), so this never shows in text-only smoke. Port: the kernel takes an mrope_axis_map ([rotary_dim//2] lane->axis) and reads positions[axis[lane], t] when positions is 2-D; MRotaryEmbedding now builds the axis map for every mrope_section style (contiguous, interleaved, GLM round-robin) instead of GLM only, while the legacy sgl_kernel call sites keep the GLM-only map via _legacy_axis_map. Unit test mirrors the kernel math bitwise for 1-D and mrope positions and checks both axis-map styles.
HanHan009527
force-pushed
the
ep_main
branch
from
September 15, 2026 04:36
8ab9652 to
f4c61f3
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Two correctness fixes for the Qwen3.5/3.8 family, ported from sgl-project where they landed after our sync point. Validated against Qwen3.8-27B (bf16) on 8xH20. No behavior change for paths we already exercise — these close gaps that surface only under image inputs or speculative decoding.
1. mrope axis handling in the fused QK-norm+RoPE+gate kernel (sgl-project sgl-project#34446)
The fused kernel loaded a single position per token, so with image inputs (
[3, T]temporal/height/width positions) every rotary lane silently read the temporal row → wrong RoPE on image tokens in all 16 full-attention layers. Text is unaffected (the three rows coincide), which is why text-only smoke can never catch this.mrope_axis_map([rotary_dim//2]lane→axis) and readspositions[axis[lane], t]when positions are 2-D.MRotaryEmbeddingbuilds the axis map for everymrope_sectionstyle (contiguous / interleaved / GLM round-robin) instead of GLM-only; legacysgl_kernelcall sites keep the GLM-only map via_legacy_axis_map.test/registered/unit/layers/attention/test_fused_qk_rmsnorm_rope_gate_mrope.py— bitwise parity vs a torch reference for 1-D and mrope positions, plus both axis-map styles. Upstream has no dedicated test for this kernel branch; this fills that hole.2. MTP / linear-attention verify correctness
qwen3_5_mtp: captured prefill pads embeddings while target hidden states keep real height; graft the real rows into an equal-height slot beforecat+fcinstead of misaligning under CUDA graph capture.gdn_backend: honor--linear-attn-verify-backend. The dispatcher re-derived the verify kernel with the auto rule and ignored the stored choice, so an explicit triton selection (required when--mamba-ssm-dtype bfloat16meets FlashInfer SM90 verify's fp32-state requirement — same class of fix as sgl-project docs(cookbook): fix Qwen3.8 Flash Next H200 MTP verify with BF16 SSM state sgl-project/sglang#36611 for H200) had no effect.linear/utils: raise when--enable-deterministic-inferenceis combined with a FlashInfer GDN prefill.Validation (H20 x1, Qwen3.8-27B bf16, local)
Scope note
Ports of upstream changes only; no new features. Follow-ups come as separate stacked PRs: GDN kernel perf pack (fused QKVZBA unpack / fused decode / PDL / MTP scatter), SM90 bf16 GEMV (opt-in), compressed-tensors KV scales, and a PD×GDN validation harness.
CI States
Latest PR Test (Base): ❌ Run #33581930370
Latest PR Test (Extra): ❌ Run #33581930143
Latest PR Test (AMD ROCm 7.2): ➖ No AMD PR run found for this commit.