Repository navigation
Conversation
Adapt SGLang PR #36990 (c6aeb8b9d9128b816777e2b64cbf6603344e8fe1) to 0bcd822, preserving the pinned top-k resolver and MLA gating. Use flashmla_kv for prefill/decode and reject HiCache for DSA DCP. Include 9 passing CPU host-logic tests, upstream integration tests, verified source overlay installer, manual Docker build/export kit, and Kubernetes manifest without speculative decoding or CUDA graphs. GPU/model validation and Docker image build remain pending.
Document eager collective waits and the separate prefill evidence gap. Generate guarded JSON patches from the live Deployment, with rollback, preserving the model PVC and parallel topology. Keep the v1 manifest as the bring-up baseline. GPU capture/replay remains to be validated on H200.
Translate logical top-k into the CP+DCP gathered KV buffer and preserve that layout through native Q8KV8 preparation. Keep flashmla_kv decode and use it for non-CP extend tails with matching metadata. Preserve ordinary non-DCP ragged routing. Add CPU dispatch/layout regression checks and a v2 Docker build kit that retains the existing launch path. Include the original 60k+15k to 1k, 20-prefix benchmark at concurrency 32. Fourteen CPU host tests pass; CUDA kernels, FP8 accuracy, graph replay, and 8-H200 performance remain unvalidated.
Prepare aligned DCP prefix layout once per forward and gather packed FP8 rows directly into the temporary KV buffer. Reuse batch- and stream-local transport storage, preserving layer-specific KV bytes. Replace the DCP interleave CP index permutation with a rank/row transpose and return independently owned output tensors. Keep Q8 prefill, flashmla_kv decode, chunk 8192 and CP8/DCP4/EP8 topology. Leave non-DCP CP and generic unaligned/non-packed DCP paths unchanged. Do not change attention math, decode graphs, scheduler or MoE collectives. Add a v3 image build kit retaining the existing launch path. 22 CPU host tests pass, including byte/order equivalence, padding, layer refresh, output lifetime, stream-key isolation and alignment gating. Cumulative and v2 delta patches apply byte-exactly; installer base hashes, idempotence, verify-only, mismatch rejection and rollback pass. CUDA/NCCL execution, H200 performance and Docker image build remain unvalidated here.
Use legacy DeepEP normal BF16 dispatch and the existing W4AFP8 CUTLASS adapter only for CP8 prefill. Preserve standard decode and its CUDA graphs; load a separate TP1 shared expert for local prefill. Allocate the transport arena before KV sizing. Include cumulative build-kit generator, transactional overlay installer, image-only Deployment patch, rollback instructions and 40 passing CPU contract checks. GPU correctness, image build and performance remain unvalidated.
DeepEP's packaged __init__ requires a usable CUDA device, so importing it from Docker RUN fails on ordinary builders even when the dependencies are installed. At build time check top-level package presence without import. Run the existing legacy API/CUTLASS import checks in launch.py before model startup when prefill DeepEP is enabled, preserving prerequisite failures. Keep the inference runtime overlay unchanged and use a new v4.1 image tag. 44 CPU checks pass, including no-GPU import regression and launcher failure propagation. Docker build and GPU execution remain unvalidated here.
The supplied v4.1 benchmark took 674.12s versus v3 644.15s, with lower output throughput. Restore the v3 tree without rewriting experiment history. Develop stock DeepEP auto and prefill CUDA graphs independently from v3.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Working CP8/DCP4 baseline pinned to upstream 0bcd822. Includes v1–v3, Q8 sparse prefill and the v3 prefix gather plan.
The experimental hybrid DeepEP v4/v4.1 changes were reverted in 2a44c21; its tree exactly matches v3 (258df59). History is preserved.
Further experiments (stock DeepEP auto and prefill breakable CUDA Graph) will be separate pull requests based on work/glm53-v3-baseline.
Baseline user benchmark: 300/300 requests, concurrency 40, 644.15 s, 465.73 output tok/s. This is a performance result, not a numerical quality evaluation.