Skip to content

GLM-5.3: CP8 + DCP4 on 0bcd822377da, v3 working baseline - #1

Draft
leoong55 wants to merge 7 commits into
mainfrom
work/glm53-cp8-dcp4-v1
Draft

leoong55 wants to merge 7 commits into
mainfrom
work/glm53-cp8-dcp4-v1

Conversation

@leoong55

@leoong55 leoong55 commented Sep 10, 2026 •

Copy link
Copy Markdown
Owner

Working CP8/DCP4 baseline pinned to upstream 0bcd822. Includes v1–v3, Q8 sparse prefill and the v3 prefix gather plan.

The experimental hybrid DeepEP v4/v4.1 changes were reverted in 2a44c21; its tree exactly matches v3 (258df59). History is preserved.

Further experiments (stock DeepEP auto and prefill breakable CUDA Graph) will be separate pull requests based on work/glm53-v3-baseline.

Baseline user benchmark: 300/300 requests, concurrency 40, 644.15 s, 465.73 output tok/s. This is a performance result, not a numerical quality evaluation.

Adapt SGLang PR #36990 (c6aeb8b9d9128b816777e2b64cbf6603344e8fe1)
to 0bcd822, preserving the pinned top-k resolver and MLA gating.
Use flashmla_kv for prefill/decode and reject HiCache for DSA DCP.

Include 9 passing CPU host-logic tests, upstream integration tests,
verified source overlay installer, manual Docker build/export kit,
and Kubernetes manifest without speculative decoding or CUDA graphs.
GPU/model validation and Docker image build remain pending.
@github-actions github-actions Bot added documentation Improvements or additions to documentation jit-kernel deepseek memory-pool labels Sep 10, 2026
Document eager collective waits and the separate prefill evidence gap. Generate guarded JSON patches from the live Deployment, with rollback, preserving the model PVC and parallel topology. Keep the v1 manifest as the bring-up baseline. GPU capture/replay remains to be validated on H200.
Translate logical top-k into the CP+DCP gathered KV buffer and preserve that layout through native Q8KV8 preparation. Keep flashmla_kv decode and use it for non-CP extend tails with matching metadata. Preserve ordinary non-DCP ragged routing.

Add CPU dispatch/layout regression checks and a v2 Docker build kit that retains the existing launch path. Include the original 60k+15k to 1k, 20-prefix benchmark at concurrency 32. Fourteen CPU host tests pass; CUDA kernels, FP8 accuracy, graph replay, and 8-H200 performance remain unvalidated.
Prepare aligned DCP prefix layout once per forward and gather packed FP8 rows directly into the temporary KV buffer. Reuse batch- and stream-local transport storage, preserving layer-specific KV bytes. Replace the DCP interleave CP index permutation with a rank/row transpose and return independently owned output tensors.

Keep Q8 prefill, flashmla_kv decode, chunk 8192 and CP8/DCP4/EP8 topology. Leave non-DCP CP and generic unaligned/non-packed DCP paths unchanged. Do not change attention math, decode graphs, scheduler or MoE collectives.

Add a v3 image build kit retaining the existing launch path. 22 CPU host tests pass, including byte/order equivalence, padding, layer refresh, output lifetime, stream-key isolation and alignment gating. Cumulative and v2 delta patches apply byte-exactly; installer base hashes, idempotence, verify-only, mismatch rejection and rollback pass. CUDA/NCCL execution, H200 performance and Docker image build remain unvalidated here.
@leoong55 leoong55 changed the title GLM-5.3: Hopper CP8/EP8/DCP4 backport for first launch without HiCache/spec GLM-5.3: Hopper CP8/DCP4 backport, Q8 prefill and prepared gather v3 Sep 10, 2026
Use legacy DeepEP normal BF16 dispatch and the existing W4AFP8 CUTLASS adapter only for CP8 prefill. Preserve standard decode and its CUDA graphs; load a separate TP1 shared expert for local prefill. Allocate the transport arena before KV sizing.

Include cumulative build-kit generator, transactional overlay installer, image-only Deployment patch, rollback instructions and 40 passing CPU contract checks. GPU correctness, image build and performance remain unvalidated.
@leoong55 leoong55 changed the title GLM-5.3: Hopper CP8/DCP4 backport, Q8 prefill and prepared gather v3 GLM-5.3: Hopper CP8/DCP4, Q8 prefill and prefill-only DeepEP v4 Sep 10, 2026
DeepEP's packaged __init__ requires a usable CUDA device, so importing it from Docker RUN fails on ordinary builders even when the dependencies are installed.

At build time check top-level package presence without import. Run the existing legacy API/CUTLASS import checks in launch.py before model startup when prefill DeepEP is enabled, preserving prerequisite failures. Keep the inference runtime overlay unchanged and use a new v4.1 image tag.

44 CPU checks pass, including no-GPU import regression and launcher failure propagation. Docker build and GPU execution remain unvalidated here.
@leoong55 leoong55 changed the title GLM-5.3: Hopper CP8/DCP4, Q8 prefill and prefill-only DeepEP v4 GLM-5.3: CP8/DCP4, Q8 and DeepEP prefill v4.1 build fix Sep 10, 2026
The supplied v4.1 benchmark took 674.12s versus v3 644.15s, with lower output throughput. Restore the v3 tree without rewriting experiment history. Develop stock DeepEP auto and prefill CUDA graphs independently from v3.
@leoong55 leoong55 changed the title GLM-5.3: CP8/DCP4, Q8 and DeepEP prefill v4.1 build fix GLM-5.3: CP8 + DCP4 on 0bcd822377da, v3 working baseline Sep 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek documentation Improvements or additions to documentation jit-kernel memory-pool

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant