Skip to content

docker: add Kimi K3 artifacts and build hpc-ops with C++20 - #33956

Merged
Fridge003 merged 4 commits into
mainfrom
codex/migrate-kimi-k3-docker-artifacts
Aug 7, 2026
Merged

Fridge003 merged 4 commits into
mainfrom
codex/migrate-kimi-k3-docker-artifacts

Conversation

@Fridge003

@Fridge003 Fridge003 commented Aug 7, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

The generic CUDA image does not include the pinned TRT-LLM generated-MoE cubin pool or the FlashInfer CuTeDSL MLA decode-context-parallel runtime patch currently installed by the Kimi K3 CUDA 12 and CUDA 13 images.

Modifications

  • Add an independent builder stage that downloads the pinned cubin archive, verifies its SHA-256 digest, extracts it, and requires exactly 1,696 cubin files.
  • Copy the verified pool into both framework and runtime images and export SGLANG_TRTLLM_GEN_MOE_CUBIN_POOL.
  • Install the build-time patch utility and apply the existing seven-file FlashInfer 0.6.15.post1 runtime patch, excluding test diffs absent from the wheel.
  • Keep the existing CUDA 12.6, CUDA 12.9, and CUDA 13.0 FlashInfer dependency-selection paths unchanged.
  • Pin HPC-Ops to upstream commit ab1a402, which switches CXX, CUDA, and NVCC to C++20 and removes the broad PyTorch includes that conflict with NVCC C++20 parsing.

Accuracy Tests

This is a Docker packaging-only change and does not alter model output code.

Validation performed:

  • pre-commit run --files docker/Dockerfile
  • Dockerfile integration contract checks for builder/copy/env/patch wiring
  • Downloaded the pinned archive and verified its SHA-256 digest and 1,696-cubin count
  • Downloaded flashinfer-python==0.6.15.post1 and dry-ran all seven runtime patch files successfully
  • Confirmed the two Kimi K3 Dockerfiles remain unchanged
  • Confirmed the pinned HPC-Ops commit contains all three C++20 settings and removes the two broad PyTorch includes that triggered the NVCC parser error

A full image build was not run locally because the workspace does not provide the Docker CLI.

Speed Tests and Profiling

Not applicable; this changes image contents and build staging, not inference execution paths.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ✅ Run #31157787110
Latest PR Test (Extra): ❌ Run #31157786806

@Fridge003
Fridge003 marked this pull request as ready for review August 7, 2026 06:31
@Fridge003 Fridge003 changed the title docker: add TRT-LLM cubins and FlashInfer MLA DCP patch docker: add Kimi K3 artifacts and build hpc-ops with C++20 Aug 7, 2026
@Fridge003
Fridge003 merged commit 8a22b83 into main Aug 7, 2026
93 of 97 checks passed
@Fridge003
Fridge003 deleted the codex/migrate-kimi-k3-docker-artifacts branch August 7, 2026 20:50
Fridge003 added a commit that referenced this pull request Aug 7, 2026
…ild hpc-ops with C++20 (#33956) (#34028)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
sagearc pushed a commit to sagearc/sglang that referenced this pull request Aug 13, 2026
…ild hpc-ops with C++20 (sgl-project#33956) (sgl-project#34028)

Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
saturn-acc pushed a commit to saturn-acc/sglang that referenced this pull request Aug 16, 2026
Atituiset pushed a commit to Atituiset/sglang that referenced this pull request Sep 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant