[AMD] Add MI355X Kimi-K2.6 tuning artifacts - #23381
Open
jhinpan wants to merge 9 commits into
Open
Conversation
Add the tuned MI355X fused-MoE config used by Kimi-K2.6 and make the DeepSeek weight loader parallelism configurable for large MoE checkpoints. Also make the fused-MoE tuner pick a stable HIP device id so the config can be regenerated from upstream sources.
jhinpan
requested review from
BBuf,
Edwardf0t1,
Fridge003,
HaiShaw,
Ying1123,
ch-wan,
fzyzcjy,
ispobock and
merrymercy
as code owners
April 21, 2026 14:34
Contributor
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
- Register SGLANG_DEEPSEEK_LOAD_MAX_WORKERS via envs (EnvInt) instead of a raw os.environ.get; default is None so ThreadPoolExecutor keeps its prior pass-through behaviour (min(32, cpu_count() + 4)) on NV hosts that do not opt in. Large MoE checkpoints on MI355X can still set SGLANG_DEEPSEEK_LOAD_MAX_WORKERS=4. - Drop inline `import os as _os` in the weight loader and route through the already-imported `envs` module; fixes the black line-length CI failure as a side effect. - Use the module-level `_is_hip` cache in the MoE tuner instead of calling `is_hip()` per BenchmarkWorker init. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Extend the fused-MoE Triton config for (E=384, N=256, MI355X, int4_w4a16) from 4 to 13 batch-size buckets so nearest-neighbor fallback is tighter for mid-M workloads (e.g. moderate concurrency, chunked prefill). Final coverage: [1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 8192], which matches Arist12's K2.5 default bucket set plus our longer-prefill 8192 bucket. Each bucket was produced by the upstream benchmark/kernels/fused_moe_triton/tuning_fused_moe_triton.py sweep over 400 candidate configs per bucket.
Collaborator
Author
jhinpan
added a commit
to amdpilot-org/amdpilot-evals
that referenced
this pull request
Apr 23, 2026
Adds a hand-authored eval instance that targets the FlyDSL fused-MoE port from AMD's Kimi-K2.5 MI300X optimization blog to Kimi-K2.6 on 4x MI355X, on top of the baseline established by sgl-project/sglang#23381. Contents: - task.yaml : optimize-type task, phase1_baseline=true, frontier_model=true, TP=4, 4xMI355X, FlyDSL env OFF at start so Phase 1 reproduces PR #23381 numbers (DSL2_ROOT + MLIR_PATH + CK_TILE_FLOAT_TO_BFLOAT16 pre-set so import works as soon as agent flips gates during a trial) - Dockerfile : layers FlyDSL + AITER dev/kimi-K2.5 on top of jhinpan/sglang-k26-mi355x:v0.5.10rc0-rocm720-20260420 - bench_flydsl_k26.sh : single bench script emitting one canonical line: output_throughput_tok_s: <v> | concurrency=40 in=10240 out=512 decode_bs1_in8k=<guard> - task_description.md : full phased plan A-G, env-var reference table, BS=1 decode guard (>= 0.98x PR #23381 ~38.05 tok/s), front-loaded supervisor-visible gating block - test_harness.py : GSM8K accuracy gate (lm_eval limit=50 >= 0.90) - metadata.json : links to GH issue amdpilot-org/sglang#2 + PR #23381 - ISSUE.md : condensed body as filed on amdpilot-org/sglang#2 Tracks GitHub issue amdpilot-org/sglang#2.
This was referenced Apr 29, 2026
jhinpan
added a commit
to amdpilot-org/amdpilot-evals
that referenced
this pull request
May 2, 2026
* add primus-qwen3-30b-mfu eval instance for issue #1 8x MI355X Qwen3-30B-A3B MoE pretraining MFU optimization. Built on ghcr.io/amdpilot-org/primus-mi355x-ready:v1. Two task.yaml variants for A/B testing the Phase 1 baseline agent: - task.yaml: phase1_baseline: true - task_nophase1.yaml: phase1_baseline: false Both pin the executor to Kimi-K2.6 at 10.235.24.154:30000. * fix(primus): use primus-mi355x-flat:v1 as base + install uv The previous base (ghcr.io/amdpilot-org/primus-mi355x-ready:v1) was a slimmed copy that did NOT include /workspace/primus_train/Primus. Phase 1 spent its full max_turns budget trying to bootstrap from scratch. Switch to primus-mi355x-flat:v1 (locally available, has Primus + Primus-Turbo pre-installed and patched for triton 3.4.0). Also install uv at /root/.local/bin in our Dockerfile so the kimi-cli runtime's source $HOME/.local/bin/env succeeds — without this, executor trials exit 137 immediately after image switch. Tag the new image primus-qwen3-30b-mfu-base:v1 so amdpilot triggers a build the first time it runs. * fix(primus): add gfx950 env vars + runtime IFNAME detection Two cumulative fixes for n08-09 (8x MI355X): 1. Without PYTORCH_ROCM_ARCH=gfx950 (and AITER_ROCM_ARCH, HSA_NO_SCRATCH_RECLAIM, HIP_FORCE_DEV_KERNARG, etc) the torch HIP runtime can't dispatch kernels and benchmarks die immediately with hipErrorInvalidDeviceFunction. These were set in xiao/baizhou's working containers but missing from amdpilot's docker run line. Add to both Dockerfile ENV and task.yaml container.env so they apply via either path. 2. Replace /workspace/detect_interface.sh with a /proc-based detector (the original needed `ip` from iproute2, unavailable in the slim base). bench_mfu.sh now auto-detects GLOO/NCCL socket IFNAME at runtime if bench_config.env doesn't pin one — without this, Megatron's distributed init fails fast (3-5s) before training starts. * feat(primus): add phase1_publish config to task.yaml Enables tag + push of the Phase 1 baseline image to docker.io/jhinpan/primus-qwen3-30b-mfu-phase1 after a successful phase1 commit. Template: {date}-{metric} (e.g. 20260422-278p80) + :latest. Other nodes can then `docker pull` that tag and skip Phase 1 entirely. * feat(primus): enable phase1_publish to ghcr.io/amdpilot-org Switch repository from docker.io/jhinpan to ghcr.io/amdpilot-org so all nodes in the org can pull the verified phase1-baseline image directly. Bump push timeout_s to 9000 (2.5h) to absorb the one-time 42 GB base-layer seed; subsequent pushes only upload the phase1-commit delta (~500 MB - 1 GB) via GHCR cross-repo mount. * feat(flydsl): sglang-kimi-k26-flydsl-mi355x eval instance Adds a hand-authored eval instance that targets the FlyDSL fused-MoE port from AMD's Kimi-K2.5 MI300X optimization blog to Kimi-K2.6 on 4x MI355X, on top of the baseline established by sgl-project/sglang#23381. Contents: - task.yaml : optimize-type task, phase1_baseline=true, frontier_model=true, TP=4, 4xMI355X, FlyDSL env OFF at start so Phase 1 reproduces PR #23381 numbers (DSL2_ROOT + MLIR_PATH + CK_TILE_FLOAT_TO_BFLOAT16 pre-set so import works as soon as agent flips gates during a trial) - Dockerfile : layers FlyDSL + AITER dev/kimi-K2.5 on top of jhinpan/sglang-k26-mi355x:v0.5.10rc0-rocm720-20260420 - bench_flydsl_k26.sh : single bench script emitting one canonical line: output_throughput_tok_s: <v> | concurrency=40 in=10240 out=512 decode_bs1_in8k=<guard> - task_description.md : full phased plan A-G, env-var reference table, BS=1 decode guard (>= 0.98x PR #23381 ~38.05 tok/s), front-loaded supervisor-visible gating block - test_harness.py : GSM8K accuracy gate (lm_eval limit=50 >= 0.90) - metadata.json : links to GH issue amdpilot-org/sglang#2 + PR #23381 - ISSUE.md : condensed body as filed on amdpilot-org/sglang#2 Tracks GitHub issue amdpilot-org/sglang#2. * fix(evals): align flydsl eval metadata and runtime env Keep FlyDSL's runtime MLIR path consistent with the Dockerfile and make new eval metadata consumable by the registry tooling. Mirror the Primus no-phase1 runtime env so the control variant does not hit known HIP temp path failures.
This was referenced May 5, 2026
This was referenced Jun 4, 2026
# Conflicts: # python/sglang/srt/environ.py
jhinpan
requested review from
JustinTong0323,
wisclmy0611 and
zijiexia
as code owners
June 9, 2026 15:13
Collaborator
Author
|
/tag-run-ci-label |
Collaborator
Author
|
/rerun-failed-ci |
Collaborator
Author
|
/tag-and-rerun-ci extra |
Collaborator
|
This is a simple PR. I have removed the |
This was referenced Jun 10, 2026
This was referenced Jun 20, 2026
This was referenced Jul 5, 2026
Closed
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR unblocks Kimi-K2.6 W4A16 MoE serving on AMD Instinct MI355X without downstream forks.
Changes:
E=384, N=256, device_name=AMD_Instinct_MI355X, dtype=int4_w4a16underpython/sglang/srt/layers/moe/moe_runner/triton_utils/configs/triton_3_6_0/.[1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 8192].SGLANG_DEEPSEEK_LOAD_MAX_WORKERSas an opt-in cap for the DeepSeek-style checkpoint loader thread pool. Default isNone, so unset preserves PythonThreadPoolExecutordefault behavior; large MoE deployments can set a smaller value such as4to reduce aggregate host I/O pressure across ranks.0inside Ray HIP workers.main.Benchmark Evidence
Validated on 4x AMD Instinct MI355X with Kimi-K2.6, Triton 3.6.0, ROCm 7.2,
--decode-attention-backend triton,--prefill-attention-backend aiter, and CUDA graph enabled.sglang.bench_one_batch_server --batch-size 1 --output-len 1024:A/B against the same launch command with only the new MoE config file removed:
The benchmark image used for the serving validation was built from this PR branch before the final current-main merge/docs/lint refresh; those final changes do not alter the tuned config or serving path.
Recommended Runtime Setting
For very large MoE checkpoints on NFS-backed storage, use an explicit cap such as:
export SGLANG_DEEPSEEK_LOAD_MAX_WORKERS=4Leaving the variable unset preserves existing SGLang behavior.
Test Plan
sgl-project/sglang:maininto the PR branch and resolved the conflict inpython/sglang/srt/environ.py.python -m ruff check benchmark/kernels/fused_moe_triton/tuning_fused_moe_triton.py python/sglang/srt/environ.py python/sglang/srt/models/deepseek_common/deepseek_weight_loader.pypython -m ruff format --check benchmark/kernels/fused_moe_triton/tuning_fused_moe_triton.py python/sglang/srt/environ.py python/sglang/srt/models/deepseek_common/deepseek_weight_loader.pypython -m py_compile python/sglang/srt/environ.py python/sglang/srt/models/deepseek_common/deepseek_weight_loader.py benchmark/kernels/fused_moe_triton/tuning_fused_moe_triton.pypython -m json.tool python/sglang/srt/layers/moe/moe_runner/triton_utils/configs/triton_3_6_0/E=384,N=256,device_name=AMD_Instinct_MI355X,dtype=int4_w4a16.jsonPYTHONPATH=pythonsmoke confirmsenvs.SGLANG_DEEPSEEK_LOAD_MAX_WORKERSdefaults toNoneand parses4throughEnvInt.override./healthreturned 200 and one chat completion round trip worked.sglang.bench_one_batch_server --batch-size 1 --input-len 1024 2048 4096 8192 16384 32768 --output-len 1024produced the table above.CI States
Latest PR Test (Base): ❌ Run #27216334917
Latest PR Test (Extra): 🚫 Run #27216335222