Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5179,7 +5179,7 @@ qwen3.5-fp8-h200-sglang-agentic-hicache-mtp:
- { tp: 8, ep: 1, spec-decoding: mtp, kv-offloading: dram, kv-offload-backend: { name: hicache }, conc-list: [2, 4, 8, 10, 12, 16, 20, 24] }

qwen3.5-fp8-h100-sglang-agentic-mtp:
image: lmsysorg/sglang:v0.5.16-cu130
image: lmsysorg/sglang:nightly-dev-cu13-20260914-4358a161
model: Qwen/Qwen3.5-397B-A17B-FP8
model-prefix: qwen3.5
runner: cluster:h100-dgxc
Expand Down
7 changes: 7 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -7956,3 +7956,10 @@
- "Update SGLang image from lmsysorg/sglang:v0.5.16-cu130 (build commit sgl-project/sglang@fdebc938f7f4d16fe6b9f55dcd9a767cf0899ea1, CUDA 13.0.1, FlashInfer 0.6.14, sgl-kernel 0.4.5) to the v0.5.19 release image lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, build commit sgl-project/sglang@0bcd822377da7b5718e674eaf9c870d349424dd1, CUDA 13.0.3, FlashInfer 0.6.18, sgl-kernel 0.4.6.post1)."
- "benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh is unchanged: nvidia/DeepSeek-R1-0528-FP4-V2 with modelopt_fp4, trtllm_mla attention, flashinfer_trtllm MoE, fp8_e4m3 KV cache, TP4/EP1 concurrency 1-32 and TP4/EP4 DP-attention concurrency 64-256 on the 8k1k workload."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2989

- config-keys:
- qwen3.5-fp8-h100-sglang-agentic-mtp
description:
- "Update SGLang image from lmsysorg/sglang:v0.5.16-cu130 to lmsysorg/sglang:nightly-dev-cu13-20260914-4358a161 (2026-09-14 upstream cu13 nightly, sgl-project/sglang@4358a161). The 2026-09-15 nightly (nightly-dev-cu13-20260915-8874c51a) cannot start any HiCache arm: sgl-project/sglang#35233 (merged 2026-09-14T22:49Z) made the host pool import get_device_accessible_ptr from a kernel wheel that image does not ship, so every kv-offloading dram point died at startup (runs 35016956476, 35016965515); sgl-project/sglang#39516 fixes it for the 2026-09-16 build, and 09-14 is the newest nightly that predates the regression. NEXTN MTP settings, HiCache arms, and the grid are unchanged"
- "将 SGLang 镜像从 lmsysorg/sglang:v0.5.16-cu130更新为 lmsysorg/sglang:nightly-dev-cu13-20260914-4358a161(2026-09-14 上游 cu13 nightly,sgl-project/sglang@4358a161)。2026-09-15 nightly(nightly-dev-cu13-20260915-8874c51a)的所有 HiCache 分支无法启动:sgl-project/sglang#35233(2026-09-14T22:49Z 合入)令主机池从镜像未附带的内核 wheel 导入 get_device_accessible_ptr,所有 kv-offloading dram 点在启动时失败(运行 35016956476、35016965515);sgl-project/sglang#39516 已修复并将随 2026-09-16 构建发布,09-14 是早于该回归的最新 nightly。NEXTN MTP 设置、HiCache 分支及网格保持不变"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3150