diff --git a/benchmarks/single_node/agentic/glm5.2_fp8_mi325x_mtp.sh b/benchmarks/single_node/agentic/glm5.2_fp8_mi325x_mtp.sh index 1201607c9d..990e110624 100755 --- a/benchmarks/single_node/agentic/glm5.2_fp8_mi325x_mtp.sh +++ b/benchmarks/single_node/agentic/glm5.2_fp8_mi325x_mtp.sh @@ -91,7 +91,9 @@ SGLANG_CMD=( --chunked-prefill-size 131072 --mem-fraction-static 0.85 --max-running-requests "$MAX_RUNNING_REQUESTS" - --cuda-graph-max-bs "$MAX_RUNNING_REQUESTS" + # SGLang v0.5.20 retired the deprecated --cuda-graph-max-bs alias + # (sgl-project/sglang#38375); the decode-phase flag is the same setting. + --cuda-graph-max-bs-decode "$MAX_RUNNING_REQUESTS" --speculative-algorithm EAGLE --speculative-num-steps 3 --speculative-eagle-topk 1 diff --git a/configs/amd-master.yaml b/configs/amd-master.yaml index a92912462a..15397953c8 100644 --- a/configs/amd-master.yaml +++ b/configs/amd-master.yaml @@ -1050,7 +1050,7 @@ dsv41flash-fp4-mi300x-vllm-agentic-dspark: # GPU-resident-KV c1/c2/c3/c4/c5/c6/c8 curve from Actions run 29657732517 # and enables EAGLE MTP with the committed thinking-on golden AL. glm5.2-fp8-mi325x-sglang-agentic-mtp: - image: lmsysorg/sglang:v0.5.19-rocm720-mi30x + image: lmsysorg/sglang:v0.5.20-rocm720-mi30x model: zai-org/GLM-5.2-FP8 model-prefix: glm5.2 runner: cluster:mi325x-amds diff --git a/perf-changelog.yaml b/perf-changelog.yaml index 542a653a1b..3f86f42195 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -8931,3 +8931,10 @@ - "Use five-token DSpark with thinking_on golden AL 3.51 for throughput and real acceptance for eval, the shipped draft, and dsml_v41 tool parsing. Replay semianalysis_cc_traces_weka_062126 for 3600 seconds per point with five warmup requests per lane." - "DSpark draft layers 37-39 keep their checkpoint precision: the recipe passes no --online_quant_config, so online_quant_config resolves to None and make_v4_quant_config builds the same per-layer quant spec for the drafter as for the target. The MTP weights (mtp.0/1/2) load inline from the same checkpoint with spec_decode=True, and config.get_layer_quant_config has no mtp.* special case, so layers 37-39 are neither re-quantized nor up/down-cast by ATOM and retain their on-disk dtype. --kv_cache_dtype bf16 and --index-cache-dtype fp8 touch cache storage only, not draft/MTP weights." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3387 + +- config-keys: + - glm5.2-fp8-mi325x-sglang-agentic-mtp + description: + - "Update SGLang ROCm image to v0.5.20-rocm720-mi30x; replace retired --cuda-graph-max-bs with --cuda-graph-max-bs-decode." + - "将 SGLang ROCm 镜像更新至 v0.5.20-rocm720-mi30x,并以 --cuda-graph-max-bs-decode 替换已移除的 --cuda-graph-max-bs。" + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3362