Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion benchmarks/single_node/agentic/glm5.2_fp8_mi325x_mtp.sh
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,9 @@ SGLANG_CMD=(
--chunked-prefill-size 131072
--mem-fraction-static 0.85
--max-running-requests "$MAX_RUNNING_REQUESTS"
--cuda-graph-max-bs "$MAX_RUNNING_REQUESTS"
# SGLang v0.5.20 retired the deprecated --cuda-graph-max-bs alias
# (sgl-project/sglang#38375); the decode-phase flag is the same setting.
--cuda-graph-max-bs-decode "$MAX_RUNNING_REQUESTS"
--speculative-algorithm EAGLE
--speculative-num-steps 3
--speculative-eagle-topk 1
Expand Down
2 changes: 1 addition & 1 deletion configs/amd-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1050,7 +1050,7 @@ dsv41flash-fp4-mi300x-vllm-agentic-dspark:
# GPU-resident-KV c1/c2/c3/c4/c5/c6/c8 curve from Actions run 29657732517
# and enables EAGLE MTP with the committed thinking-on golden AL.
glm5.2-fp8-mi325x-sglang-agentic-mtp:
image: lmsysorg/sglang:v0.5.19-rocm720-mi30x
image: lmsysorg/sglang:v0.5.20-rocm720-mi30x
model: zai-org/GLM-5.2-FP8
model-prefix: glm5.2
runner: cluster:mi325x-amds
Expand Down
7 changes: 7 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8931,3 +8931,10 @@
- "Use five-token DSpark with thinking_on golden AL 3.51 for throughput and real acceptance for eval, the shipped draft, and dsml_v41 tool parsing. Replay semianalysis_cc_traces_weka_062126 for 3600 seconds per point with five warmup requests per lane."
- "DSpark draft layers 37-39 keep their checkpoint precision: the recipe passes no --online_quant_config, so online_quant_config resolves to None and make_v4_quant_config builds the same per-layer quant spec for the drafter as for the target. The MTP weights (mtp.0/1/2) load inline from the same checkpoint with spec_decode=True, and config.get_layer_quant_config has no mtp.* special case, so layers 37-39 are neither re-quantized nor up/down-cast by ATOM and retain their on-disk dtype. --kv_cache_dtype bf16 and --index-cache-dtype fp8 touch cache storage only, not draft/MTP weights."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3387

- config-keys:
- glm5.2-fp8-mi325x-sglang-agentic-mtp
description:
- "Update SGLang ROCm image to v0.5.20-rocm720-mi30x; replace retired --cuda-graph-max-bs with --cuda-graph-max-bs-decode."
- "将 SGLang ROCm 镜像更新至 v0.5.20-rocm720-mi30x,并以 --cuda-graph-max-bs-decode 替换已移除的 --cuda-graph-max-bs。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3362
Loading