Repository navigation
[GLM-5.2] Fix GLM-5.2 AMD image pin and mark MI325X verified - #1036
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
Code Review
This pull request updates the GLM-5.2 model configuration to verify compatibility with AMD MI325X GPUs and pins the AMD Docker image to v0.30.0 to avoid regressions introduced in newer ROCm builds. It also adds documentation regarding performance on MI325X and troubleshooting instructions for issues with newer vLLM builds. The review feedback correctly points out that the file paths referenced in the troubleshooting section should be updated to reflect their actual location under 'vllm/model_executor/models/' instead of 'vllm/models/'.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Correct the inaccurate ROCm warning and validate the pinned image before claiming MI325X verification.
Review effort: Lite
Findings: 1
Open (2)
What changed in this PR
Updates the GLM-5.2 AMD recipe to use a stable ROCm image and document MI325X support.
Changes:
- Pins the AMD image to
v0.30.0. - Marks MI325X verified and expands AMD guidance.
- Adds ROCm troubleshooting notes.
| File | Summary |
|---|---|
models/zai-org/GLM-5.2.yaml |
Updates the AMD image, hardware metadata, guide, and troubleshooting documentation. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
9487a32 to
c690be8
Compare
The pinned AMD image tag nightly-4c626633159887b0f2c962058c17c78f1434556d no longer resolves on Docker Hub, which prunes nightly-<sha> tags. Point it at vllm/vllm-openai-rocm:nightly, as most recipes do. Verified on 8x MI325X: the AMD FP8 command serves the full 524288 context with a 1,564,016-token KV cache and 2.98x maximum concurrency, and MTP's mean acceptance length is 4.00 of 5. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Jin Tao <jin.tao@amd.com>
c690be8 to
e7da1ec
Compare


What this changes
Two updates to
models/zai-org/GLM-5.2.yaml:docker_image.amdpointed atvllm/vllm-openai-rocm:nightly-4c626633159887b0f2c962058c17c78f1434556d, which no longer exists on Docker Hub;nightly-<sha>tags are pruned after about a week. It now points atvllm/vllm-openai-rocm:nightly, as most recipes do.mi325x: verified, added MI325X to the AMD FP8 section title, and noted what the command achieves on 8x MI325X.The serve command is unchanged.
Validation on 8x MI325X
Ran the AMD FP8 command as written:
--max-model-len 524288at--gpu-memory-utilization 0.80--reasoning-parser glm47populatesreasoning, andchat_template_kwargs: {"enable_thinking": false}returns a plain answer withreasoning_tokens: 0