Skip to content

[GLM-5.2] Fix GLM-5.2 AMD image pin and mark MI325X verified - #1036

Merged
esmeetu merged 1 commit into
vllm-project:mainfrom
jin-amd:update-glm-5.2-rocm-mi325x
Sep 28, 2026
Merged

esmeetu merged 1 commit into
vllm-project:mainfrom
jin-amd:update-glm-5.2-rocm-mi325x

Conversation

@jin-amd

@jin-amd jin-amd commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

What this changes

Two updates to models/zai-org/GLM-5.2.yaml:

  1. AMD image. docker_image.amd pointed at vllm/vllm-openai-rocm:nightly-4c626633159887b0f2c962058c17c78f1434556d, which no longer exists on Docker Hub; nightly-<sha> tags are pruned after about a week. It now points at vllm/vllm-openai-rocm:nightly, as most recipes do.
  2. MI325X. Marked mi325x: verified, added MI325X to the AMD FP8 section title, and noted what the command achieves on 8x MI325X.

The serve command is unchanged.

Validation on 8x MI325X

Ran the AMD FP8 command as written:

  • Full --max-model-len 524288 at --gpu-memory-utilization 0.80
  • KV cache 1,564,016 tokens, 2.98x maximum concurrency
  • MTP mean acceptance length 4.00 of 5 drafted tokens
  • --reasoning-parser glm47 populates reasoning, and chat_template_kwargs: {"enable_thinking": false} returns a plain answer with reasoning_tokens: 0

@vercel

vercel Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
vllm-recipes Ready Ready Preview Sep 28, 2026 8:09am UTC

Request Review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the GLM-5.2 model configuration to verify compatibility with AMD MI325X GPUs and pins the AMD Docker image to v0.30.0 to avoid regressions introduced in newer ROCm builds. It also adds documentation regarding performance on MI325X and troubleshooting instructions for issues with newer vLLM builds. The review feedback correctly points out that the file paths referenced in the troubleshooting section should be updated to reflect their actual location under 'vllm/model_executor/models/' instead of 'vllm/models/'.

Comment thread models/zai-org/GLM-5.2.yaml Outdated
@jin-amd
jin-amd marked this pull request as ready for review September 27, 2026 12:05
Copilot AI lite review requested due to automatic review settings September 27, 2026 12:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Correct the inaccurate ROCm warning and validate the pinned image before claiming MI325X verification.

Review effort: Lite
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

Updates the GLM-5.2 AMD recipe to use a stable ROCm image and document MI325X support.

Changes:

  • Pins the AMD image to v0.30.0.
  • Marks MI325X verified and expands AMD guidance.
  • Adds ROCm troubleshooting notes.
File Summary
models/​zai-org/​GLM-5.2.yaml Updates the AMD image, hardware metadata, guide, and troubleshooting documentation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread models/zai-org/GLM-5.2.yaml Outdated
Comment thread models/zai-org/GLM-5.2.yaml
The pinned AMD image tag nightly-4c626633159887b0f2c962058c17c78f1434556d
no longer resolves on Docker Hub, which prunes nightly-<sha> tags. Point
it at vllm/vllm-openai-rocm:nightly, as most recipes do.

Verified on 8x MI325X: the AMD FP8 command serves the full 524288
context with a 1,564,016-token KV cache and 2.98x maximum concurrency,
and MTP's mean acceptance length is 4.00 of 5.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Jin Tao <jin.tao@amd.com>
@jin-amd
jin-amd force-pushed the update-glm-5.2-rocm-mi325x branch from c690be8 to e7da1ec Compare September 28, 2026 08:07
@jin-amd jin-amd changed the title Fix GLM-5.2 AMD image pin and mark MI325X verified [GLM-5.2]Fix GLM-5.2 AMD image pin and mark MI325X verified Sep 28, 2026
@jin-amd jin-amd changed the title [GLM-5.2]Fix GLM-5.2 AMD image pin and mark MI325X verified [GLM-5.2] Fix GLM-5.2 AMD image pin and mark MI325X verified Sep 28, 2026
@esmeetu
esmeetu merged commit 985e84a into vllm-project:main Sep 28, 2026
4 checks passed

This branch was successfully deployed

1 active deployment
Preview — e7da1ec1 Deployed Sep 28, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants