docs: record why GLM-5.2 and DeepSeek-V4 cannot run yet - #30
Merged
Conversation
Both need model code that no mlx-lm release has. GLM-5.2's IndexShare shares an attention indexer across every four layers, so stock mlx-lm looks for 285 parameters the checkpoint does not carry; the support PR (#1410) is open, not merged. deepseek_v4 has no module at all. Decision: stay on the z.ai cloud for GLM, keep the Developer on Qwen3.6-27B, and revisit when the PRs land. Weights and launchers stay in place; the GLM LaunchAgent is uninstalled so nothing crashloops at login. GLM-5-4bit is noted as a stock-compatible fallback. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes out the two blocked moves from #27 with what was actually measured.
GLM-5.2 downloaded and verified (91/91 shards), the GPU wired-memory ceiling was raised to 491520 MB, and the server started — but the model fails to load on mlx-lm 0.31.3:
GLM-5.2's IndexShare shares one attention indexer across every four layers, so the checkpoint carries indexer weights on 21 of 78 layers while stock mlx-lm builds one per layer (57 × 5 = 285). Support is mlx-lm PR #1410 — open, not merged.
DeepSeek-V4-Flash has no
deepseek_v4module in any mlx-lm release.Decision: stay on cloud GLM, keep the Developer on Qwen3.6-27B, wait for upstream. Weights, serve script, LaunchAgent template and the meter's env flip all stay; the rendered GLM LaunchAgent is uninstalled so nothing crashloops at login.
mlx-community/GLM-5-4bitis recorded as a stock-mlx-lm-compatible fallback (verified: indexer on every layer).🤖 Generated with Claude Code