fix: make the 512gb tier's glm-5.2 actually serve - #19
Merged
Merged
Conversation
…ding The hardcoded copies drifted from the shipped presets twice — most recently 0e3fefb updated vision and tts in lib/models.py only, leaving the 512gb smoke map pointing at qwen3.5:122b (never a served model) and the retired fishaudio TTS. Deriving from DEFAULT_PROFILES makes the smoke suites always exercise what the presets actually ship. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r cold loads A cold on-demand load of glm-5.2 (390GB, 80-140s) plus a thinking-model generation exceeds the 300s read timeout, so /api/test translate failed while MLX was still working on a valid request. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…b tier GLM-5.2's glm_moe_dsa arch puts DSA indexer weights on 21 of 78 layers (config.indexer_types); released mlx-lm builds an indexer per layer and fails strict load with 'Missing 285 parameters'. The fix (upstream PR ml-explore/mlx-lm#1463, unmerged) needs mlx>=0.31.2, which still has the thread-local-stream hang (#1256) — so bin/apply-mlx-glm52-patch.sh ports the PR's two model files onto the installed mlx-lm 0.31.1, pinned to the PR head sha and failing closed if the ref moves. The script also fixes two mlx-openai-server bugs (present through 1.8.1) that break a 390GB on-demand model in practice: the hardcoded 300s handler-readiness timeout, and the warm-request refcount bypass that let the idle timer unload the model every 300s regardless of traffic — thrash-reloading 390GB every ~5 minutes against the menu bar's 240s keep-warm pings. install.sh applies it automatically on 512GB machines. Idempotent; self-disables once upstream mlx-lm ships indexer_types support. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
jerrytalton
force-pushed
the
fix/512gb-glm52
branch
from
July 9, 2026 04:19
9b55a5c to
f97b166
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
v1.2.0's 512gb tier points general/reasoning/long_context/translation at
glm-5.2, but the model had never loaded on the 512GB Studio: mlx-lm (all released versions and main) can't load GLM-5.2's shared-indexer layout (config.indexer_types) and fails strict load with "Missing 285 parameters". The release gate didn't catch it because the correctness tests skip when a model isn't pulled — and the release was cut on a machine that can't hold the 390GB checkpoint. On top of that, two mlx-openai-server bugs (still present in 1.8.1) meant even a loadable 390GB on-demand model couldn't stay resident: a hardcoded 300s readiness timeout, and a warm-request fast path that bypasses refcounting so the idle timer unloads the model every 300s regardless of traffic — thrash-reloading 390GB every ~5 minutes against the menu bar's 240s keep-warm pings.Changes
bin/apply-mlx-glm52-patch.sh— idempotent script that ports the two model files from unmerged upstream PR GLM-5.2: full/shared indexer typing for glm_moe_dsa (DSA schedule + interleaved indexer rope) ml-explore/mlx-lm#1463 onto the installed mlx-lm 0.31.1 (pinned to the PR head sha, fails closed if the ref moves; installing the PR branch wholesale is blocked by the mlx ≥0.31.2 stream-hang, mlx-lm#1256), raises the readiness timeout to 1800s, and fixes the warm-request refcount bypass. Self-disables once upstream shipsindexer_typessupport.install.shruns it automatically on 512GB machines.app/profile-server.py— chat read timeout 300→600s; a cold frontier-model load stacking with a thinking-model generation exceeded 300s on valid requests.tests/test_tools_smoke_*.py— tier model maps now derive fromDEFAULT_PROFILESinstead of hardcoded copies that drifted twice.Verification (on the 512GB Studio)
index_topk=2048sparse-attention engagement point where mlx-lm#1453 reported gibberish): correct recall from start/middle/end.🤖 Generated with Claude Code