Skip to content

fix: make the 512gb tier's glm-5.2 actually serve - #19

Merged
jerrytalton merged 3 commits into
mainfrom
fix/512gb-glm52
Jul 9, 2026
Merged

jerrytalton merged 3 commits into
mainfrom
fix/512gb-glm52

Conversation

@jerrytalton

Copy link
Copy Markdown
Owner

Problem

v1.2.0's 512gb tier points general/reasoning/long_context/translation at glm-5.2, but the model had never loaded on the 512GB Studio: mlx-lm (all released versions and main) can't load GLM-5.2's shared-indexer layout (config.indexer_types) and fails strict load with "Missing 285 parameters". The release gate didn't catch it because the correctness tests skip when a model isn't pulled — and the release was cut on a machine that can't hold the 390GB checkpoint. On top of that, two mlx-openai-server bugs (still present in 1.8.1) meant even a loadable 390GB on-demand model couldn't stay resident: a hardcoded 300s readiness timeout, and a warm-request fast path that bypasses refcounting so the idle timer unloads the model every 300s regardless of traffic — thrash-reloading 390GB every ~5 minutes against the menu bar's 240s keep-warm pings.

Changes

  • bin/apply-mlx-glm52-patch.sh — idempotent script that ports the two model files from unmerged upstream PR GLM-5.2: full/shared indexer typing for glm_moe_dsa (DSA schedule + interleaved indexer rope) ml-explore/mlx-lm#1463 onto the installed mlx-lm 0.31.1 (pinned to the PR head sha, fails closed if the ref moves; installing the PR branch wholesale is blocked by the mlx ≥0.31.2 stream-hang, mlx-lm#1256), raises the readiness timeout to 1800s, and fixes the warm-request refcount bypass. Self-disables once upstream ships indexer_types support. install.sh runs it automatically on 512GB machines.
  • app/profile-server.py — chat read timeout 300→600s; a cold frontier-model load stacking with a thinking-model generation exceeded 300s on valid requests.
  • tests/test_tools_smoke_*.py — tier model maps now derive from DEFAULT_PROFILES instead of hardcoded copies that drifted twice.
  • docs — troubleshooting entry (symptoms, root causes, fix, single-on-demand-slot limitation and the whisper-static mitigation) and CLAUDE.md notes.

Verification (on the 512GB Studio)

  • glm-5.2 strict-loads clean and answers correctly; PR #1463's unit tests pass against the ported files.
  • Long-context probe at 5,108 prompt tokens (past the index_topk=2048 sparse-attention engagement point where mlx-lm#1453 reported gibberish): correct recall from start/middle/end.
  • Warm-keeping verified: idle timer re-arms on every 240s keep-warm ping; no more unload/reload cycle.
  • Full correctness suite: 13/13 (one image_gen Metal GPU-timeout flake reproduced only while a 390GB reload was in flight — the thrash this PR eliminates — and passed cleanly on retry).
  • Patch script exercised through full path, fast path, and idempotent re-run; installed files byte-identical to the PR head.

🤖 Generated with Claude Code

jerrytalton and others added 3 commits July 8, 2026 21:46
…ding

The hardcoded copies drifted from the shipped presets twice — most
recently 0e3fefb updated vision and tts in lib/models.py only, leaving
the 512gb smoke map pointing at qwen3.5:122b (never a served model) and
the retired fishaudio TTS. Deriving from DEFAULT_PROFILES makes the
smoke suites always exercise what the presets actually ship.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r cold loads

A cold on-demand load of glm-5.2 (390GB, 80-140s) plus a thinking-model
generation exceeds the 300s read timeout, so /api/test translate failed
while MLX was still working on a valid request.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…b tier

GLM-5.2's glm_moe_dsa arch puts DSA indexer weights on 21 of 78 layers
(config.indexer_types); released mlx-lm builds an indexer per layer and
fails strict load with 'Missing 285 parameters'. The fix (upstream PR
ml-explore/mlx-lm#1463, unmerged) needs mlx>=0.31.2, which still has the
thread-local-stream hang (#1256) — so bin/apply-mlx-glm52-patch.sh ports
the PR's two model files onto the installed mlx-lm 0.31.1, pinned to the
PR head sha and failing closed if the ref moves.

The script also fixes two mlx-openai-server bugs (present through 1.8.1)
that break a 390GB on-demand model in practice: the hardcoded 300s
handler-readiness timeout, and the warm-request refcount bypass that let
the idle timer unload the model every 300s regardless of traffic —
thrash-reloading 390GB every ~5 minutes against the menu bar's 240s
keep-warm pings.

install.sh applies it automatically on 512GB machines. Idempotent;
self-disables once upstream mlx-lm ships indexer_types support.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@jerrytalton
jerrytalton merged commit f8a3637 into main Jul 9, 2026
@jerrytalton
jerrytalton deleted the fix/512gb-glm52 branch July 9, 2026 04:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant