fix: select the text-model family by prefix instead of one exact string - #686
Merged
waybarrios merged 2 commits intoAug 11, 2026
Merged
Conversation
`_import_text_model_classes` matched `model_type == "gemma4_text"` exactly and
sent everything else to `qwen3_5.TextModel`. Gemma 4 reports `gemma4_text` on
some checkpoints and `gemma4_unified_text` on others, so the latter reached
`qwen3_5.TextModelArgs`, which leaves `num_experts` as None, and construction
died on `args.num_experts > 0`:
TypeError: '>' not supported between instances of 'NoneType' and 'int'
A wrong guess does not fail where it is made. It fails deep inside the chosen
constructor with an error naming neither the model nor the class,
`build_text_model` catches it and returns None, and the engine goes on to
report itself loaded with `_text_model=None` — a route quietly losing its
backend while the log shows one line that reads like a warning.
Families are now matched by prefix, longest first, so one entry covers a
family's variants. The generic fallback stays: `qwen3_5.TextModel` handles
dense and MoE natively, and making unknown types raise would be a regression
for every model that works through it today. It is now logged rather than
silent, and a failed build names both the `model_type` and the class that was
selected, with the traceback preserved.
Verified on `gemma-4-12B-it` (`text_config.model_type: gemma4_unified_text`,
mlx_vlm conversion, M3 Ultra): the TextModel now builds as
`mlx_lm.models.gemma4_text` and the error is gone.
Scope note, since waybarrios#685 reports two symptoms: this fixes the dispatch, and it
does **not** by itself change the empty-chunk symptom on
`stream_generate(prompt=...)`. That route calls the mlx_vlm path regardless of
whether a TextModel exists, and the raw prompt is the real problem there — I
measured `mlx_lm` driving the correctly-built Gemma 4 TextModel with the same
raw prompt and it produces the same garbage (`'-......1.1___-______'`) as the
mlx_vlm path. So it is an instruct model being handed a bare completion
prompt, which is what `/v1/completions` is meant to do, not a routing bug.
Applying a chat template inside `stream_generate` would break that endpoint's
semantics, so I have left it alone and corrected the issue.
Repo suite: 2305 passed. Confirmed by mutation — restoring the exact-match
dispatch fails four of the new tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Thump604
approved these changes
Aug 8, 2026
Thump604
left a comment
Collaborator
There was a problem hiding this comment.
This is the right split. The family dispatch now handles Gemma 4 variants without changing the existing fallback, and failures identify the selected class with a traceback. Focused dispatch/build tests pass at the exact head. The remaining blank-chunk behavior stays correctly scoped to #685.
Owner
|
I noticed the new test file was missing from the Apple Silicon CI list, so I added it. This way the Gemma dispatch tests will actually run on every PR. Small oversight, but it’s covered now. |
TimotejLabsky
added a commit
to TimotejLabsky/vllm-mlx
that referenced
this pull request
Aug 18, 2026
…ng, wanted deltas, stream-global restore engine/simple.py was kept wholesale over upstream's c65c356/d7c1c98 rewrites (PATCHES.md 2026-08-17 rebase note). This commit restores the upstream deltas we do want and reconciles the plumbing that assumed upstream's engine: - accept upstream waybarrios#574's prefix_trie_cache* kwargs fail-closed (raise if enabled — the fork's system-KV supersedes the trie; no silent no-op) - upstream waybarrios#681: stamp finish_reason=length at the token-budget cutoff when the backend chunk carries None - upstream waybarrios#686 folded into fork semantics: _import_text_model_classes candidate chain with gemma4/qwen3 family rules; unknown families raise instead of guessing qwen3_5 - engine_core: keep upstream's owns_worker teardown guard AND the fork's waybarrios#49 SSD flush; getattr-guard close_ssd_tier for duck-typed schedulers - fix(streams): snapshot the pre-bind generation-stream globals at first worker bind and restore them in SimpleEngine.stop() — a retired worker otherwise leaves mlx_lm/mlx_vlm generation_stream naming a dead thread's stream (surfaced by upstream waybarrios#702's parity tests) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TimotejLabsky
added a commit
to TimotejLabsky/vllm-mlx
that referenced
this pull request
Aug 23, 2026
…ng, wanted deltas, stream-global restore engine/simple.py was kept wholesale over upstream's c65c356/d7c1c98 rewrites (PATCHES.md 2026-08-17 rebase note). This commit restores the upstream deltas we do want and reconciles the plumbing that assumed upstream's engine: - accept upstream waybarrios#574's prefix_trie_cache* kwargs fail-closed (raise if enabled — the fork's system-KV supersedes the trie; no silent no-op) - upstream waybarrios#681: stamp finish_reason=length at the token-budget cutoff when the backend chunk carries None - upstream waybarrios#686 folded into fork semantics: _import_text_model_classes candidate chain with gemma4/qwen3 family rules; unknown families raise instead of guessing qwen3_5 - engine_core: keep upstream's owns_worker teardown guard AND the fork's waybarrios#49 SSD flush; getattr-guard close_ssd_tier for duck-typed schedulers - fix(streams): snapshot the pre-bind generation-stream globals at first worker bind and restore them in SimpleEngine.stop() — a retired worker otherwise leaves mlx_lm/mlx_vlm generation_stream naming a dead thread's stream (surfaced by upstream waybarrios#702's parity tests) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TimotejLabsky
added a commit
to TimotejLabsky/vllm-mlx
that referenced
this pull request
Aug 27, 2026
…ng, wanted deltas, stream-global restore engine/simple.py was kept wholesale over upstream's c65c356/d7c1c98 rewrites (PATCHES.md 2026-08-17 rebase note). This commit restores the upstream deltas we do want and reconciles the plumbing that assumed upstream's engine: - accept upstream waybarrios#574's prefix_trie_cache* kwargs fail-closed (raise if enabled — the fork's system-KV supersedes the trie; no silent no-op) - upstream waybarrios#681: stamp finish_reason=length at the token-budget cutoff when the backend chunk carries None - upstream waybarrios#686 folded into fork semantics: _import_text_model_classes candidate chain with gemma4/qwen3 family rules; unknown families raise instead of guessing qwen3_5 - engine_core: keep upstream's owns_worker teardown guard AND the fork's waybarrios#49 SSD flush; getattr-guard close_ssd_tier for duck-typed schedulers - fix(streams): snapshot the pre-bind generation-stream globals at first worker bind and restore them in SimpleEngine.stop() — a retired worker otherwise leaves mlx_lm/mlx_vlm generation_stream naming a dead thread's stream (surfaced by upstream waybarrios#702's parity tests) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TimotejLabsky
added a commit
to TimotejLabsky/vllm-mlx
that referenced
this pull request
Sep 23, 2026
…ng, wanted deltas, stream-global restore engine/simple.py was kept wholesale over upstream's c65c356/d7c1c98 rewrites (PATCHES.md 2026-08-17 rebase note). This commit restores the upstream deltas we do want and reconciles the plumbing that assumed upstream's engine: - accept upstream waybarrios#574's prefix_trie_cache* kwargs fail-closed (raise if enabled — the fork's system-KV supersedes the trie; no silent no-op) - upstream waybarrios#681: stamp finish_reason=length at the token-budget cutoff when the backend chunk carries None - upstream waybarrios#686 folded into fork semantics: _import_text_model_classes candidate chain with gemma4/qwen3 family rules; unknown families raise instead of guessing qwen3_5 - engine_core: keep upstream's owns_worker teardown guard AND the fork's waybarrios#49 SSD flush; getattr-guard close_ssd_tier for duck-typed schedulers - fix(streams): snapshot the pre-bind generation-stream globals at first worker bind and restore them in SimpleEngine.stop() — a retired worker otherwise leaves mlx_lm/mlx_vlm generation_stream naming a dead thread's stream (surfaced by upstream waybarrios#702's parity tests) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TimotejLabsky
added a commit
to TimotejLabsky/vllm-mlx
that referenced
this pull request
Sep 25, 2026
…ng, wanted deltas, stream-global restore engine/simple.py was kept wholesale over upstream's c65c356/d7c1c98 rewrites (PATCHES.md 2026-08-17 rebase note). This commit restores the upstream deltas we do want and reconciles the plumbing that assumed upstream's engine: - accept upstream waybarrios#574's prefix_trie_cache* kwargs fail-closed (raise if enabled — the fork's system-KV supersedes the trie; no silent no-op) - upstream waybarrios#681: stamp finish_reason=length at the token-budget cutoff when the backend chunk carries None - upstream waybarrios#686 folded into fork semantics: _import_text_model_classes candidate chain with gemma4/qwen3 family rules; unknown families raise instead of guessing qwen3_5 - engine_core: keep upstream's owns_worker teardown guard AND the fork's waybarrios#49 SSD flush; getattr-guard close_ssd_tier for duck-typed schedulers - fix(streams): snapshot the pre-bind generation-stream globals at first worker bind and restore them in SimpleEngine.stop() — a retired worker otherwise leaves mlx_lm/mlx_vlm generation_stream naming a dead thread's stream (surfaced by upstream waybarrios#702's parity tests) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the first half of #685.
Problem
_import_text_model_classesmatchedmodel_type == "gemma4_text"exactly and defaulted everything else toqwen3_5.TextModel:Gemma 4 reports
gemma4_texton some checkpoints andgemma4_unified_texton others, so the latter reachedqwen3_5.TextModelArgs, which leavesnum_expertsas None, and construction died onargs.num_experts > 0:A wrong guess does not fail where it is made. It fails deep inside the chosen constructor with an error naming neither the model nor the class,
build_text_modelcatches it and returns None, and the engine then reports itself loaded with_text_model=None. The whole visible trace is:That reads like a warning rather than a route losing its backend, which is why it sat there.
Changes
qwen3_5.TextModelhandles dense and MoE natively, and making unknown types raise would be a regression for every model that works through it today — but the fallback is now logged instead of silent.model_typeand the class that was selected, withexc_info=Trueso the traceback survives.Verification
gemma-4-12B-it(text_config.model_type: gemma4_unified_text, mlx_vlm conversion, M3 Ultra):build_text_modelNone+ TypeErrorengine._text_modelNonemlx_lm.models.gemma4_textRepo suite: 2305 passed (four local failures reproduce on pristine
main— three Python 3.14, one missingffmpeg). Confirmed by mutation: restoring the exact-match dispatch fails four of the new tests, including thegemma4_unified_textcase specifically.Scope: this does not fix the other half of #685, and I want to be clear about why
#685 also reports
stream_generate(prompt=...)returning empty chunks, and in the issue I guessed that the missing chat template on the mlx_vlm fallback was to blame. That guess was wrong and I have corrected the issue._stream_generate_implcalls the mlx_vlm path regardless of whether a TextModel exists, so a working TextModel does not change that route. More to the point, I drovemlx_lmagainst the now-correctly-built Gemma 4 TextModel with the same raw prompt:'-..1.1______''-......1.1___-______''The three primary colors are:\n\n1. **Red**'Both backends agree, so this is an instruct-tuned model being handed a bare completion prompt — which is exactly what
/v1/completionsis for. Applying a chat template insidestream_generatewould break that endpoint's semantics for every VLM, so I have deliberately left it alone.The one thing I still cannot explain is why the engine surfaces empty text where a direct call on the same model instance surfaces the garbage above; the leading chunks are empty in both and the pump stops at
completion_tokens >= max_tokens, which accounts for part of it but not the differing chunk counts. Both are wrong outputs from a wrong-shaped prompt, so I have left that noted in #685 rather than guessing at a fix.