fix(qwen4-exp): keep text inference on mlx-vlm path - #741
Merged
Merged
Conversation
janhilgard
approved these changes
Aug 29, 2026
janhilgard
left a comment
Collaborator
There was a problem hiding this comment.
Correct fix for the right reason. Returning None so SimpleEngine stays on the loaded mlx-vlm model is better than the alternative here: a mechanically compatible skeleton that loads cleanly and emits wrong logits is far harder to diagnose than a model that refuses the fast path, because nothing fails — the output is just quietly worse.
Two details I checked rather than assumed:
str.startswithaccepts a tuple, so_VLM_ONLY_TEXT_MODEL_PREFIXESworks as written and coversqwen4_exp_textfrom the single entry — the parametrised test pins both spellings, which is what makes the prefix form safe to extend later- the check sits before
_import_text_model_classes, so an unregistered architecture short-circuits instead of falling through to theqwen3_5default that #686 made prefix-based
The logger.info naming the model type matters more than it looks: this path silently changes which engine serves text, and without that line the only symptom would be a throughput difference nobody attributes to dispatch.
Approving.
waybarrios
approved these changes
Sep 3, 2026
Owner
|
all set and ready to go |
TimotejLabsky
added a commit
to TimotejLabsky/vllm-mlx
that referenced
this pull request
Sep 23, 2026
…aybarrios#740 implicit-<think> on latch sites, system-KV index pin Rebase onto waybarrios/vllm-mlx ec8e493 (11 commits past 22efb47): 172 fork commits replayed, 0 dropped, 4 conflict stops. Outside the four resolutions the rebase's net delta matches upstream's window hunk-for-hunk, so this commit carries only what the replay could not: - waybarrios#741 (761abfe): re-add the qwen4_exp VLM-only text-path guard, lost when the text_model_from_vlm.py stops were resolved to the fork's candidate-chain dispatch. Load-bearing for waybarrios#97 once mlx-lm ships a qwen4_exp module (resident PLE table instead of the SSD-backed one). - waybarrios#740 (d843a76): the Responses/Anthropic streaming latch sites (#27/waybarrios#47) build their reasoning parser directly, so they now pass implicit_mode from _detect_implicit_thinking, gated on thinking ON like upstream. GLM-4.7-Flash opens <think> in its generation prompt, so this changes that route's streamed reasoning split. One test stub updated to the new reset_state signature. - waybarrios#740: SSDIndex._SCHEMA_VERSION 1 -> 2 would purge every system-KV SSD spill on first start although our entries don't use its serializers. Pin the system-KV store's index at 1 via _SystemKVIndex. - waybarrios#729 (d80db21) partially retires patch waybarrios#95 (bench_command site stays). PATCHES.md rebase note + README base pin updated. Suite 3601 passed / 31 skipped / 30 deselected; ruff clean. New tests mutation-checked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TimotejLabsky
added a commit
to TimotejLabsky/vllm-mlx
that referenced
this pull request
Sep 25, 2026
…aybarrios#740 implicit-<think> on latch sites, system-KV index pin Rebase onto waybarrios/vllm-mlx ec8e493 (11 commits past 22efb47): 172 fork commits replayed, 0 dropped, 4 conflict stops. Outside the four resolutions the rebase's net delta matches upstream's window hunk-for-hunk, so this commit carries only what the replay could not: - waybarrios#741 (761abfe): re-add the qwen4_exp VLM-only text-path guard, lost when the text_model_from_vlm.py stops were resolved to the fork's candidate-chain dispatch. Load-bearing for waybarrios#97 once mlx-lm ships a qwen4_exp module (resident PLE table instead of the SSD-backed one). - waybarrios#740 (d843a76): the Responses/Anthropic streaming latch sites (#27/waybarrios#47) build their reasoning parser directly, so they now pass implicit_mode from _detect_implicit_thinking, gated on thinking ON like upstream. GLM-4.7-Flash opens <think> in its generation prompt, so this changes that route's streamed reasoning split. One test stub updated to the new reset_state signature. - waybarrios#740: SSDIndex._SCHEMA_VERSION 1 -> 2 would purge every system-KV SSD spill on first start although our entries don't use its serializers. Pin the system-KV store's index at 1 via _SystemKVIndex. - waybarrios#729 (d80db21) partially retires patch waybarrios#95 (bench_command site stays). PATCHES.md rebase note + README base pin updated. Suite 3601 passed / 31 skipped / 30 deselected; ruff clean. New tests mutation-checked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Qwen4-Exp is not compatible with the generic Qwen3.5 TextModel fallback. Keep text requests on the loaded mlx-vlm language model instead of constructing a mechanically compatible model that produces incorrect logits.
Checks: