Skip to content

patches: serve-model-path-match — accept the model path /v1/models advertises - #167

Merged
mhenrichsen merged 1 commit into
syv-ai:mainfrom
TyroneNel:serve-model-path-match
Sep 22, 2026
Merged

mhenrichsen merged 1 commit into
syv-ai:mainfrom
TyroneNel:serve-model-path-match

Conversation

@TyroneNel

@TyroneNel TyroneNel commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

What

Adds patches/serve-model-path-match.patch, its patches/series line and its PATCHES.md row. is_base_model now accepts the served name, the full model path, or the basename of that path, all as exact matches.

Upstream: vllm-project/vllm#58026.

Why

/v1/models publishes the checkpoint path in root, but the name check matched the served name only, so echoing root back returned 404. Measured on a live server from the current image:

model="qwen3.8-27b"                                -> 200
model="/app/models/Qwen3.8-27B-Uncensored-W4A16"   -> 404
model="Qwen3.8-27B-Uncensored-W4A16"               -> 404

This repo is a steady source of that mismatch: --model is a path everywhere and the served name is qwen3.8-27b everywhere, so any client that discovers the model from /v1/models picks the path.

There is no prefix match and no fuzzy match, so an unknown name still returns 404.

Of the five patches, this is the one I would drop first if you would rather keep the name space strict: #165 and #166 already make the failure self-explaining.

Verification

  • patch integrity (the workflow's git-apply job) applies the whole series to a pristine vllm-project/vllm checkout at the pin: passes on this PR.
  • patch -p1 --fuzz 0 --dry-run of this file against the installed tree in ghcr.io/syv-ai/hyperqwen:latest (06150174): applies, no fuzz.
  • verify.sh needs no new entry: its loop reads patches/series, and patches/_check_applied.py parses the patch file itself.

Not done: I have not rebuilt the image and restarted a server on this patch, so the runtime evidence above comes from the code path, not from a rebuilt server.

@TyroneNel
TyroneNel force-pushed the serve-model-path-match branch from 0ea5994 to dbba103 Compare September 21, 2026 22:41
@mhenrichsen
mhenrichsen force-pushed the serve-model-path-match branch 2 times, most recently from a815160 to b51f8e6 Compare September 22, 2026 17:14
…vertises

is_base_model matched the served name only, so a client echoing the root
that /v1/models itself publishes got a 404. The patch accepts the served
name, the full checkpoint path, or the path's basename — exact matches
only. Independent of every other patch, appended at the series' block
boundary. Cut from the extended cpuchip/vllm qwen38/0.28 branch, topic
commit [qwen38] serve-model-path-match; kind: fix, retires when upstream
takes it.
@mhenrichsen
mhenrichsen force-pushed the serve-model-path-match branch from b51f8e6 to c9414df Compare September 22, 2026 17:14
@mhenrichsen

Copy link
Copy Markdown
Contributor

Merged (rebased for you).

I did take your offer to consider dropping this one seriously, since #165 and #166 make the failure self-explaining, and I am keeping it for a reason worth writing down: the mismatch is this repo's doing, not the client's. Every launcher here passes a checkpoint path as --model and qwen3.8-27b as --served-model-name, so /v1/models advertises root as a path that the server then refuses. A client that discovers the model instead of hardcoding it is doing the right thing and gets a 404 for it. A better error message documents that; accepting the name removes it.

Checked against the pinned source: show_available_models builds each card as id=base_model.name, root=base_model.model_path from the same base_model_paths list is_base_model walks, so the three accepted forms are exactly what the server publishes, and the basename is the only one that is not published verbatim. Exact matches, no prefix matching, unknown names still 404 — which is what keeps this from being a name-space loosening.

@mhenrichsen
mhenrichsen merged commit 056c660 into syv-ai:main Sep 22, 2026
1 check passed
mhenrichsen pushed a commit that referenced this pull request Sep 22, 2026
The cell said 'none yet'; the upstream PR is vllm-project/vllm#58026. The link was lost in a rebase of #167.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants