Skip to content

patches: serve-404-served-names — the model-not-found 404 lists the served names - #166

Merged
mhenrichsen merged 2 commits into
syv-ai:mainfrom
TyroneNel:serve-404-served-names
Sep 22, 2026
Merged

mhenrichsen merged 2 commits into
syv-ai:mainfrom
TyroneNel:serve-404-served-names

Conversation

@TyroneNel

@TyroneNel TyroneNel commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

What

Adds patches/serve-404-served-names.patch, its patches/series line and its PATCHES.md row. The model-not-found 404 now names the served models:

The model `x` does not exist. Served models: qwen3.8-27b.

Upstream: vllm-project/vllm#58025.

Why

The current body gives the rejected name and nothing else, so the reader must call /v1/models or read the server log. Every launcher here passes --served-model-name qwen3.8-27b while --model is a checkpoint path, so a client that uses the path gets this, captured on a live server from the current image:

{"error":{"message":"The model `/app/models/Qwen3.8-27B-W4A16-AutoRound` does not exist.","type":"NotFoundError","param":"model","code":404}}

That reads like a missing route. It is a name mismatch. The list is already in base_model_paths, which is what /v1/models renders, so the message costs one f-string.

Status code, error type and param do not change.

Verification

  • patch integrity (the workflow's git-apply job) applies the whole series to a pristine vllm-project/vllm checkout at the pin: passes on this PR.
  • patch -p1 --fuzz 0 --dry-run of this file against the installed tree in ghcr.io/syv-ai/hyperqwen:latest (06150174): applies, no fuzz.
  • verify.sh needs no new entry: its loop reads patches/series, and patches/_check_applied.py parses the patch file itself.

Not done: I have not rebuilt the image and restarted a server on this patch, so the runtime evidence above comes from the code path, not from a rebuilt server.

…erved names

One hunk in BaseServing._check_model: the 404 body appends the names the
server actually serves, so a misnamed model is a one-read response body
instead of a log hunt. Independent of every other patch (nothing else
touches entrypoints/serve/engine/serving.py), appended at the series'
block boundary. Cut from the extended cpuchip/vllm qwen38/0.28 branch,
topic commit [qwen38] serve-404-served-names; kind: fix, retires when
upstream takes it. The verify.sh contract row (unknown name -> 404 + names)
turns green with this patch installed.
@mhenrichsen
mhenrichsen force-pushed the serve-404-served-names branch from 9c7658a to ca4226f Compare September 22, 2026 17:13
@mhenrichsen

Copy link
Copy Markdown
Contributor

Merged (rebased onto main for you — #165 moved the tail of patches/series and PATCHES.md).

Verified against the pinned source rather than only against the patch: OpenAIServingModels.base_model_paths is a list[BaseModelPath] with name and model_path (entrypoints/openai/models/protocol.py:9), and show_available_models renders id=name, root=model_path from the same list — so the 404 body and /v1/models cannot disagree by construction. Whole series applies to a pristine v0.28.0 checkout at fuzz 0 with all five stacked.

One thing you may want for the upstream PR, not a blocker here: entrypoints/openai/models/serving.py has a second producer of the same string, OpenAIServingModels.check_model, and entrypoints/pooling/base/serving.py has a third inside its own _check_model. Yours is the one every route this repo serves goes through — /v1/chat/completions, /v1/completions, /tokenize and /detokenize all reach OpenAIServing._check_model at entrypoints/serve/engine/serving.py:62 — and OpenAIServingModels.check_model has no in-tree callers at all. But a reviewer at vllm-project will ask, and "three copies, one of them dead" is a better answer than fixing one and leaving the others to drift.

@mhenrichsen
mhenrichsen merged commit 41f3470 into syv-ai:main Sep 22, 2026
1 check failed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants