patches: serve-404-served-names — the model-not-found 404 lists the served names - #166
Conversation
f46a673 to
9c7658a
Compare
…erved names One hunk in BaseServing._check_model: the 404 body appends the names the server actually serves, so a misnamed model is a one-read response body instead of a log hunt. Independent of every other patch (nothing else touches entrypoints/serve/engine/serving.py), appended at the series' block boundary. Cut from the extended cpuchip/vllm qwen38/0.28 branch, topic commit [qwen38] serve-404-served-names; kind: fix, retires when upstream takes it. The verify.sh contract row (unknown name -> 404 + names) turns green with this patch installed.
9c7658a to
ca4226f
Compare
|
Merged (rebased onto main for you — #165 moved the tail of Verified against the pinned source rather than only against the patch: One thing you may want for the upstream PR, not a blocker here: |
What
Adds
patches/serve-404-served-names.patch, itspatches/seriesline and itsPATCHES.mdrow. The model-not-found 404 now names the served models:Upstream: vllm-project/vllm#58025.
Why
The current body gives the rejected name and nothing else, so the reader must call
/v1/modelsor read the server log. Every launcher here passes--served-model-name qwen3.8-27bwhile--modelis a checkpoint path, so a client that uses the path gets this, captured on a live server from the current image:{"error":{"message":"The model `/app/models/Qwen3.8-27B-W4A16-AutoRound` does not exist.","type":"NotFoundError","param":"model","code":404}}That reads like a missing route. It is a name mismatch. The list is already in
base_model_paths, which is what/v1/modelsrenders, so the message costs one f-string.Status code, error type and
paramdo not change.Verification
patch integrity(the workflow'sgit-applyjob) applies the whole series to a pristinevllm-project/vllmcheckout at the pin: passes on this PR.patch -p1 --fuzz 0 --dry-runof this file against the installed tree inghcr.io/syv-ai/hyperqwen:latest(06150174): applies, no fuzz.verify.shneeds no new entry: its loop readspatches/series, andpatches/_check_applied.pyparses the patch file itself.Not done: I have not rebuilt the image and restarted a server on this patch, so the runtime evidence above comes from the code path, not from a rebuilt server.