From 8aeabf4fc07b10c609506556e404c20bcea861a1 Mon Sep 17 00:00:00 2001 From: Tyrone <71038642+TyroneNel@users.noreply.github.com> Date: Mon, 21 Sep 2026 21:15:20 +0000 Subject: [PATCH 1/2] =?UTF-8?q?patches:=20serve-404-served-names=20?= =?UTF-8?q?=E2=80=94=20the=20model-not-found=20404=20lists=20the=20served?= =?UTF-8?q?=20names?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One hunk in BaseServing._check_model: the 404 body appends the names the server actually serves, so a misnamed model is a one-read response body instead of a log hunt. Independent of every other patch (nothing else touches entrypoints/serve/engine/serving.py), appended at the series' block boundary. Cut from the extended cpuchip/vllm qwen38/0.28 branch, topic commit [qwen38] serve-404-served-names; kind: fix, retires when upstream takes it. The verify.sh contract row (unknown name -> 404 + names) turns green with this patch installed. --- PATCHES.md | 1 + patches/series | 4 ++++ patches/serve-404-served-names.patch | 31 ++++++++++++++++++++++++++++ 3 files changed, 36 insertions(+) create mode 100644 patches/serve-404-served-names.patch diff --git a/PATCHES.md b/PATCHES.md index be8994ad..add41e42 100644 --- a/PATCHES.md +++ b/PATCHES.md @@ -46,6 +46,7 @@ build by name instead of landing by guess. Regenerate a file with `bash scripts/ | qwen3_5-embed-quant | fix | pass `quant_config` to the token embedding (main model and MTP module) | none yet | 0.28.0 | upstream PR | | qwen3_5-mtp-draft-vocab | feature | vocab-truncated draft head for MTP | none | 0.28.0 | upstreamed | | sampler-small-topk-fast-softmax | feature | sort-free top-k/top-p for small k, multi-block row softmax; registers `VLLM_DRAFT_TOPK_TOPP` and `VLLM_DRAFT_TEMP_SCALE` (both read once at import) | none | 0.28.0 | upstreamed or superseded | +| serve-404-served-names | fix | the model-not-found 404 lists the served names (`Served models: ...`) so a misnamed model is a one-read response body | none yet | 0.28.0 | upstream PR | | spec-decode-attn | feature | split-KV verify attention on FLASH_ATTN with query-row tiling; registers `VLLM_SPEC_DECODE_ATTN`, `VLLM_SPEC_DECODE_ATTN_QMAX`, `VLLM_SPEC_ATTN_BLOCK_M` (#114) | none | 0.28.0 | upstreamed | | spec-decode-int4-kv-mq3d | feature | multi-query 3D int4 verify path | none | 0.28.0 | rides with int4-kv-per-token-head | | spec-decode-int8-kv | feature | split-KV verify attention over an int8 per-token-head cache | none | 0.28.0 | rides with spec-decode-attn | diff --git a/patches/series b/patches/series index f1b29b41..38edf7ec 100644 --- a/patches/series +++ b/patches/series @@ -51,4 +51,8 @@ engine-stall-sentinel.patch sse-keep-alive.patch int4-mq3d-envs.patch triton-spec-attn-fp8-kv.patch +<<<<<<< HEAD bench-probe-errors.patch +======= +serve-404-served-names.patch +>>>>>>> 48e63ae (patches: serve-404-served-names — the model-not-found 404 lists the served names) diff --git a/patches/serve-404-served-names.patch b/patches/serve-404-served-names.patch new file mode 100644 index 00000000..2ec4fba8 --- /dev/null +++ b/patches/serve-404-served-names.patch @@ -0,0 +1,31 @@ +The model-not-found 404 now lists the served names. + +"The model `X` does not exist." sends the reader to the server log to find +out what WOULD have been accepted. /v1/models already knows; the 404 body +now carries the same list, so a misnamed model becomes a one-read response +body instead of an incident: + + The model `X` does not exist. Served models: qwen3.8-27b. + +--- exported from cpuchip/vllm 1c5d490 (serve-404-served-names); regenerate with scripts/export-patch.sh, do not edit --- + +diff --git a/entrypoints/serve/engine/serving.py b/entrypoints/serve/engine/serving.py +index a31e47d..ada47d6 100644 +--- a/entrypoints/serve/engine/serving.py ++++ b/entrypoints/serve/engine/serving.py +@@ -60,8 +60,14 @@ + ): + error_response = load_result + ++ served_names = ", ".join( ++ model.name for model in self.models.base_model_paths ++ ) + return error_response or self.create_error_response( +- message=f"The model `{request.model}` does not exist.", ++ message=( ++ f"The model `{request.model}` does not exist. " ++ f"Served models: {served_names}." ++ ), + err_type="NotFoundError", + status_code=HTTPStatus.NOT_FOUND, + param="model", From ca4226fcb4e1f06316eca23b7448ba54b8316c2a Mon Sep 17 00:00:00 2001 From: Tyrone <71038642+TyroneNel@users.noreply.github.com> Date: Tue, 22 Sep 2026 00:26:52 +0200 Subject: [PATCH 2/2] PATCHES.md: link the upstream vLLM PR for serve-404-served-names (vllm-project/vllm#58025) --- PATCHES.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/PATCHES.md b/PATCHES.md index add41e42..430b6b6a 100644 --- a/PATCHES.md +++ b/PATCHES.md @@ -46,7 +46,7 @@ build by name instead of landing by guess. Regenerate a file with `bash scripts/ | qwen3_5-embed-quant | fix | pass `quant_config` to the token embedding (main model and MTP module) | none yet | 0.28.0 | upstream PR | | qwen3_5-mtp-draft-vocab | feature | vocab-truncated draft head for MTP | none | 0.28.0 | upstreamed | | sampler-small-topk-fast-softmax | feature | sort-free top-k/top-p for small k, multi-block row softmax; registers `VLLM_DRAFT_TOPK_TOPP` and `VLLM_DRAFT_TEMP_SCALE` (both read once at import) | none | 0.28.0 | upstreamed or superseded | -| serve-404-served-names | fix | the model-not-found 404 lists the served names (`Served models: ...`) so a misnamed model is a one-read response body | none yet | 0.28.0 | upstream PR | +| serve-404-served-names | fix | the model-not-found 404 lists the served names (`Served models: ...`) so a misnamed model is a one-read response body | vllm #58025 | 0.28.0 | upstream PR | | spec-decode-attn | feature | split-KV verify attention on FLASH_ATTN with query-row tiling; registers `VLLM_SPEC_DECODE_ATTN`, `VLLM_SPEC_DECODE_ATTN_QMAX`, `VLLM_SPEC_ATTN_BLOCK_M` (#114) | none | 0.28.0 | upstreamed | | spec-decode-int4-kv-mq3d | feature | multi-query 3D int4 verify path | none | 0.28.0 | rides with int4-kv-per-token-head | | spec-decode-int8-kv | feature | split-KV verify attention over an int8 per-token-head cache | none | 0.28.0 | rides with spec-decode-attn |