Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions PATCHES.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ build by name instead of landing by guess. Regenerate a file with `bash scripts/
| qwen3_5-embed-quant | fix | pass `quant_config` to the token embedding (main model and MTP module) | none yet | 0.28.0 | upstream PR |
| qwen3_5-mtp-draft-vocab | feature | vocab-truncated draft head for MTP | none | 0.28.0 | upstreamed |
| sampler-small-topk-fast-softmax | feature | sort-free top-k/top-p for small k, multi-block row softmax; registers `VLLM_DRAFT_TOPK_TOPP` and `VLLM_DRAFT_TEMP_SCALE` (both read once at import) | none | 0.28.0 | upstreamed or superseded |
| serve-404-served-names | fix | the model-not-found 404 lists the served names (`Served models: ...`) so a misnamed model is a one-read response body | vllm #58025 | 0.28.0 | upstream PR |
| spec-decode-attn | feature | split-KV verify attention on FLASH_ATTN with query-row tiling; registers `VLLM_SPEC_DECODE_ATTN`, `VLLM_SPEC_DECODE_ATTN_QMAX`, `VLLM_SPEC_ATTN_BLOCK_M` (#114) | none | 0.28.0 | upstreamed |
| spec-decode-int4-kv-mq3d | feature | multi-query 3D int4 verify path | none | 0.28.0 | rides with int4-kv-per-token-head |
| spec-decode-int8-kv | feature | split-KV verify attention over an int8 per-token-head cache | none | 0.28.0 | rides with spec-decode-attn |
Expand Down
4 changes: 4 additions & 0 deletions patches/series
Original file line number Diff line number Diff line change
Expand Up @@ -51,4 +51,8 @@ engine-stall-sentinel.patch
sse-keep-alive.patch
int4-mq3d-envs.patch
triton-spec-attn-fp8-kv.patch
<<<<<<< HEAD
bench-probe-errors.patch
=======
serve-404-served-names.patch
>>>>>>> 48e63ae (patches: serve-404-served-names — the model-not-found 404 lists the served names)
31 changes: 31 additions & 0 deletions patches/serve-404-served-names.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
The model-not-found 404 now lists the served names.

"The model `X` does not exist." sends the reader to the server log to find
out what WOULD have been accepted. /v1/models already knows; the 404 body
now carries the same list, so a misnamed model becomes a one-read response
body instead of an incident:

The model `X` does not exist. Served models: qwen3.8-27b.

--- exported from cpuchip/vllm 1c5d490 (serve-404-served-names); regenerate with scripts/export-patch.sh, do not edit ---

diff --git a/entrypoints/serve/engine/serving.py b/entrypoints/serve/engine/serving.py
index a31e47d..ada47d6 100644
--- a/entrypoints/serve/engine/serving.py
+++ b/entrypoints/serve/engine/serving.py
@@ -60,8 +60,14 @@
):
error_response = load_result

+ served_names = ", ".join(
+ model.name for model in self.models.base_model_paths
+ )
return error_response or self.create_error_response(
- message=f"The model `{request.model}` does not exist.",
+ message=(
+ f"The model `{request.model}` does not exist. "
+ f"Served models: {served_names}."
+ ),
err_type="NotFoundError",
status_code=HTTPStatus.NOT_FOUND,
param="model",
Loading