bench: pass the served model name, keep the checkpoint as --tokenizer - #170
Conversation
vllm bench serve posts --model verbatim as the request's model in its /tokenize alignment probe, so passing the checkpoint path 404s the model check there and every run logs "WARNING: /tokenize unavailable, skipping alignment" (314 of 602 campaign logs carry it). --model is now the served name (qwen3.8-27b) on every bench client invocation, and --tokenizer loads the tokenizer from the checkpoint dir, which is what the path was for. Zero risk, no server restart: requests already carried the served name via --served-model-name, only the probe was fed the path. No existing row is invalidated, so nothing needs re-running: the random dataset re-encodes prompts with the same local tokenizer either way, the alignment step the warning skips would have early-returned on token-count agreement, 0 of 602 logs show the client's own tokenizer-mismatch warning, and the custom cohorts never reach the probe. The campaign tables stand.
1a91a1c to
d71235d
Compare
|
Merged. Confirmed in the pinned client rather than taking it on trust: The "no published number changes" section is why this merged without a re-run. The argument is complete: Two notes for the record, neither blocking:
|
What
bench/run_benchmarks.sh:25,bench/prefill_ab.sh:23,bench/real_rep.sh:11andbench/warmup.sh:50pass the served name as--modeland the checkpoint dir as--tokenizer:Four lines plus one comment each.
manifest.jsonstill records the checkpoint path inmodel_arg.Why
vllm bench serveuses--modelfor two jobs: it loads the tokenizer from it, and it posts it as themodelfield of the tokenizer-alignment probe (0.28.0:model_id = args.modelat line 2043, probe at 96-104, call at 2126, gated todataset_name in ("random", "prefix_repetition")).Our
--modelis a path and the server knows onlyqwen3.8-27b, so the probe 404s on the model check and the client printsWARNING: /tokenize unavailable, skipping alignment.Measured on a live server from the current image:--tokenizerexists in the pinned client (vllm/benchmarks/serve.py:1627), so the two jobs separate cleanly.No published number changes
Stated explicitly, because "the benchmark skipped a step" invites a re-run and I do not think one row needs it:
RandomDataset.samplealready decodes and re-encodes locally to force the prompt length, andprompt_lenis that local count.len(first_tokens) == expectedearly return.The warning was cosmetic. The fix is worth having so the next reader does not have to prove that again, and because on a keyed server the same probe fails for a second reason (no
Authorizationheader; fixed upstream in vllm-project/vllm#58024, carried in #165).Verification
Ran the fixed invocation in the container against a live server from
ghcr.io/syv-ai/hyperqwen:latest:/tokenize unavailableline (present in the same invocation with--model <path>)Successful requests: 2Total input tokens: 256(2 x 128 exactly, so alignment ran and agreed)bash -nclean on all four scripts.