Skip to content

bench: pass the served model name, keep the checkpoint as --tokenizer - #170

Merged
mhenrichsen merged 1 commit into
syv-ai:mainfrom
TyroneNel:bench-served-model-name
Sep 22, 2026
Merged

mhenrichsen merged 1 commit into
syv-ai:mainfrom
TyroneNel:bench-served-model-name

Conversation

@TyroneNel

@TyroneNel TyroneNel commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

What

bench/run_benchmarks.sh:25, bench/prefill_ab.sh:23, bench/real_rep.sh:11 and bench/warmup.sh:50 pass the served name as --model and the checkpoint dir as --tokenizer:

-B="venv/bin/vllm bench serve --host $HOST --port $PORT --model $MODEL --served-model-name qwen3.8-27b"
+B="venv/bin/vllm bench serve --host $HOST --port $PORT --model qwen3.8-27b --tokenizer $MODEL --served-model-name qwen3.8-27b"

Four lines plus one comment each. manifest.json still records the checkpoint path in model_arg.

Why

vllm bench serve uses --model for two jobs: it loads the tokenizer from it, and it posts it as the model field of the tokenizer-alignment probe (0.28.0: model_id = args.model at line 2043, probe at 96-104, call at 2126, gated to dataset_name in ("random", "prefix_repetition")).

Our --model is a path and the server knows only qwen3.8-27b, so the probe 404s on the model check and the client prints WARNING: /tokenize unavailable, skipping alignment. Measured on a live server from the current image:

POST /tokenize {"model":"qwen3.8-27b"}                             -> 200 {"count":1,"tokens":[14556]}
POST /tokenize {"model":"/app/models/Qwen3.8-27B-W4A16-AutoRound"} -> 404 {"message":"The model `...` does not exist.","type":"NotFoundError"}

--tokenizer exists in the pinned client (vllm/benchmarks/serve.py:1627), so the two jobs separate cleanly.

No published number changes

Stated explicitly, because "the benchmark skipped a step" invites a re-run and I do not think one row needs it:

  • RandomDataset.sample already decodes and re-encodes locally to force the prompt length, and prompt_len is that local count.
  • The client tokenizer is the same checkpoint the server serves, so alignment would have hit its len(first_tokens) == expected early return.
  • Residual drift would raise the client's own "more/fewer tokens than expected after decoding and re-encoding" warning. That string appears in 0 of 602 benchmark logs here.

The warning was cosmetic. The fix is worth having so the next reader does not have to prove that again, and because on a keyed server the same probe fails for a second reason (no Authorization header; fixed upstream in vllm-project/vllm#58024, carried in #165).

Verification

Ran the fixed invocation in the container against a live server from ghcr.io/syv-ai/hyperqwen:latest:

venv/bin/vllm bench serve --host 127.0.0.1 --port 18020 \
  --model qwen3.8-27b --tokenizer /app/models/<checkpoint> --served-model-name qwen3.8-27b \
  --dataset-name random --random-input-len 128 --random-output-len 16 --num-prompts 2 \
  --max-concurrency 1 --ignore-eos
  • no /tokenize unavailable line (present in the same invocation with --model <path>)
  • Successful requests: 2
  • Total input tokens: 256 (2 x 128 exactly, so alignment ran and agreed)

bash -n clean on all four scripts.

vllm bench serve posts --model verbatim as the request's model in its
/tokenize alignment probe, so passing the checkpoint path 404s the model
check there and every run logs "WARNING: /tokenize unavailable, skipping
alignment" (314 of 602 campaign logs carry it). --model is now the served
name (qwen3.8-27b) on every bench client invocation, and --tokenizer loads
the tokenizer from the checkpoint dir, which is what the path was for.

Zero risk, no server restart: requests already carried the served name via
--served-model-name, only the probe was fed the path.

No existing row is invalidated, so nothing needs re-running: the random
dataset re-encodes prompts with the same local tokenizer either way, the
alignment step the warning skips would have early-returned on token-count
agreement, 0 of 602 logs show the client's own tokenizer-mismatch warning,
and the custom cohorts never reach the probe. The campaign tables stand.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Merged.

Confirmed in the pinned client rather than taking it on trust: --tokenizer is real (vllm/benchmarks/serve.py:1627, "Name or path of the tokenizer, if not using the default tokenizer"), and the probe at lines 76-120 posts model_id — which is args.model — straight into /tokenize with raise_for_status() wrapped in a bare except Exception that prints the one line. So the two jobs really are entangled in one flag and --tokenizer really does separate them.

The "no published number changes" section is why this merged without a re-run. The argument is complete: RandomDataset.sample decodes and re-encodes locally and prompt_len is that local count, the client tokenizer is the server's checkpoint so alignment would have early-returned, and the one warning that would fire on residual drift appears in 0 of 602 logs. Proving a fix changes nothing measured is more work than the fix and it is the part that keeps a benchmark trustworthy — I would rather have this than a quiet edit.

Two notes for the record, neither blocking:

@mhenrichsen
mhenrichsen merged commit 15db29f into syv-ai:main Sep 22, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants