Skip to content

verify: contract-test the tokenize endpoint at the HTTP boundary - #171

Merged
mhenrichsen merged 4 commits into
syv-ai:mainfrom
TyroneNel:tokenize-contract-verify
Sep 23, 2026
Merged

mhenrichsen merged 4 commits into
syv-ai:mainfrom
TyroneNel:tokenize-contract-verify

Conversation

@TyroneNel

@TyroneNel TyroneNel commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

What

Six rows in verify.sh's live-server section that test the /tokenize contract, plus one environment guard:

row asserts
served id the ids equal the checkpoint tokenizer's encode
model omitted the same ids (the field is optional; the server falls back to the served model)
chat form, enable_thinking:false the ids equal apply_chat_template
unknown name 404 whose body lists the served names
keyless 401
/v1/tokenize 200 with the same ids
guard FAIL if VLLM_SKIP_MODEL_NAME_VALIDATION is set

Why

/tokenize exists to prove that the client's tokenizer and the server's agree, and nothing asserted it, although bench/labd_accept.py:150 builds its teacher-forced prompts from it and every random-dataset benchmark run probes it. The rows catch gotcha 32 (a dir with no tokenizer.json encodes everything to []) and chat-template drift at the HTTP boundary.

The guard exists because VLLM_SKIP_MODEL_NAME_VALIDATION looks like the fix for a model-name 404 and is not: it disables the check on /v1/chat/completions too, so a typo'd name would silently serve the wrong checkpoint.

The last three rows are graded, not asserted

They describe behaviour that three patches add (#166, #168, #169). As plain fail rows they would make verify.sh exit 1 on every live box between merging those patches and rebuilding the image, including anyone on the published image. So each greps the installed tree for its patch:

graded() { # patch file-under-$SP marker message — FAIL if installed, else WARN
  if grep -q "$3" "$SP/$2" 2>/dev/null; then fail "$4"
  else warn "$4 [pending: $1 is not in $SP — rebuild the image]"; fi
}

Markers: Served models: in entrypoints/serve/engine/serving.py, UNGUARDED_PATHS in entrypoints/serve/utils/server_utils.py, prefix="/v1" in entrypoints/serve/tokenize/api_router.py. $SP is the path verify.sh resolves at line 28.

If you would rather this waited for those three PRs and then asserted unconditionally, I will rebase it that way. The first three rows and the guard stand on their own either way.

Verification

Against a live server from the current image, which has none of the three patches:

  PASS  /tokenize with the served id returns the checkpoint tokenizer's ids
  PASS  /tokenize with the model omitted returns the same ids
  PASS  /tokenize chat form (enable_thinking:false) matches apply_chat_template
  WARN  unknown model name -> 404, body lists nothing usable ... [pending: serve-404-served-names is not in <SP>]
  WARN  keyless /tokenize -> 200 ... [pending: auth-deny-default is not in <SP>]
  WARN  /v1/tokenize unavailable or disagrees ('') [pending: tokenize-v1-route is not in <SP>]

graded itself was exercised both ways with a stub $SP: marker absent → WARN, file missing → WARN, marker present → FAIL with the failure counted.

Three mechanics I got wrong first and fixed before proposing this, in case they read oddly in the diff:

  1. the local-tokenizer heredocs read the prompt as sys.argv[2], so it must be passed as an argument; stderr is dropped so a tokenizer that will not load leaves the row empty instead of printing a traceback;
  2. apply_chat_template(tokenize=True) returns a BatchEncoding on the pinned transformers==5.15.0, so the chat row renders with tokenize=False and encodes the text, the form bench/bugb_sweep.py:64 and bench/residue_sweep.py:78 already use (identical ids on 4.x and 5.x, checked);
  3. the model-omitted and /v1/tokenize rows require a non-empty server answer, otherwise empty == empty passes and the /v1 row reports green against a server with no such route.

bash -n verify.sh clean. verify.sh --install is unaffected: at build time there is no server, so the block is skipped.

Review follow-up: a rejected chat template warns, it does not fail

The chat row compared /tokenize's chat form against a local apply_chat_template and failed on any empty server answer. A checkpoint whose template raise_exceptions on a kwarg it does not know (gotcha 58) then reads as tokenizer drift: the server answers 400 and the local render raises, so both sides are empty for a reason that has nothing to do with the two tokenizers disagreeing — and docs/third-party-checkpoints.md now lists four community checkpoints with their own templates.

Those two cases are now WARN, with the server's body; FAIL is kept for the case the row exists to catch. Exercised against a stub server, four behaviours:

server local render row
ids match the template ok PASS
ids differ ok FAIL chat form disagrees (server='9,9,9' template='1,2,3,4')
400 from the template ok WARN this checkpoint's chat template rejected the request (400), which is not tokenizer drift: {"object": "error", "message": "Unknown effort value …
200 raises locally WARN apply_chat_template raised locally …, so there is nothing to compare the server against

The graded mechanism is unchanged, and the keyless row stays graded while #169 is held.

Nothing asserted client/server tokenizer agreement, which is the endpoint's
only purpose. New rows against the live server:
  - served id -> token ids equal the checkpoint tokenizer's encode
  - model omitted -> identical ids
  - chat form with enable_thinking:false -> equal apply_chat_template
  - unknown name -> 404 whose body lists the served names
  - keyless -> 401
  - /v1/tokenize -> 200 with the same ids
These catch gotcha 32 (empty-vocab dirs) and chat-template drift where the
bench client's alignment probe actually operates. The last three rows go
green when the serving patches (serve-404-served-names, auth-deny-default,
tokenize-v1-route) are in the installed tree and the server is rebuilt;
until then they state the contract this repo expects, fail-closed.

Also FAIL if VLLM_SKIP_MODEL_NAME_VALIDATION is set: it looks like a fix for
model-name 404s but disables the check on /v1/chat/completions too, so a
typo'd name would silently serve the wrong checkpoint and benchmark
provenance dies with it.

Row mechanics, fixed before this lands (they were broken as first written):
  - the heredocs took the prompt as sys.argv[2] but only $MODEL was passed,
    so every local-tokenizer row raised IndexError, compared against an empty
    string and failed; the prompt is now an argument and stderr is dropped;
  - apply_chat_template(tokenize=True) returns a BatchEncoding on the pinned
    transformers 5.15, so the chat row compared ids against the literal
    "input_ids,attention_mask"; it now renders with tokenize=False and
    encodes the text, the form bench/bugb_sweep.py and bench/residue_sweep.py
    already use (identical ids on 4.x and 5.x);
  - the model-omitted and /v1/tokenize rows compared to $TOK_LOC with no
    non-empty guard, so empty == empty passed: the /v1 row reported green
    against a server that has no such route. Both now require a non-empty
    server answer, i.e. they fail closed.

Verified against a live server: the three in-tree rows pass (ids match the
checkpoint tokenizer for the plain, model-omitted and chat forms) and the
three forward-looking rows fail with the reasons above, no tracebacks.

The three forward-looking rows are graded against the installed tree rather
than asserted unconditionally: they FAIL once their patch is in $SP (marker
grep: "Served models:", UNGUARDED_PATHS, prefix="/v1") and WARN
"[pending: <patch> is not in $SP]" until then. Otherwise verify.sh exits 1 on
every live box — including anyone on the published image — for the whole
window between merging the serving patches and rebuilding, and a check that
is red for reasons the operator cannot act on stops being read.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Three of the four things this depends on are now on main — #166 (Served models:), #167 and #168 (/v1 mount) merged today; #169 is held for one boot and a trailing-slash fix, so the keyless row stays graded for now. Keep the graded mechanism rather than rebasing to unconditional asserts: it is the right answer for a repo whose users run a published image, and it stays right the next time a patch lands before a rebuild.

I am holding this one on a live-server run, which needs the GPUs back. Two things I want to check when it runs, both about rows failing on servers that are fine:

The chat row can fail for a reason that is not drift. It compares /tokenize's chat form against local apply_chat_template, and a checkpoint whose template raises on an unexpected kwarg gives you a 400 from the server and an exception locally — the row then sees an empty server answer and, by your own rule 3, fails. That is not hypothetical here: gotcha 58 is exactly this shape (a template that raise_exceptions on an effort value it does not know, 400ing the request), and docs/third-party-checkpoints.md now lists four community checkpoints with their own templates. A 400 on the chat form should probably WARN with the body, not FAIL — "this checkpoint's template rejected the request" is a different finding from "the tokenizers disagree", and only the second one should stop a boot.

The enable_thinking:false and tokenize=True/BatchEncoding notes are the sort of thing I would have got wrong too, and writing down which three mechanics you fixed before proposing is more useful than the diff. The transformers==5.15.0 return-type change in particular will bite the next person editing those heredocs.

The VLLM_SKIP_MODEL_NAME_VALIDATION guard I would take on its own today if it were a separate PR — it is a genuine trap (it looks like the fix for a name 404 and silently disables the check on /v1/chat/completions too, so a typo serves the wrong checkpoint), and it has no live-server dependency at all. No need to split it; just saying it is not what is holding this.

Queued behind the GPU, not behind a disagreement.

The chat row compared /tokenize's chat form against a local
apply_chat_template, and any empty server answer failed the row. A checkpoint
whose template raise_exception()s on a kwarg it does not know (gotcha 58) then
reads as tokenizer drift: the server answers 400 and the local render raises,
so both sides are empty for a reason that has nothing to do with the two
tokenizers disagreeing.

Grade those two cases as WARN, with the server's body, and keep FAIL for the
case the row exists to catch: both sides answered and the ids differ.
@TyroneNel

Copy link
Copy Markdown
Contributor Author

The chat row now warns instead of failing when the template is what rejected the request. You were right that it is a different finding, and gotcha 58 makes it a live one rather than a hypothetical.

Two cases moved to WARN — a 400 from the server (with the body), and apply_chat_template raising locally — and FAIL is kept for the case the row exists to catch. Exercised against a stub server:

server local render row
ids match ok PASS
ids differ ok FAIL chat form disagrees (server='9,9,9' template='1,2,3,4')
400 from the template ok WARN this checkpoint's chat template rejected the request (400), which is not tokenizer drift: …
200 raises locally WARN apply_chat_template raised locally …, so there is nothing to compare the server against

The graded mechanism stays exactly as it is — agreed that it is the right answer for a repo whose users run a published image, and the keyless row stays graded while #169 is held.

I have left the VLLM_SKIP_MODEL_NAME_VALIDATION guard in this PR since you said it is not what is holding it; happy to split it out the moment you want it sooner.

Queued behind the GPU is fine. Thanks for writing down which mechanics were worth keeping — the transformers==5.15.0 return-type note is in the body precisely so the next person editing those heredocs does not rediscover it.

@mhenrichsen

Copy link
Copy Markdown
Contributor

Ran it live, against a server carrying the patches its graded rows are waiting for — today's four merged ones plus #169's branch, on the reference 3090, keyed with api_key.txt:

== live server (127.0.0.1:18020)
  PASS  /health 200
  PASS  chat completion answers ('København')
  PASS  /tokenize with the served id returns the checkpoint tokenizer's ids
  PASS  /tokenize with the model omitted returns the same ids
  PASS  /tokenize chat form (enable_thinking:false) matches apply_chat_template
  PASS  unknown model name -> 404 whose body lists the served names
  PASS  keyless /tokenize -> 401
  PASS  /v1/tokenize answers with the same ids (OpenAI-SDK base_url .../v1)

So the graded mechanism works in both directions: on the published image those last three are warnings, and here, with the markers present in the installed tree, they asserted and passed. The chat-template row matching apply_chat_template on the real checkpoint is the one I most wanted to see on real hardware, and it holds.

For completeness: the run reported one failure outside these rows — patches/series out of sync with patches/. That is my harness, not your PR. I dropped this branch's verify.sh into the box's serving checkout, which is a month behind main, so its patches/ directory predates today's patches. Nothing in your diff touches that check.

Merging once #169 lands — the keyless row asserts against #169's behaviour, and I would rather the two go in together than have this assert a guard main does not have yet. If #169 slips, I will take this as-is: the row is graded and warns cleanly without it.

@mhenrichsen

Copy link
Copy Markdown
Contributor

Still queued behind #169, which now needs a 0.29 cut (main moved to vLLM 0.29.0 with #148). Your branch's verify.sh rows need nothing: they grep the installed tree for their markers, so they stay graded until #169's 0.29 patch lands, then assert.

…eware/authenticate.py on 0.29

The 0.28 path, entrypoints/serve/utils/server_utils.py, does not exist on
the 0.29 pin, so the graded row would have stayed a warning forever once
syv-ai#169's 0.29 cut landed.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Merged. A correction and a fix first.

My last comment was wrong about one row. I said your rows needed nothing on 0.29. Two of the three graded markers are fine: Served models: in entrypoints/serve/engine/serving.py and prefix="/v1" in entrypoints/serve/tokenize/api_router.py both exist on the 0.29 tree, and both patches are on main now. But the auth-deny-default marker grepped entrypoints/serve/utils/server_utils.py, which 0.29 moved to entrypoints/serve/middleware/authenticate.py. After #169's 0.29 cut, that row would have stayed a warning forever. I pushed a one-line commit pointing it at the new file, on top of a merge of today's main.

Live on vLLM 0.29, reference 3090, the merged verify.sh against a keyed server from main's 0.29 tree (which carries the four merged patches and not #169):

  PASS  /health 200
  PASS  chat completion answers ('København')
  PASS  /tokenize with the served id returns the checkpoint tokenizer's ids
  PASS  /tokenize with the model omitted returns the same ids
  PASS  /tokenize chat form (enable_thinking:false) matches apply_chat_template
  PASS  unknown model name -> 404 whose body lists the served names
  WARN  keyless /tokenize -> 200 ... [pending: auth-deny-default is not in .../vllm — rebuild the image]
  PASS  /v1/tokenize answers with the same ids (OpenAI-SDK base_url .../v1)
verify: OK (0 failures)

The first seven rows are the same as last night on 0.28. The WARN is the graded mechanism doing its job, and it becomes an assertion the day #169 lands. Thanks for this, and for fixing the chat row so it reports a template rejection separately from tokenizer drift.

@mhenrichsen
mhenrichsen merged commit b82bb08 into syv-ai:main Sep 23, 2026
3 checks passed
cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
TyroneNel added a commit to TyroneNel/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
Takes upstream's 0.29.0 series wholesale (patches/series, PATCHES.md, the
re-cut patches, KVarN 0.29.0), including the four patches syv-ai#148 retired
(int4-mq3d-envs, sse-keep-alive, vllm-pr54282-draft-gumbel-salt,
xgrammar-spec-terminated).

Drops the fork's auth-deny-default.patch from the series and the tree: it is
cut against 0.28.0, both of its target files moved in 0.29.0
(serve/utils/server_utils.py -> serve/middleware/authenticate.py,
openai/cli_args.py -> launchers/cli_args.py), so it cannot apply at fuzz 0.
It returns with the 0.29 port in syv-ai#169.

Keeps from the fork: manual-only image builds and the fork's own buildcache
ref (docker-image.yml), the guarded resolver source in the bench scripts,
the F12 digest-pinned base, and F04 copy-only variant writes in
drafter/gptq_lm_head.py (upstream's syv-ai#181 file handling, copy semantics).
verify.sh equals upstream: every fork change to it landed via syv-ai#158/syv-ai#171/syv-ai#172.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants