verify: contract-test the tokenize endpoint at the HTTP boundary - #171
Conversation
Nothing asserted client/server tokenizer agreement, which is the endpoint's
only purpose. New rows against the live server:
- served id -> token ids equal the checkpoint tokenizer's encode
- model omitted -> identical ids
- chat form with enable_thinking:false -> equal apply_chat_template
- unknown name -> 404 whose body lists the served names
- keyless -> 401
- /v1/tokenize -> 200 with the same ids
These catch gotcha 32 (empty-vocab dirs) and chat-template drift where the
bench client's alignment probe actually operates. The last three rows go
green when the serving patches (serve-404-served-names, auth-deny-default,
tokenize-v1-route) are in the installed tree and the server is rebuilt;
until then they state the contract this repo expects, fail-closed.
Also FAIL if VLLM_SKIP_MODEL_NAME_VALIDATION is set: it looks like a fix for
model-name 404s but disables the check on /v1/chat/completions too, so a
typo'd name would silently serve the wrong checkpoint and benchmark
provenance dies with it.
Row mechanics, fixed before this lands (they were broken as first written):
- the heredocs took the prompt as sys.argv[2] but only $MODEL was passed,
so every local-tokenizer row raised IndexError, compared against an empty
string and failed; the prompt is now an argument and stderr is dropped;
- apply_chat_template(tokenize=True) returns a BatchEncoding on the pinned
transformers 5.15, so the chat row compared ids against the literal
"input_ids,attention_mask"; it now renders with tokenize=False and
encodes the text, the form bench/bugb_sweep.py and bench/residue_sweep.py
already use (identical ids on 4.x and 5.x);
- the model-omitted and /v1/tokenize rows compared to $TOK_LOC with no
non-empty guard, so empty == empty passed: the /v1 row reported green
against a server that has no such route. Both now require a non-empty
server answer, i.e. they fail closed.
Verified against a live server: the three in-tree rows pass (ids match the
checkpoint tokenizer for the plain, model-omitted and chat forms) and the
three forward-looking rows fail with the reasons above, no tracebacks.
The three forward-looking rows are graded against the installed tree rather
than asserted unconditionally: they FAIL once their patch is in $SP (marker
grep: "Served models:", UNGUARDED_PATHS, prefix="/v1") and WARN
"[pending: <patch> is not in $SP]" until then. Otherwise verify.sh exits 1 on
every live box — including anyone on the published image — for the whole
window between merging the serving patches and rebuilding, and a check that
is red for reasons the operator cannot act on stops being read.
003b9fc to
fba3892
Compare
|
Three of the four things this depends on are now on I am holding this one on a live-server run, which needs the GPUs back. Two things I want to check when it runs, both about rows failing on servers that are fine: The chat row can fail for a reason that is not drift. It compares The The Queued behind the GPU, not behind a disagreement. |
The chat row compared /tokenize's chat form against a local apply_chat_template, and any empty server answer failed the row. A checkpoint whose template raise_exception()s on a kwarg it does not know (gotcha 58) then reads as tokenizer drift: the server answers 400 and the local render raises, so both sides are empty for a reason that has nothing to do with the two tokenizers disagreeing. Grade those two cases as WARN, with the server's body, and keep FAIL for the case the row exists to catch: both sides answered and the ids differ.
|
The chat row now warns instead of failing when the template is what rejected the request. You were right that it is a different finding, and gotcha 58 makes it a live one rather than a hypothetical. Two cases moved to WARN — a 400 from the server (with the body), and
The I have left the Queued behind the GPU is fine. Thanks for writing down which mechanics were worth keeping — the |
|
Ran it live, against a server carrying the patches its graded rows are waiting for — today's four merged ones plus #169's branch, on the reference 3090, keyed with So the For completeness: the run reported one failure outside these rows — Merging once #169 lands — the keyless row asserts against #169's behaviour, and I would rather the two go in together than have this assert a guard |
…eware/authenticate.py on 0.29 The 0.28 path, entrypoints/serve/utils/server_utils.py, does not exist on the 0.29 pin, so the graded row would have stayed a warning forever once syv-ai#169's 0.29 cut landed.
|
Merged. A correction and a fix first. My last comment was wrong about one row. I said your rows needed nothing on 0.29. Two of the three graded markers are fine: Live on vLLM 0.29, reference 3090, the merged The first seven rows are the same as last night on 0.28. The WARN is the graded mechanism doing its job, and it becomes an assertion the day #169 lands. Thanks for this, and for fixing the chat row so it reports a template rejection separately from tokenizer drift. |
…vcc, syv-ai#171 tokenize contract test)
Takes upstream's 0.29.0 series wholesale (patches/series, PATCHES.md, the re-cut patches, KVarN 0.29.0), including the four patches syv-ai#148 retired (int4-mq3d-envs, sse-keep-alive, vllm-pr54282-draft-gumbel-salt, xgrammar-spec-terminated). Drops the fork's auth-deny-default.patch from the series and the tree: it is cut against 0.28.0, both of its target files moved in 0.29.0 (serve/utils/server_utils.py -> serve/middleware/authenticate.py, openai/cli_args.py -> launchers/cli_args.py), so it cannot apply at fuzz 0. It returns with the 0.29 port in syv-ai#169. Keeps from the fork: manual-only image builds and the fork's own buildcache ref (docker-image.yml), the guarded resolver source in the bench scripts, the F12 digest-pinned base, and F04 copy-only variant writes in drafter/gptq_lm_head.py (upstream's syv-ai#181 file handling, copy semantics). verify.sh equals upstream: every fork change to it landed via syv-ai#158/syv-ai#171/syv-ai#172.
What
Six rows in
verify.sh's live-server section that test the/tokenizecontract, plus one environment guard:encodeenable_thinking:falseapply_chat_template/v1/tokenizeVLLM_SKIP_MODEL_NAME_VALIDATIONis setWhy
/tokenizeexists to prove that the client's tokenizer and the server's agree, and nothing asserted it, althoughbench/labd_accept.py:150builds its teacher-forced prompts from it and every random-dataset benchmark run probes it. The rows catch gotcha 32 (a dir with notokenizer.jsonencodes everything to[]) and chat-template drift at the HTTP boundary.The guard exists because
VLLM_SKIP_MODEL_NAME_VALIDATIONlooks like the fix for a model-name 404 and is not: it disables the check on/v1/chat/completionstoo, so a typo'd name would silently serve the wrong checkpoint.The last three rows are graded, not asserted
They describe behaviour that three patches add (#166, #168, #169). As plain
failrows they would makeverify.shexit 1 on every live box between merging those patches and rebuilding the image, including anyone on the published image. So each greps the installed tree for its patch:Markers:
Served models:inentrypoints/serve/engine/serving.py,UNGUARDED_PATHSinentrypoints/serve/utils/server_utils.py,prefix="/v1"inentrypoints/serve/tokenize/api_router.py.$SPis the pathverify.shresolves at line 28.If you would rather this waited for those three PRs and then asserted unconditionally, I will rebase it that way. The first three rows and the guard stand on their own either way.
Verification
Against a live server from the current image, which has none of the three patches:
gradeditself was exercised both ways with a stub$SP: marker absent → WARN, file missing → WARN, marker present → FAIL with the failure counted.Three mechanics I got wrong first and fixed before proposing this, in case they read oddly in the diff:
sys.argv[2], so it must be passed as an argument; stderr is dropped so a tokenizer that will not load leaves the row empty instead of printing a traceback;apply_chat_template(tokenize=True)returns aBatchEncodingon the pinnedtransformers==5.15.0, so the chat row renders withtokenize=Falseand encodes the text, the formbench/bugb_sweep.py:64andbench/residue_sweep.py:78already use (identical ids on 4.x and 5.x, checked);/v1/tokenizerows require a non-empty server answer, otherwise empty == empty passes and the/v1row reports green against a server with no such route.bash -n verify.shclean.verify.sh --installis unaffected: at build time there is no server, so the block is skipped.Review follow-up: a rejected chat template warns, it does not fail
The chat row compared
/tokenize's chat form against a localapply_chat_templateand failed on any empty server answer. A checkpoint whose templateraise_exceptions on a kwarg it does not know (gotcha 58) then reads as tokenizer drift: the server answers 400 and the local render raises, so both sides are empty for a reason that has nothing to do with the two tokenizers disagreeing — anddocs/third-party-checkpoints.mdnow lists four community checkpoints with their own templates.Those two cases are now WARN, with the server's body; FAIL is kept for the case the row exists to catch. Exercised against a stub server, four behaviours:
chat form disagrees (server='9,9,9' template='1,2,3,4')this checkpoint's chat template rejected the request (400), which is not tokenizer drift: {"object": "error", "message": "Unknown effort value …apply_chat_template raised locally …, so there is nothing to compare the server againstThe
gradedmechanism is unchanged, and the keyless row stays graded while #169 is held.