patches: tokenize-v1-route — /tokenize and /detokenize under /v1 as well - #168
Conversation
d17acfe to
9e8f305
Compare
4126a2e to
1bbad4b
Compare
Any OpenAI-SDK client (base_url ending in /v1) and most gateways resolve every endpoint relative to /v1, so the root-only tokenization routes were a genuine 404 for them. The patch includes the router twice — once at the root, once with prefix="/v1" — with operation ids staying unique by FastAPI's name+path+method derivation. Independent of every other patch, appended at the series' block boundary. Cut from the extended cpuchip/vllm qwen38/0.28 branch, topic commit [qwen38] tokenize-v1-route; kind: feature, retires when upstream takes it. The verify.sh /v1/tokenize row turns green with this patch installed.
1bbad4b to
c38436d
Compare
|
Merged (rebased for you). The operation-id check is what made this easy to take. Mounting one router twice is the classic way to get duplicate ids and an invalid OpenAPI schema, and "I checked instead of assuming" with the four ids printed is the difference between a two-line patch I can merge and one I have to go verify myself. The Concretely useful here beyond SDK clients: |
…, the four serving and bench patches from syv-ai#165-syv-ai#168, the triton message fix)
What
Adds
patches/tokenize-v1-route.patch, itspatches/seriesline and itsPATCHES.mdrow./tokenizeand/detokenizeare served under/v1as well as at the root.Upstream: vllm-project/vllm#58027.
Why
vLLM registers the tokenization endpoints at the root only. OpenAI SDK clients and gateways set
base_urlto a URL ending in/v1and resolve endpoints against it, so/v1/tokenizeis a real 404 for them. Measured on a live server from the current image:POST /tokenize200,POST /v1/tokenize404, and/openapi.jsonlists/tokenizeonly.That matters here because
bench/labd_accept.py:150builds its teacher-forced prompts through/tokenize, and the headroom proxy indocker-compose.ymlspeaks an OpenAI base URL. One extra mount, and both work from one base URL.Operation ids
Mounting one router twice is the usual way to get duplicate
operationIds, so I checked instead of assuming. FastAPI 0.141.1, the same twoinclude_routercalls:Note: with
--enable-tokenizer-info-endpoint,/tokenizer_infois on the same router and also gets the/v1mount. This repo leaves that flag off.Verification
patch integrity(the workflow'sgit-applyjob) applies the whole series to a pristinevllm-project/vllmcheckout at the pin: passes on this PR.patch -p1 --fuzz 0 --dry-runof this file against the installed tree inghcr.io/syv-ai/hyperqwen:latest(06150174): applies, no fuzz.verify.shneeds no new entry: its loop readspatches/series, andpatches/_check_applied.pyparses the patch file itself.Not done: I have not rebuilt the image and restarted a server on this patch, so the runtime evidence above comes from the code path, not from a rebuilt server.