Skip to content

patches: tokenize-v1-route — /tokenize and /detokenize under /v1 as well - #168

Merged
mhenrichsen merged 2 commits into
syv-ai:mainfrom
TyroneNel:tokenize-v1-route
Sep 22, 2026
Merged

mhenrichsen merged 2 commits into
syv-ai:mainfrom
TyroneNel:tokenize-v1-route

Conversation

@TyroneNel

@TyroneNel TyroneNel commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

What

Adds patches/tokenize-v1-route.patch, its patches/series line and its PATCHES.md row. /tokenize and /detokenize are served under /v1 as well as at the root.

Upstream: vllm-project/vllm#58027.

Why

vLLM registers the tokenization endpoints at the root only. OpenAI SDK clients and gateways set base_url to a URL ending in /v1 and resolve endpoints against it, so /v1/tokenize is a real 404 for them. Measured on a live server from the current image: POST /tokenize 200, POST /v1/tokenize 404, and /openapi.json lists /tokenize only.

That matters here because bench/labd_accept.py:150 builds its teacher-forced prompts through /tokenize, and the headroom proxy in docker-compose.yml speaks an OpenAI base URL. One extra mount, and both work from one base URL.

Operation ids

Mounting one router twice is the usual way to get duplicate operationIds, so I checked instead of assuming. FastAPI 0.141.1, the same two include_router calls:

('/tokenize',      'post', 'tokenize_tokenize_post')
('/detokenize',    'post', 'detokenize_detokenize_post')
('/v1/tokenize',   'post', 'tokenize_v1_tokenize_post')
('/v1/detokenize', 'post', 'detokenize_v1_detokenize_post')
unique operation ids: True (4 routes)

Note: with --enable-tokenizer-info-endpoint, /tokenizer_info is on the same router and also gets the /v1 mount. This repo leaves that flag off.

Verification

  • patch integrity (the workflow's git-apply job) applies the whole series to a pristine vllm-project/vllm checkout at the pin: passes on this PR.
  • patch -p1 --fuzz 0 --dry-run of this file against the installed tree in ghcr.io/syv-ai/hyperqwen:latest (06150174): applies, no fuzz.
  • verify.sh needs no new entry: its loop reads patches/series, and patches/_check_applied.py parses the patch file itself.

Not done: I have not rebuilt the image and restarted a server on this patch, so the runtime evidence above comes from the code path, not from a rebuilt server.

Any OpenAI-SDK client (base_url ending in /v1) and most gateways resolve
every endpoint relative to /v1, so the root-only tokenization routes were
a genuine 404 for them. The patch includes the router twice — once at the
root, once with prefix="/v1" — with operation ids staying unique by
FastAPI's name+path+method derivation. Independent of every other patch,
appended at the series' block boundary. Cut from the extended cpuchip/vllm
qwen38/0.28 branch, topic commit [qwen38] tokenize-v1-route; kind:
feature, retires when upstream takes it. The verify.sh /v1/tokenize row
turns green with this patch installed.
@mhenrichsen

Copy link
Copy Markdown
Contributor

Merged (rebased for you).

The operation-id check is what made this easy to take. Mounting one router twice is the classic way to get duplicate ids and an invalid OpenAPI schema, and "I checked instead of assuming" with the four ids printed is the difference between a two-line patch I can merge and one I have to go verify myself.

The /tokenizer_info note is worth keeping in view: it rides the same router, so --enable-tokenizer-info-endpoint gets the /v1 mount too. This repo leaves that flag off, and after #169 it would need the key anyway, but anyone copying this patch onto a server that enables it is publishing a second path to it.

Concretely useful here beyond SDK clients: bench/labd_accept.py:150 builds its teacher-forced prompts through /tokenize, and the headroom proxy in docker-compose.yml speaks an OpenAI base URL — one extra mount and both work from one base URL instead of two.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants