Repository navigation
[Doc] Show how to get Responses prompt token IDs - #55912
Merged
DarkLight1337 merged 3 commits intoSep 16, 2026
Merged
Conversation
Contributor
|
Documentation preview: https://vllm--55912.org.readthedocs.build/en/55912/ |
franciscojavierarceo
force-pushed
the
responses-tokenize
branch
from
September 8, 2026 16:06
65d4575 to
664d7fc
Compare
Document using the existing Responses render endpoint to inspect prompt IDs before routing, with server setup, a curl/jq example, authentication, and a link from the tokenization guide. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Francisco Javier Arceo <farceo@redhat.com>
franciscojavierarceo
force-pushed
the
responses-tokenize
branch
from
September 8, 2026 17:19
664d7fc to
1af5b60
Compare
franciscojavierarceo
marked this pull request as ready for review
September 8, 2026 19:44
Describe general preprocessing capabilities in the renderer overview and place the Responses tokenization guidance in the Responses API section. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Francisco Javier Arceo <farceo@redhat.com>
DarkLight1337
approved these changes
Sep 14, 2026
DarkLight1337
enabled auto-merge (squash)
September 14, 2026 17:59
1 task done
Contributor
|
/ci run |
|
✅ Triggered Buildkite CI #89229 for commit |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
frontends can already get Responses prompt token IDs from
/v1/responses/render, added in #50195. this documents how to use it before choosing a model replica: enable the endpoint, send the full request, and extract token IDs or their count withjq.it expands the renderer overview to explain its preprocessing capabilities and links the example from the Responses API docs, where it clarifies that a separate
/tokenizecall is not needed. the guide covers API-key authentication, matching renderer configuration, stateless history, and retaining multimodal payloads for generation.follow-up to #40556. #50195 provides the implementation; this adds the token-inspection workflow. duplicate checks on September 8 found no other open PR documenting this workflow.
Validation
pre-commit passed:
both Bash examples pass
bash -n; the example JSON validates as aResponsesRequest.both
jqfilters return the expected IDs/count on a sample payload; the added links point to existing headings.git diff --checkpassed.only documentation changes. no model evaluation is needed. the live example remains unverified on this macOS host because vLLM cannot infer a supported device; a full docs build was not run locally.
AI assistance: OpenAI Codex helped edit and check the documentation.