Repository navigation
Add checkpoint-driven response-template adapters - #40479
yonigozlan wants to merge 23 commits into
Conversation
b3b4804 to
635181a
Compare
| tool_detector = FunctionCallParser.ToolCallParserEnum.get(self.tool_call_parser) | ||
| return ( | ||
| request.tool_choice != "none" | ||
| and bool(self._effective_tools(request)) |
There was a problem hiding this comment.
Could we think about handling ResponsesRequest without reading request.messages.
I believe if default tool_choice is set to "auto" and no response-template reasoning detector is used, this calls the chat-only helper and returns 500.
| @@ -1195,6 +1208,8 @@ def _convert_to_internal_request( | |||
|
|
|||
| # Process messages and apply chat template | |||
| processed_messages = self._process_messages(request, is_multimodal) | |||
| if self._requires_response_template_detokenization(request): | |||
| configure_response_template_request(request) | |||
There was a problem hiding this comment.
Could we also set skip_special_tokens=False and spaces_between_special_tokens=False for Responses requests using this parser?
| tokenizer=None, | ||
| response_template: dict | None = None, | ||
| prefix: str | None = None, | ||
| **_kwargs, |
There was a problem hiding this comment.
Could we honor force_nonempty_content here? With this option enabled, a non-streaming response containing only reasoning still returns empty content.
| if thinking and thinking.close_literals | ||
| else self._default_think_end | ||
| ) | ||
| self.think_start_self_label = "" |
There was a problem hiding this comment.
Could we define think_excluded_tokens=None here? With --enable-strict-thinking, the grammar backend reads this missing attribute and raises an error
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
A `repeats` field appends into the default list object owned by the compiled template (and, through the shallow copy, into the caller's spec dict), so one parse's tool calls leak into every later parse sharing the template. Deep-copy defaults at template load and at parser init, and add a regression test.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Delimiter text comes from the input offsets, and open events keep named captures. Unmatched text around XML tags no longer fails the parse.
Keep the event contract to what adapters need: open, close and malformed events carry the span of input_text they consumed, explicit opens carry their captures, and malformed events carry the original error. - Drop per-chunk offsets, prefix_end, close_start and closed; the prefix boundary is len(input_text) before the first feed(). - parse_response re-raises the original exception, like Transformers. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
`parse_response` only checked events from the generated text, so a region the prefix left malformed was silently ignored. Include `initial_events` in the check. Co-authored-by: Cursor <cursoragent@cursor.com>
Connect checkpoint `response_template` metadata to SGLang reasoning and tool parsing through the `chat_parsing` streaming events. Co-authored-by: Cursor <cursoragent@cursor.com>
- Configure response-template detokenization in `_process_messages`, so the Responses API gets the same special-token settings as Chat and no longer inspects `ResponsesRequest` as a chat request. - Honor `force_nonempty_content` in the reasoning detector. - Define `think_excluded_tokens` and `get_think_end_token_ids`, which strict thinking and scheduler startup read from every reasoning detector. - Register `response_template` in the CLI parser name lists. Co-authored-by: Cursor <cursoragent@cursor.com>
f75b214 to
406d136
Compare
With tool_choice "required" or a named tool, a response-template tool parser has no native constraint, so serving uses the generic JSON schema and parses the JSON array the model writes after its reasoning. Two things broke that path: - When every field of the template has an opener, the reasoning parser has no field for text outside them, so it dropped the JSON. - Detokenization keeps a tool-call closer that stops generation, so a JSON array ended by one failed to parse, or streamed the closer as content. Serving now records on the request whether a grammar governs the output (tool JSON schema, response_format, regex, ebnf or structural tag). Only then does the reasoning parser pass text outside the template's fields through, and a kept closer is trimmed like any other matched stop. This covers Chat and Responses, streaming and non-streaming.
Serving now records where an output grammar takes over from the response template: after the reasoning when the grammar backend gates on it, or from the first token. The response-template reasoning parser stops at the reasoning closer, as the grammar backend does, and passes the rest through unparsed, so a JSON string containing template delimiters is no longer rewritten. Unconstrained requests keep the template's own parsing. JsonArrayParser also gains finish(). An increment emits a call's name before its arguments, so a call that arrived whole in the last chunk lost its arguments; Chat and Responses streams now flush it at stream end like the other tool parsers.
The grammar backend defers an output grammar until the reasoning ends only when the reasoning parser's think_end_token encodes to token ids; otherwise the grammar applies from the first token. A response-template thinking field that closes with a pattern has no such token, yet a request asking for reasoning was still marked to parse reasoning first, so the JSON the grammar produced from the start was dropped. Check the same condition the scheduler checks before recording where the grammar takes over.
Only JSON-schema output replaces the response template's framing. A structural tag, regex or EBNF constraint can spell the template's own delimiters, so its output keeps the template's parsing: content wrappers are removed and a native tool call that ends on its kept closer is still returned. The JSON repairs for required and named tool choice and response_format are unchanged. Also shorten the related comments and give the regression tests black-box docstrings in place of assertions on private state.
The two non-streaming Gemma 4 tool-call parity tests differed only in their input, so they are now subtests of one method, as the reasoning parity test already is. _collect_tool_stream repeated its call accumulation for the finish() result; one loop now covers every result. The Gemma 4 call's expected arguments were written out three times and are now named once next to the call. Every input, detector and assertion is kept.
The result of feeding "}" was assigned to closed and immediately overwritten by the "</call>" result; only the call itself matters.
`has_tool_call` streamed the text through the adapter, which holds back an opener whose pattern could still grow. A response cut off right after a call header was then treated as plain content and returned the raw opener. Parse the text as complete instead, so the cut-off call is dropped like any other malformed call. Signed-off-by: Yoni Gozlan <yonigozlan@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Connect checkpoint
response_templatemetadata to SGLang reasoning and tool parsing, so new response formats work without adding another parser implementation.Context
Response templates are the output-side counterpart to chat templates. They describe how raw generated text becomes structured reasoning, content, and tool calls. The first two PRs in this stack provide the generic parser and its streaming event contract; this PR connects them to SGLang serving.
Parser selection
--reasoning-parser response_templateand--tool-call-parser response_template.--chat-templateunset, checkpoint metadata fills only the parser slots that existing chat-template detection left unresolved. Invalid metadata is ignored with a warning.response_templateargument, then checkpoint metadata, then a detector subclass's built-in fallback. Named aliases stay thin subclasses instead of separate implementations.Integration flow
ReasoningParserandFunctionCallParserconstruct response-template detectors through their existing registries.ResponseTemplateStreamAdapterfeeds generated text intochat_parsing.ResponseParserand routes its region events into SGLang reasoning, content, and tool-call results.Reasoning and content
thinkingchunks to reasoning andcontentchunks to assistant content. Ordinary text is preserved when a template has only an explicit thinking field.Tool calls
tool_callsand passes all other text through unchanged.SGLANG_FORWARD_UNKNOWN_TOOLSis set, like other SGLang detectors.incomplete. Chat Completions keeps the engine finish reason instead of promoting it totool_calls.Detokenization and prompt
ReasoningParserandFunctionCallParseralso accept an explicitprefix.Validation and constraints
load_response_templatevalidates the generic grammar. Serving adds a check for what SGLang can safely route:thinking,content, andtool_callsfields. The tool detector requires atool_callsfield.parallel_tool_calls=Falsewith automatic tool choice, since the template cannot enforce them during generation.Compatibility coverage
The adapter is compared with the registered Gemma 4 reasoning and tool parsers across non-streaming and streaming output, repeated calls, prompt prefills, early tool-name emission, malformed or cut-off output, and arbitrary chunk boundaries. Existing explicit parsers keep precedence during
autodetection.Stacked on Enrich chat_parsing streaming events.
JSON-schema output and stream finalization
Required or named tool requests using the generic JSON fallback, and JSON
response_formatrequests, can produce bare JSON even when every template field has an explicit opener. The adapter preserves that JSON after the reasoning boundary, or from the start when the grammar backend does not wait for reasoning. Template delimiters inside JSON strings remain data. Structural-tag, regex and EBNF output keeps the existing template parsing.For JSON output, Chat and Responses trim a retained tool-call closer only when it is the matched stop. Chat's
no_stop_trimoption is honored.The shared
JsonArrayParsernow flushes buffered calls at stream end. This fixes lost arguments when a complete call arrives in the final chunk. Regression coverage includes native framing, JSON literals, both API paths, multiple buffered calls and repeated finalization.CI States
Latest PR Test (Base): Not run yet⚠️ Not enabled -- add
Latest PR Test (Extra):
run-ci-extralabel to opt in.Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.