Skip to content

Add checkpoint-driven response-template adapters - #40479

Open
yonigozlan wants to merge 23 commits into
sgl-project:mainfrom
yonigozlan:feat/response-template-adapters
Open

yonigozlan wants to merge 23 commits into
sgl-project:mainfrom
yonigozlan:feat/response-template-adapters

Conversation

@yonigozlan

@yonigozlan yonigozlan commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Connect checkpoint response_template metadata to SGLang reasoning and tool parsing, so new response formats work without adding another parser implementation.

Context

Response templates are the output-side counterpart to chat templates. They describe how raw generated text becomes structured reasoning, content, and tool calls. The first two PRs in this stack provide the generic parser and its streaming event contract; this PR connects them to SGLang serving.

Parser selection

  • Explicit: --reasoning-parser response_template and --tool-call-parser response_template.
  • Auto: with --chat-template unset, checkpoint metadata fills only the parser slots that existing chat-template detection left unresolved. Invalid metadata is ignored with a warning.
  • Template precedence: an explicit response_template argument, then checkpoint metadata, then a detector subclass's built-in fallback. Named aliases stay thin subclasses instead of separate implementations.

Integration flow

  1. ReasoningParser and FunctionCallParser construct response-template detectors through their existing registries.
  2. ResponseTemplateStreamAdapter feeds generated text into chat_parsing.ResponseParser and routes its region events into SGLang reasoning, content, and tool-call results.
  3. Chat Completions and Responses keep their normal parser lifecycle and wire-format builders.

Reasoning and content

  • Derive a streaming template where supported fields are optional, allowing reasoning-only, content-only, or mixed output.
  • Route thinking chunks to reasoning and content chunks to assistant content. Ordinary text is preserved when a template has only an explicit thinking field.
  • Keep reasoning policy and delimiters configurable on detector subclasses.
  • Keep one-shot parsing isolated from streaming state, so both paths can share a detector.

Tool calls

  • Derive a tool-focused template that parses tool_calls and passes all other text through unchanged.
  • Stream a tool name as soon as its opener determines it, then its arguments when the region closes.
  • Drop policy: a call is emitted only when it ends with its closer, parses, and names a requested tool. Otherwise it is dropped with a warning. If its name was already streamed, that call stays incomplete, because a streamed name cannot be retracted.
  • A call cut off right after its opener is still detected and dropped, instead of returning the raw opener as content.
  • Unknown tool names are dropped unless SGLANG_FORWARD_UNKNOWN_TOOLS is set, like other SGLang detectors.
  • Incomplete calls: Responses marks the function-call item incomplete. Chat Completions keeps the engine finish reason instead of promoting it to tool_calls.
  • Support repeated and parallel calls, stable tool indices, calls opened by the prompt prefix, and schema coercion of string arguments.

Detokenization and prompt

  • While a response-template parser is active, requests keep special tokens and their exact spacing.
  • The detokenizer keeps a tool-call closer when it is the stop token, so a call that ends at its closer is not mistaken for a cut-off one.
  • Parsers receive the rendered assistant prefill, either the exact prompt text or the decoded prompt ids, so they can continue regions opened in the prompt. ReasoningParser and FunctionCallParser also accept an explicit prefix.

Validation and constraints

load_response_template validates the generic grammar. Serving adds a check for what SGLang can safely route:

  • Accept thinking, content, and tool_calls fields. The tool detector requires a tool_calls field.
  • Allow structured parsing and transforms for tool calls, whose values are emitted only after validation.
  • Reject structured or transformed semantics for reasoning and content, rather than silently losing streamed data.
  • Reuse SGLang's generic JSON-schema constraints for required or named tool choice.
  • Reject strict tools and parallel_tool_calls=False with automatic tool choice, since the template cannot enforce them during generation.

Compatibility coverage

The adapter is compared with the registered Gemma 4 reasoning and tool parsers across non-streaming and streaming output, repeated calls, prompt prefills, early tool-name emission, malformed or cut-off output, and arbitrary chunk boundaries. Existing explicit parsers keep precedence during auto detection.

Stacked on Enrich chat_parsing streaming events.

JSON-schema output and stream finalization

Required or named tool requests using the generic JSON fallback, and JSON response_format requests, can produce bare JSON even when every template field has an explicit opener. The adapter preserves that JSON after the reasoning boundary, or from the start when the grammar backend does not wait for reasoning. Template delimiters inside JSON strings remain data. Structural-tag, regex and EBNF output keeps the existing template parsing.

For JSON output, Chat and Responses trim a retained tool-call closer only when it is the matched stop. Chat's no_stop_trim option is honored.

The shared JsonArrayParser now flushes buffered calls at stream end. This fixes lost arguments when a complete call arrives in the final chunk. Regression coverage includes native framing, JSON literals, both API paths, multiple buffered calls and repeated finalization.


CI States

Latest PR Test (Base): Not run yet
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

@github-actions github-actions Bot added dependencies Pull requests that update a dependency file npu labels Sep 20, 2026
@yonigozlan
yonigozlan force-pushed the feat/response-template-adapters branch 3 times, most recently from b3b4804 to 635181a Compare September 22, 2026 19:09
tool_detector = FunctionCallParser.ToolCallParserEnum.get(self.tool_call_parser)
return (
request.tool_choice != "none"
and bool(self._effective_tools(request))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we think about handling ResponsesRequest without reading request.messages.

I believe if default tool_choice is set to "auto" and no response-template reasoning detector is used, this calls the chat-only helper and returns 500.

@@ -1195,6 +1208,8 @@ def _convert_to_internal_request(

# Process messages and apply chat template
processed_messages = self._process_messages(request, is_multimodal)
if self._requires_response_template_detokenization(request):
configure_response_template_request(request)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we also set skip_special_tokens=False and spaces_between_special_tokens=False for Responses requests using this parser?

tokenizer=None,
response_template: dict | None = None,
prefix: str | None = None,
**_kwargs,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we honor force_nonempty_content here? With this option enabled, a non-streaming response containing only reasoning still returns empty content.

if thinking and thinking.close_literals
else self._default_think_end
)
self.think_start_self_label = ""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we define think_excluded_tokens=None here? With --enable-strict-thinking, the grammar backend reads this missing attribute and raises an error

yonigozlan and others added 13 commits September 27, 2026 15:51
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
A `repeats` field appends into the default list object owned by the
compiled template (and, through the shallow copy, into the caller's
spec dict), so one parse's tool calls leak into every later parse
sharing the template. Deep-copy defaults at template load and at
parser init, and add a regression test.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Delimiter text comes from the input offsets, and open events keep named
captures. Unmatched text around XML tags no longer fails the parse.
Keep the event contract to what adapters need: open, close and malformed
events carry the span of input_text they consumed, explicit opens carry
their captures, and malformed events carry the original error.

- Drop per-chunk offsets, prefix_end, close_start and closed; the prefix
  boundary is len(input_text) before the first feed().
- parse_response re-raises the original exception, like Transformers.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
`parse_response` only checked events from the generated text, so a
region the prefix left malformed was silently ignored. Include
`initial_events` in the check.

Co-authored-by: Cursor <cursoragent@cursor.com>
Connect checkpoint `response_template` metadata to SGLang reasoning and
tool parsing through the `chat_parsing` streaming events.

Co-authored-by: Cursor <cursoragent@cursor.com>
- Configure response-template detokenization in `_process_messages`, so
  the Responses API gets the same special-token settings as Chat and no
  longer inspects `ResponsesRequest` as a chat request.
- Honor `force_nonempty_content` in the reasoning detector.
- Define `think_excluded_tokens` and `get_think_end_token_ids`, which
  strict thinking and scheduler startup read from every reasoning
  detector.
- Register `response_template` in the CLI parser name lists.

Co-authored-by: Cursor <cursoragent@cursor.com>
@yonigozlan
yonigozlan force-pushed the feat/response-template-adapters branch from f75b214 to 406d136 Compare September 27, 2026 20:23
With tool_choice "required" or a named tool, a response-template tool
parser has no native constraint, so serving uses the generic JSON schema
and parses the JSON array the model writes after its reasoning. Two things
broke that path:

- When every field of the template has an opener, the reasoning parser has
  no field for text outside them, so it dropped the JSON.
- Detokenization keeps a tool-call closer that stops generation, so a JSON
  array ended by one failed to parse, or streamed the closer as content.

Serving now records on the request whether a grammar governs the output
(tool JSON schema, response_format, regex, ebnf or structural tag). Only
then does the reasoning parser pass text outside the template's fields
through, and a kept closer is trimmed like any other matched stop. This
covers Chat and Responses, streaming and non-streaming.
Serving now records where an output grammar takes over from the response
template: after the reasoning when the grammar backend gates on it, or
from the first token. The response-template reasoning parser stops at the
reasoning closer, as the grammar backend does, and passes the rest through
unparsed, so a JSON string containing template delimiters is no longer
rewritten. Unconstrained requests keep the template's own parsing.

JsonArrayParser also gains finish(). An increment emits a call's name
before its arguments, so a call that arrived whole in the last chunk lost
its arguments; Chat and Responses streams now flush it at stream end like
the other tool parsers.
The grammar backend defers an output grammar until the reasoning ends only
when the reasoning parser's think_end_token encodes to token ids; otherwise
the grammar applies from the first token. A response-template thinking
field that closes with a pattern has no such token, yet a request asking
for reasoning was still marked to parse reasoning first, so the JSON the
grammar produced from the start was dropped. Check the same condition the
scheduler checks before recording where the grammar takes over.
@Jiminator Jiminator added the run-ci CI: run the baseline test suite on this PR label Sep 28, 2026
Jiminator and others added 4 commits September 28, 2026 00:07
Only JSON-schema output replaces the response template's framing. A
structural tag, regex or EBNF constraint can spell the template's own
delimiters, so its output keeps the template's parsing: content wrappers
are removed and a native tool call that ends on its kept closer is still
returned. The JSON repairs for required and named tool choice and
response_format are unchanged.

Also shorten the related comments and give the regression tests
black-box docstrings in place of assertions on private state.
The two non-streaming Gemma 4 tool-call parity tests differed only in
their input, so they are now subtests of one method, as the reasoning
parity test already is. _collect_tool_stream repeated its call
accumulation for the finish() result; one loop now covers every result.
The Gemma 4 call's expected arguments were written out three times and
are now named once next to the call.

Every input, detector and assertion is kept.
The result of feeding "}" was assigned to closed and immediately
overwritten by the "</call>" result; only the call itself matters.
`has_tool_call` streamed the text through the adapter, which holds back
an opener whose pattern could still grow. A response cut off right
after a call header was then treated as plain content and returned the
raw opener. Parse the text as complete instead, so the cut-off call is
dropped like any other malformed call.

Signed-off-by: Yoni Gozlan <yonigozlan@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file npu run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants