Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 30 additions & 10 deletions docs/features/tool_calling.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,18 +109,22 @@ vLLM supports the `tool_choice='none'` option in the chat completion API. When t

## Constrained Decoding Behavior

Whether vLLM enforces the tool parameter schema during generation depends on the `tool_choice` mode and the per-tool `strict` field:
Structural-tag parsers resolve call obligation, grammar activation, and argument-schema enforcement separately. With structural-tag enforcement enabled:

| `tool_choice` value | Schema-constrained decoding | Behavior |
| `tool_choice` value | Call obligation | Structural-tag activation |
| --- | --- | --- |
| Named function | Yes (via structured outputs backend) | Arguments are guaranteed to be valid JSON conforming to the function's parameter schema. |
| `"required"` | Yes (via structured outputs backend) | Same as named function. The model must produce at least one tool call. |
| `"auto"` | Only when `strict: true` is set on at least one tool | Structural-tag parsers constrain tool-call arguments when a tool opts in with `strict: true`. Without it, the model generates freely and tool calls are extracted from raw text. |
| `"none"` | N/A | No tool calls are produced. |
| Named function | Call the selected function | Always |
| `"required"` | Produce at least one tool call | Always |
| `"auto"` | Tool calls are optional | When at least one tool sets `strict: true`, or `--tool-strict-level` is `function` or `parameter` |
| `"none"` | Tool calling is disabled | Disabled |

When a structural tag applies, each tool's declared parameter schema is enforced only when that tool sets `strict: true` or the server uses `--tool-strict-level parameter`. Tools with omitted or false `strict` receive broad argument-syntax constraints at levels `auto` and `function`, including for required and named calls.

For parsers using schema-derived JSON constraints, required and named calls continue to enforce the declared parameter schemas.

### Strict Mode

For `tool_choice="required"` or named function calling, structural-tag constraints are always applied regardless of the `strict` field. For `tool_choice="auto"`, setting `strict: true` on at least one tool opts in to structural-tag constraints; without it, the model generates freely and tool calls are extracted from raw text. The `strict` field is supported across all three API surfaces: Chat Completion, Responses, and Anthropic Messages.
Structural-tag constraints have two layers. The tool-call *envelope* (markup and function name) is constrained whenever a structural tag applies: always for `tool_choice="required"` and named function calling, and for `tool_choice="auto"` once at least one tool sets `strict: true` or the server raises the floor. The *argument schema* of an individual tool is pinned only when that tool sets `strict: true` (or under `--tool-strict-level parameter`); a tool that omits `strict` keeps its arguments unconstrained, even when another tool in the same request is strict. The `strict` field is supported across all three API surfaces: Chat Completion, Responses, and Anthropic Messages.

For best compatibility with strict schema enforcement, define tool parameter schemas in the OpenAI strict-schema style:

Expand All @@ -134,6 +138,22 @@ vLLM also provides a global toggle via the `VLLM_ENFORCE_STRICT_TOOL_CALLING` en
VLLM_ENFORCE_STRICT_TOOL_CALLING=false vllm serve ...
```

### Server-Side Strictness Floor

Most OpenAI-compatible clients and agent frameworks never set `strict` on their tools, so with `tool_choice="auto"` the model generates tool calls without any grammar and malformed markup can leak into the response. The `--tool-strict-level` option lets the server operator raise the floor for every request that carries tools, independently of what the client declares:

| Value | Behavior |
| --- | --- |
| `auto` (default) | Follow the request's tool choice and per-tool strictness. Required/named choices activate structural tags; `tool_choice="auto"` activates them when at least one tool sets `strict: true`. |
| `function` | Constrain the tool-call envelope (markup and the function name) for every request with tools, leaving argument contents free unless the client marked the tool `strict: true`. |
| `parameter` | Additionally pin argument schemas for every tool, as if every tool had `strict: true`. |

```bash
vllm serve ... --tool-strict-level function
```

The floor never relaxes a constraint the request would already receive: tools the client marked `strict: true` keep their schemas at every level. With `tool_choice="auto"`, the grammar does not force a tool call; a plain text response stays valid. `VLLM_ENFORCE_STRICT_TOOL_CALLING=false` disables structural tags entirely and takes precedence over this option.

## Automatic Function Calling

To enable this feature, you should set the following flags:
Expand All @@ -152,7 +172,7 @@ from HuggingFace; and you can find an example of this in a `tokenizer_config.jso
If your favorite tool-calling model is not supported, please feel free to contribute a parser & tool use chat template!

!!! note
With `tool_choice="auto"`, schema-level constraint requires both `VLLM_ENFORCE_STRICT_TOOL_CALLING=true` (the default) and at least one tool with `strict: true`. When these conditions are met and the selected parser supports structural tags, vLLM constrains tool-call arguments. Otherwise, vLLM extracts tool calls from raw text, so arguments may occasionally be malformed or violate the function's parameter schema.
With `tool_choice="auto"`, structural-tag constraints require both `VLLM_ENFORCE_STRICT_TOOL_CALLING=true` (the default) and at least one tool with `strict: true`, or a server-side floor set via `--tool-strict-level`. When these conditions are met and the selected parser supports structural tags, vLLM constrains the tool-call envelope and pins the argument schema of each tool that sets `strict: true` (or of every tool under `--tool-strict-level parameter`). Otherwise, vLLM extracts tool calls from raw text, so arguments may occasionally be malformed or violate the function's parameter schema.

### Hermes Models (`hermes`)

Expand All @@ -162,8 +182,8 @@ All Nous Research Hermes-series models newer than Hermes 2 Pro should be support
* `NousResearch/Hermes-2-Theta-*`
* `NousResearch/Hermes-3-*`

_Note that the Hermes 2 **Theta** models are known to have degraded tool call quality and capabilities due to the merge
step in their creation_.
*Note that the Hermes 2 **Theta** models are known to have degraded tool call quality and capabilities due to the merge
step in their creation*.

Flags: `--tool-call-parser hermes`

Expand Down
2 changes: 2 additions & 0 deletions rust/src/chat/src/backend/hf.rs
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,7 @@ impl ChatBackend for HfChatBackend {
self.tokenizer.clone(),
options.tool_call_parser,
options.reasoning_parser,
options.tool_strict_level,
)?))
}
}
Expand Down Expand Up @@ -343,6 +344,7 @@ mod tests {
let error = match backend.new_chat_output_processor(
&mut request,
NewChatOutputProcessorOptions {
tool_strict_level: crate::ToolStrictLevel::Auto,
tool_call_parser: &ParserSelection::Explicit("json".to_string()),
reasoning_parser: &ParserSelection::Auto,
},
Expand Down
3 changes: 2 additions & 1 deletion rust/src/chat/src/backend/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -13,14 +13,15 @@ use crate::multimodal::{MmLimitPerPrompt, MultimodalModelInfo};
use crate::output::DynChatOutputProcessor;
use crate::renderer::DynChatRenderer;
use crate::request::ChatRequest;
use crate::{ChatTemplateContentFormatOption, ParserSelection, RendererSelection};
use crate::{ChatTemplateContentFormatOption, ParserSelection, RendererSelection, ToolStrictLevel};

pub mod hf;

/// Options for creating a new chat output processor.
pub struct NewChatOutputProcessorOptions<'a> {
pub tool_call_parser: &'a ParserSelection,
pub reasoning_parser: &'a ParserSelection,
pub tool_strict_level: ToolStrictLevel,
}

/// Minimal prompt-processing backend needed by `vllm-chat`.
Expand Down
19 changes: 18 additions & 1 deletion rust/src/chat/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ pub use parser::reasoning::{
ReasoningDelta, ReasoningError, ReasoningParser, ReasoningParserFactory,
};
pub use parser::tool::{ToolParser, ToolParserError, ToolParserFactory};
pub use parser::{ParserSelection, validate_parser_overrides};
pub use parser::{ParserSelection, ToolStrictLevel, validate_parser_overrides};
pub use reasoning::EffortValue;
pub use renderer::hf::ChatTemplateContentFormatOption;
pub use renderer::{
Expand Down Expand Up @@ -74,6 +74,8 @@ pub struct ChatRequestProcessor {
tool_call_parser: ParserSelection,
/// Reasoning parser selection used when preparing generation requests.
reasoning_parser: ParserSelection,
/// Server-side floor for tool-call structural tags.
tool_strict_level: ToolStrictLevel,
}

impl ChatRequestProcessor {
Expand All @@ -85,6 +87,7 @@ impl ChatRequestProcessor {
model_dtype: Some(model_dtype),
tool_call_parser: ParserSelection::Auto,
reasoning_parser: ParserSelection::Auto,
tool_strict_level: ToolStrictLevel::Auto,
}
}

Expand All @@ -95,6 +98,7 @@ impl ChatRequestProcessor {
model_dtype: None,
tool_call_parser: ParserSelection::Auto,
reasoning_parser: ParserSelection::Auto,
tool_strict_level: ToolStrictLevel::Auto,
}
}

Expand All @@ -109,6 +113,12 @@ impl ChatRequestProcessor {
self
}

/// Configure the server-side floor for tool-call structural tags.
pub fn with_tool_strict_level(mut self, tool_strict_level: ToolStrictLevel) -> Self {
self.tool_strict_level = tool_strict_level;
self
}

async fn finalize_rendered_prompt(
&self,
request: &ChatRequest,
Expand Down Expand Up @@ -207,6 +217,7 @@ impl ChatRequestProcessor {
NewChatOutputProcessorOptions {
tool_call_parser: &self.tool_call_parser,
reasoning_parser: &self.reasoning_parser,
tool_strict_level: self.tool_strict_level,
},
)?;
let text_request = self.prepare_text_request(request).await?;
Expand Down Expand Up @@ -255,6 +266,12 @@ impl ChatLlm {
self
}

/// Set the server-side floor for tool-call structural tags.
pub fn with_tool_strict_level(mut self, tool_strict_level: ToolStrictLevel) -> Self {
self.processor.tool_strict_level = tool_strict_level;
self
}

/// Override the effective model dtype used for multimodal tensor encoding.
pub fn with_model_dtype(mut self, model_dtype: ModelDtype) -> Self {
self.processor.model_dtype = Some(model_dtype);
Expand Down
15 changes: 12 additions & 3 deletions rust/src/chat/src/output/default/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -18,10 +18,10 @@ use self::unified::unified_event_stream;
use super::structured::structured_chat_event_stream;
use crate::error::Result;
use crate::output::{ChatOutputProcessor, DynChatEventStream, DynDecodedTextEventStream};
use crate::parser::ParserSelection;
use crate::parser::reasoning::{ReasoningParser, ReasoningParserFactory};
use crate::parser::tool::{ToolParser, ToolParserFactory};
use crate::parser::unified::UnifiedParserFactory;
use crate::parser::{ParserSelection, ToolStrictLevel};
use crate::request::{ChatRequest, ChatTool};
use crate::{Error, Result as ChatResult};

Expand Down Expand Up @@ -49,6 +49,7 @@ impl DefaultChatOutputProcessor {
tokenizer: DynTokenizer,
tool_call_parser: &ParserSelection,
reasoning_parser: &ParserSelection,
tool_strict_level: ToolStrictLevel,
) -> ChatResult<Self> {
let parser = if let Some(parser) = Self::resolve_optional_unified_parser(
request.tools(),
Expand All @@ -73,7 +74,11 @@ impl DefaultChatOutputProcessor {
Box::new(CombinedParser::new(reasoning_parser, tool_parser)) as Box<dyn UnifiedParser>
};

apply_structural_tag_constraint(request, parser.structural_tag_builder())?;
apply_structural_tag_constraint(
request,
parser.structural_tag_builder(),
tool_strict_level,
)?;

if parser.preserve_special_tokens() {
request.decode_options.skip_special_tokens = false;
Expand Down Expand Up @@ -191,7 +196,7 @@ mod tests {
use vllm_tokenizer::test_utils::TestTokenizer;

use super::DefaultChatOutputProcessor;
use crate::parser::ParserSelection;
use crate::parser::{ParserSelection, ToolStrictLevel};
use crate::request::ChatRequest;

fn tokenizer() -> Arc<TestTokenizer> {
Expand All @@ -213,6 +218,7 @@ mod tests {
tokenizer(),
&selection,
&selection,
ToolStrictLevel::Auto,
)
.unwrap();
}
Expand All @@ -227,6 +233,7 @@ mod tests {
tokenizer(),
&ParserSelection::Auto,
&ParserSelection::Auto,
ToolStrictLevel::Auto,
)
.unwrap();
}
Expand All @@ -244,6 +251,7 @@ mod tests {
tokenizer(),
tool,
reasoning,
ToolStrictLevel::Auto,
)
.unwrap();
}
Expand All @@ -258,6 +266,7 @@ mod tests {
tokenizer(),
&ParserSelection::Auto,
&ParserSelection::Explicit("gemma4".to_string()),
ToolStrictLevel::Auto,
) {
Ok(_) => panic!("expected mixed Gemma4 parser selection to fail"),
Err(error) => error,
Expand Down
Loading
Loading