Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,9 @@ jobs:
tests/test_gemma4_openai_format.py \
tests/test_gemma4_streaming_edge.py \
tests/test_gemma4_tool_parser.py \
tests/test_deepseek_v4_encoding.py \
tests/test_deepseek_v4_reasoning.py \
tests/test_deepseek_v4_tool_parser.py \
tests/test_minimax_tool_calling.py \
tests/test_qwen3_xml_parser.py \
tests/test_qwen3_xml_registration.py \
Expand Down
4 changes: 2 additions & 2 deletions README.es.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ claude
### APIs
- **Compatible con OpenAI**: `/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`, `/v1/rerank`, `/v1/responses`
- **Compatible con Anthropic**: `/v1/messages` (streaming, tool use, system prompts)
- **MCP Tool Calling**: 12 parsers (OpenAI, Anthropic, Gemini, Qwen, DeepSeek, Gemma y más)
- **MCP Tool Calling**: 19 parsers (OpenAI, Anthropic, Gemini, Qwen, DeepSeek, Gemma y más)
- **Salida estructurada**: JSON Schema vía `response_format` (lm-format-enforcer)

### Throughput y memoria
Expand All @@ -64,7 +64,7 @@ claude
- **STT**: familia Whisper con RTF hasta 197x en M4 Max

### Razonamiento y avanzado
- **Extracción de razonamiento**: Qwen3, DeepSeek-R1 (`--reasoning-parser`)
- **Extracción de razonamiento**: Qwen3, DeepSeek-R1, DeepSeek-V4 (`--reasoning-parser`)
- **Reducción de expertos MoE**: `--moe-top-k` para +7-16% en Qwen3-30B-A3B
- **Decodificación especulativa**: `--mtp` para Qwen3-Next
- **Prefill disperso**: `--spec-prefill` basado en atención para reducir TTFT
Expand Down
4 changes: 2 additions & 2 deletions README.fr.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ claude
### APIs
- **Compatible OpenAI** : `/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`, `/v1/rerank`, `/v1/responses`
- **Compatible Anthropic** : `/v1/messages` (streaming, tool use, system prompts)
- **MCP Tool Calling** : 12 parsers (OpenAI, Anthropic, Gemini, Qwen, DeepSeek, Gemma et plus)
- **MCP Tool Calling** : 19 parsers (OpenAI, Anthropic, Gemini, Qwen, DeepSeek, Gemma et plus)
- **Sortie structurée** : JSON Schema via `response_format` (lm-format-enforcer)

### Débit et mémoire
Expand All @@ -64,7 +64,7 @@ claude
- **STT** : famille Whisper avec RTF jusqu'à 197x sur M4 Max

### Raisonnement et avancé
- **Extraction du raisonnement** : Qwen3, DeepSeek-R1 (`--reasoning-parser`)
- **Extraction du raisonnement** : Qwen3, DeepSeek-R1, DeepSeek-V4 (`--reasoning-parser`)
- **Réduction d'experts MoE** : `--moe-top-k` pour +7-16% sur Qwen3-30B-A3B
- **Décodage spéculatif** : `--mtp` pour Qwen3-Next
- **Prefill creux** : `--spec-prefill` basé sur l'attention pour réduire le TTFT
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ claude
### APIs
- **OpenAI-compatible**: `/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`, `/v1/rerank`, `/v1/responses`
- **Anthropic-compatible**: `/v1/messages` (streaming, tool use, system prompts)
- **MCP Tool Calling**: 12 parsers (OpenAI, Anthropic, Gemini, Qwen, DeepSeek, Gemma, and more)
- **MCP Tool Calling**: 19 parsers (OpenAI, Anthropic, Gemini, Qwen, DeepSeek, Gemma, and more)
- **Structured output**: JSON Schema via `response_format` (lm-format-enforcer)

### Throughput & memory
Expand All @@ -64,7 +64,7 @@ claude
- **STT**: Whisper family with RTF up to 197x on M4 Max

### Reasoning & advanced
- **Reasoning extraction**: Qwen3, DeepSeek-R1 (`--reasoning-parser`)
- **Reasoning extraction**: Qwen3, DeepSeek-R1, DeepSeek-V4 (`--reasoning-parser`)
- **MoE expert reduction**: `--moe-top-k` for +7-16% on Qwen3-30B-A3B
- **Speculative decoding**: `--mtp` for Qwen3-Next
- **Sparse prefill**: attention-based `--spec-prefill` for TTFT reduction
Expand Down
4 changes: 2 additions & 2 deletions README.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ claude
### API
- **兼容 OpenAI**:`/v1/chat/completions`、`/v1/completions`、`/v1/embeddings`、`/v1/rerank`、`/v1/responses`
- **兼容 Anthropic**:`/v1/messages`(流式、工具调用、system prompts)
- **MCP 工具调用**:12 种解析器(OpenAI、Anthropic、Gemini、Qwen、DeepSeek、Gemma 等)
- **MCP 工具调用**:19 种解析器(OpenAI、Anthropic、Gemini、Qwen、DeepSeek、Gemma 等)
- **结构化输出**:通过 `response_format` 的 JSON Schema(基于 lm-format-enforcer)

### 吞吐与内存
Expand All @@ -64,7 +64,7 @@ claude
- **STT**:Whisper 系列,M4 Max 上 RTF 最高可达 197 倍

### 推理与高级功能
- **思维链提取**:Qwen3、DeepSeek-R1(`--reasoning-parser`)
- **思维链提取**:Qwen3、DeepSeek-R1、DeepSeek-V4(`--reasoning-parser`)
- **MoE 专家裁剪**:`--moe-top-k`,Qwen3-30B-A3B 上 +7-16%
- **投机解码**:`--mtp`,用于 Qwen3-Next
- **稀疏 prefill**:基于注意力的 `--spec-prefill`,降低 TTFT
Expand Down
214 changes: 214 additions & 0 deletions benchmarks/bench_deepseek_v4.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,214 @@
"""Benchmark: DeepSeek-V4 prompt encoding and streaming parser overhead.

Covers both serving paths. Single-stream is one sequence decoded token by
token, where the parser sits directly in the latency path. Batch is N
concurrent sequences interleaved, the shape continuous batching produces: each
request carries its own parser state, so per-token cost is paid N times per
decode step and any per-token work that scales with output length compounds.

The thing to watch for is quadratic scaling. A parser that rescans the whole
accumulated text on every delta is O(N²) over a generation, which is invisible
on short replies and dominant on long ones — so the ms/tok column is the one
that matters, not the totals. It is currently flat for the DSML tool parser and
grows for the reasoning path; see the notes the script prints at the end.

Usage:
python benchmarks/bench_deepseek_v4.py
"""

import time

from vllm_mlx.reasoning.deepseek_v4_parser import DeepSeekV4ReasoningParser
from vllm_mlx.tool_parsers.deepseek_v4_tool_parser import DeepSeekV4ToolParser
from vllm_mlx.utils.deepseek_v4_encoding import apply_chat_template

D = "|DSML|"

TOOLS = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string"},
"days": {"type": "integer"},
},
"required": ["city"],
},
},
}
]


def reasoning_tokens(n: int) -> list[str]:
"""A plain thinking turn: N reasoning tokens, then a short answer."""
return (
[f"step{i} " for i in range(n)]
+ ["</think>"]
+ [f"word{i} " for i in range(10)]
)


def tool_call_tokens(n: int) -> list[str]:
"""A tool-calling turn: N reasoning tokens, then DSML markup.

The markup is split the way the tokenizer splits it — |DSML| has an id of
its own, the surrounding punctuation does not — so the parser sees the same
fragmentation it sees in production.
"""
tokens = [f"step{i} " for i in range(n)] + ["</think>", "\n\n"]
tokens += ["<", D, "tool_calls", ">", "\n"]
tokens += ["<", D, "invoke", ' name="get_weather"', ">", "\n"]
tokens += [
"<",
D,
"parameter",
' name="city"',
' string="true"',
">",
"Prague",
"</",
D,
"parameter",
">",
"\n",
]
tokens += ["</", D, "invoke", ">", "\n"]
tokens += ["</", D, "tool_calls", ">"]
return tokens


def bench_stream(make_tokens, n_tokens: int, streams: int) -> tuple[float, int]:
"""Interleave `streams` concurrent sequences, one delta each per step.

Returns (total ms, total deltas). With streams=1 this is the single-stream
latency path; above that it is the batched decode step.
"""
token_lists = [make_tokens(n_tokens) for _ in range(streams)]
parsers = []
for _ in range(streams):
reasoner = DeepSeekV4ReasoningParser()
reasoner.reset_state()
tools = DeepSeekV4ToolParser()
tools.reset()
parsers.append((reasoner, tools, {"acc": "", "tool_acc": ""}))

steps = max(len(t) for t in token_lists)
deltas = 0
start = time.perf_counter()
for step in range(steps):
for stream in range(streams):
tokens = token_lists[stream]
if step >= len(tokens):
continue
reasoner, tools, state = parsers[stream]
delta = tokens[step]
previous, state["acc"] = state["acc"], state["acc"] + delta
deltas += 1

message = reasoner.extract_reasoning_streaming(
previous, state["acc"], delta
)
if message is None or not message.content:
continue
prev_tool = state["tool_acc"]
state["tool_acc"] = prev_tool + message.content
tools.extract_tool_calls_streaming(
prev_tool, state["tool_acc"], message.content
)
return (time.perf_counter() - start) * 1000, deltas


def bench_tool_parser_only(make_tokens, n_tokens: int) -> tuple[float, int]:
"""The DSML parser without the reasoning parser in front of it.

Isolates how much of the per-token cost is the tool parser's own work.
"""
tokens = make_tokens(n_tokens)
parser = DeepSeekV4ToolParser()
parser.reset()

accumulated = ""
start = time.perf_counter()
for delta in tokens:
previous, accumulated = accumulated, accumulated + delta
parser.extract_tool_calls_streaming(previous, accumulated, delta)
return (time.perf_counter() - start) * 1000, len(tokens)


def bench_encoder(turns: int, repeats: int = 200) -> float:
"""Prompt build cost for a conversation of `turns` user/assistant pairs."""
conversation = [{"role": "system", "content": "You are a helpful assistant."}]
for i in range(turns):
conversation.append({"role": "user", "content": f"Question number {i}?"})
conversation.append(
{
"role": "assistant",
"content": f"Answer number {i}.",
"reasoning_content": f"Thinking about question {i} at some length.",
}
)
conversation.append({"role": "user", "content": "And finally?"})

start = time.perf_counter()
for _ in range(repeats):
apply_chat_template(conversation, tools=TOOLS)
return (time.perf_counter() - start) * 1000 / repeats


def main():
print("DeepSeek-V4 encoder and parser benchmark")
print("=" * 68)

print("\nPrompt encoding (per call, tools attached)")
for turns in (1, 4, 16, 64):
ms = bench_encoder(turns)
print(f" {turns * 2 + 2:>4} messages -> {ms:>8.3f} ms")

for label, make_tokens in (
("plain thinking turn", reasoning_tokens),
("tool-calling turn", tool_call_tokens),
):
print(f"\nSingle stream, {label}")
for n in (100, 500, 1000, 2000, 5000):
ms, deltas = bench_stream(make_tokens, n, streams=1)
print(
f" {n:>5} reasoning tokens -> {ms:>8.2f} ms total, "
f"{ms / deltas:>7.4f} ms/tok"
)

print("\nDSML tool parser alone, tool-calling turn")
for n in (100, 500, 1000, 2000, 5000):
ms, deltas = bench_tool_parser_only(tool_call_tokens, n)
print(
f" {n:>5} reasoning tokens -> {ms:>8.2f} ms total, "
f"{ms / deltas:>7.4f} ms/tok"
)

print("\nBatched decode, tool-calling turn, 1000 reasoning tokens each")
for streams in (1, 2, 4, 8, 16):
ms, deltas = bench_stream(tool_call_tokens, 1000, streams=streams)
print(
f" {streams:>3} concurrent -> {ms:>8.2f} ms total, "
f"{ms / deltas:>7.4f} ms/tok"
)

print("\nReading the numbers:")
print(" At 50 tok/s the per-token budget is 20 ms, so anything under")
print(" 0.1 ms/tok is noise. Batching adds no per-token cost — each")
print(" request carries independent parser state and the totals scale")
print(" linearly with the number of streams.")
print()
print(" The single-stream ms/tok does grow with output length. That comes")
print(" from BaseThinkingReasoningParser, which searches the accumulated")
print(" text for its start and end tags on every delta while the reasoning")
print(" block is open; the DSML tool parser on its own stays flat. It is")
print(" 0.2% of the decode budget even at 5000 tokens, but it is quadratic,")
print(" so it is worth fixing in the base class rather than per model.")


if __name__ == "__main__":
main()
2 changes: 1 addition & 1 deletion docs/development/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -135,7 +135,7 @@ vllm_mlx/
├── models/
│ ├── llm.py # MLXLanguageModel
│ └── mllm.py # MLXMultimodalLM
├── tool_parsers/ # Tool call parsers (12 formats)
├── tool_parsers/ # Tool call parsers (19 formats)
├── reasoning_parsers/ # Reasoning parsers (qwen3, deepseek_r1)
├── server.py # FastAPI server
├── engine_core.py # AsyncEngineCore
Expand Down
2 changes: 1 addition & 1 deletion docs/es/guides/mcp-tools.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,7 +228,7 @@ python examples/mcp_chat.py

## Formatos de herramientas soportados

vllm-mlx soporta 12 tool call parsers que cubren todas las familias de modelos principales. Consulta [Tool Calling](tool-calling.md) para ver la lista completa de parsers, alias y ejemplos.
vllm-mlx soporta 19 tool call parsers que cubren todas las familias de modelos principales. Consulta [Tool Calling](tool-calling.md) para ver la lista completa de parsers, alias y ejemplos.

## Seguridad

Expand Down
2 changes: 1 addition & 1 deletion docs/fr/guides/mcp-tools.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,7 +228,7 @@ python examples/mcp_chat.py

## Formats d'outils pris en charge

vllm-mlx prend en charge 12 tool call parsers couvrant toutes les grandes familles de modèles. Voir [Tool Calling](tool-calling.md) pour la liste complète des parsers, alias et exemples.
vllm-mlx prend en charge 19 tool call parsers couvrant toutes les grandes familles de modèles. Voir [Tool Calling](tool-calling.md) pour la liste complète des parsers, alias et exemples.

## Sécurité

Expand Down
2 changes: 1 addition & 1 deletion docs/guides/mcp-tools.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,7 +228,7 @@ python examples/mcp_chat.py

## Supported Tool Formats

vllm-mlx supports 12 tool call parsers covering all major model families. See [Tool Calling](tool-calling.md) for the full list of parsers, aliases, and examples.
vllm-mlx supports 19 tool call parsers covering all major model families. See [Tool Calling](tool-calling.md) for the full list of parsers, aliases, and examples.

## Security

Expand Down
18 changes: 18 additions & 0 deletions docs/guides/reasoning.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,6 +131,24 @@ For DeepSeek-R1 models that may omit the opening `<think>` tag.
vllm-mlx serve mlx-community/DeepSeek-R1-Distill-Qwen-7B-4bit --reasoning-parser deepseek_r1
```

### DeepSeek-V4 Parser (`deepseek_v4`)

DeepSeek V4 Flash can end an implicit reasoning block by opening a DSML tool
call, even when `</think>` is absent. Use the V4 reasoning parser together with
the V4 tool parser so the DSML tail is routed to structured tool-call output:

```bash
vllm-mlx serve deepseek-ai/DeepSeek-V4-Flash-0731 \
--reasoning-parser deepseek_v4 \
--enable-auto-tool-choice \
--tool-call-parser deepseek_v4
```

The encoder detects the published preview and 0731 reasoning-effort profiles.
For OpenAI requests, unspecified effort selects `high`; `minimal`, `low`, and
`medium` select `low`; `high` and `xhigh` select `high`; and `max` selects
`max`. The preview profile normalizes `low` to its `high` behavior.

## How It Works

The reasoning parser uses text-based detection to identify thinking tags in the model output. During streaming, it tracks the current position in the output to correctly route each token to either `reasoning` or `content`.
Expand Down
Loading
Loading