Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions docs/cookbook/autoregressive/Qwen/Qwen3.5.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -777,6 +777,37 @@ Tool Call: None
Arguments: }
```

#### 4.3.3 Semantic Decisions

Qwen3.5-35B-A3B can answer typed questions about an input with a probability for every option through `/v1/decisions`, without generating text. The server turns thinking off for these requests, so the answer is read after a closed think block, and the same server keeps serving chat traffic. No launch flag is needed. Until a release contains `/v1/decisions`, install a [nightly build](/docs/get-started/install#nightly-builds). With a server started by `sglang serve --model-path Qwen/Qwen3.5-35B-A3B --host 127.0.0.1 --port 30000`:

```python Example
import requests

response = requests.post(
"http://127.0.0.1:30000/v1/decisions",
json={
"input": "I've been trying to connect my Stripe account for 3 days and the integration keeps failing.",
"questions": [
{
"id": "team",
"type": "choice",
"question": "Which team should handle this ticket?",
"options": [{"name": "billing"}, {"name": "technical"}, {"name": "sales"}],
},
{"id": "urgent", "type": "yes_no", "question": "The customer needs an answer today."},
],
},
timeout=60,
)
response.raise_for_status()
answers = response.json()["answers"]
print(answers["team"]["choice"], answers["team"]["probabilities"])
print(answers["urgent"]["probabilities"]["yes"])
```

See [Decision models](/docs/supported-models/decision_models) for the request reference, how the probabilities are computed, and the limits.

## 5. Benchmark

### 5.1 Accuracy Benchmark
Expand Down
31 changes: 31 additions & 0 deletions docs/cookbook/autoregressive/Qwen/Qwen3.8-27B.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -541,3 +541,34 @@ providers:
```

</Accordion>

## 4. Semantic Decisions

Qwen3.8-27B can answer typed questions about an input with a probability for every option through `/v1/decisions`, without generating text. The server turns thinking off for these requests, so the answer is read after a closed think block, and the same server keeps serving chat traffic. No launch flag is needed. Until a release contains `/v1/decisions`, install a [nightly build](/docs/get-started/install#nightly-builds). With a server started by `sglang serve --model-path Qwen/Qwen3.8-27B --host 127.0.0.1 --port 30000`:

```python Example
import requests

response = requests.post(
"http://127.0.0.1:30000/v1/decisions",
json={
"input": "I've been trying to connect my Stripe account for 3 days and the integration keeps failing.",
"questions": [
{
"id": "team",
"type": "choice",
"question": "Which team should handle this ticket?",
"options": [{"name": "billing"}, {"name": "technical"}, {"name": "sales"}],
},
{"id": "urgent", "type": "yes_no", "question": "The customer needs an answer today."},
],
},
timeout=60,
)
response.raise_for_status()
answers = response.json()["answers"]
print(answers["team"]["choice"], answers["team"]["probabilities"])
print(answers["urgent"]["probabilities"]["yes"])
```

See [Decision models](/docs/supported-models/decision_models) for the request reference, how the probabilities are computed, and the limits.
3 changes: 2 additions & 1 deletion docs/docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -998,7 +998,8 @@
{
"group": "Specialized Models",
"pages": [
"docs/supported-models/reward_models"
"docs/supported-models/reward_models",
"docs/supported-models/decision_models"
]
},
{
Expand Down
17 changes: 17 additions & 0 deletions docs/docs/basic_usage/native_api.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ Apart from the OpenAI compatible APIs, the SGLang Runtime also provides its nati
- `/encode`(embedding model)
- `/v1/rerank`(cross encoder rerank model)
- `/v1/score`(decoder-only scoring)
- `/v1/decisions`(typed decisions on a chat model)
- `/classify`(reward model)
- `/start_expert_distribution_record`
- `/stop_expert_distribution_record`
Expand Down Expand Up @@ -256,6 +257,8 @@ Parameters:
- `item_first`: Whether items come first in concatenation order (default: False)
- `model`: Model name

`query` and `items` also accept token ids, with `query` set to `[]` when each item is a complete prompt. `label_token_ids` can be one list per item. `temperature` divides the label logits before the softmax and needs `apply_softmax`, and `return_token_logprobs` adds `token_logprobs`, the full-vocabulary log-probabilities of the labels.

The response contains `scores` - a list of probability lists, one per item, each in the order of `label_token_ids`.

```python Example
Expand Down Expand Up @@ -293,6 +296,20 @@ for item, scores in zip(items, response_json["scores"]):
terminate_process(score_process)
```

## v1/decisions (typed decisions on a chat model)

Answer typed questions about an input with per-option probabilities, without generating. The server renders each question with the model's chat template and thinking turned off, assigns one-token answer labels, and scores them through the same path as `v1/score`.

Parameters:
- `input`: Text, object, or array the questions are about
- `questions`: List of questions, each with a unique `id` and a `type` of `choice` (with `options`), `score` (with `levels`), or `yes_no`
- `temperature`: Divides the label logits before the softmax (default: 1)
- `chat_template_kwargs`: Extra chat template arguments
- `prompt_format_version`: Optional pin of the server-owned prompt wording
- `return_prompt_token_ids`: Return the scored prompt and label token ids for replay through `v1/score` (default: False)

The response contains `answers` keyed by question id, each with `type`, `probabilities`, and `label_mass`, plus `choice` for a choice question or `score` for a score question. A yes or no answer is read from `probabilities["yes"]`. See [Decision models](/docs/supported-models/decision_models) for examples, the full reference, and error cases.

## Classify (reward model)

SGLang Runtime also supports reward models. Here we use a reward model to classify the quality of pairwise generations.
Expand Down
9 changes: 9 additions & 0 deletions docs/docs/supported-models.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -84,4 +84,13 @@ SGLang supports model families across text generation, retrieval, and reward wor
>
RLHF and reward scoring pipelines optimized for production latency.
</Card>
<Card
title="Decision models"
mode="card"
className="max-w-sm mx-auto"
href="./supported-models/decision_models"
icon="list-check"
>
Typed choice, score, and yes or no answers with a probability for every option.
</Card>
</CardGroup>
Loading
Loading