Skip to content

[Feature] Add /v1/decisions for typed decisions on candidate scoring - #40992

Closed
rwang5203 wants to merge 1 commit into
sgl-project:mainfrom
rwang5203:jev-decisions
Closed

rwang5203 wants to merge 1 commit into
sgl-project:mainfrom
rwang5203:jev-decisions

Conversation

@rwang5203

@rwang5203 rwang5203 commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Semantic decisions ask a model a fixed question about some input and need one answer plus a probability for every option, without parsing generated text. #40826 added the per-item candidate scoring that gives these probabilities, but each client still has to render the chat prompt, find answer labels that are single tokens, and keep thinking templates out of the way, and mistakes there change decisions without any error.

Changes

  • POST /v1/decisions answers typed questions (choice, score, and yes_no) about an input. The server renders each question with the served chat template and thinking off, checks that its answer labels (A to Z, 0 to 9, or yes and no) are single tokens at the answer position, and scores all questions in one score_prompts call without generating. Each answer returns the probability of every option and label_mass, plus choice for a choice question or score for a score question.
  • The server owns and versions the prompt wording, and return_prompt_token_ids lets any answer be replayed through /v1/score.
  • Templates that would put the answer inside a reasoning block, and setups the route cannot serve faithfully, such as --enable-mis, --dllm-algorithm, a LoRA adapter named in model, or labels that are not single tokens, return an actionable 400.
  • Three fixes found while validating. Scoring batched with plain generation no longer stops the scheduler, a grammar that fails to compile with logprob fields returns its 400 instead of stopping the server, and streaming chat returns each token's own top_logprobs inside a multi-token chunk.
  • A Decision models page documents the route, and the Qwen3.8-27B and Qwen3.5 model pages link to it.

Validation

On main, Qwen3.8-27B and Qwen3.5-35B-A3B in BF16, one H200 each, with the documented launch command.

  • Decisions are bitwise equal to /v1/score replays of the returned ids, cold cache against cold cache, for 23 tested questions, and all 144 SemIf rows return well formed answers on both models (138 and 130 correct, a sanity check only).
  • 56 concurrent mixed requests across decisions, /v1/score, and /generate return 200 on both models, on both model pages' deploy commands, and on the 27B under NEXTN. Main stops the scheduler on the same load without decisions.
  • Grammar failures with logprob fields return 400 on both models, and on the 35B streamed top_logprobs match the non-streamed ones for every streamed token, where main repeats the first token's list.
  • Every tested refusal returns 400 naming the question, field, or unsupported feature, followed by a healthy request. The Decision models examples run on both models, and each model page example on its own model.
  • On CPU, a refusal survey of 537 local chat tokenizer directories, copies counted separately, accepts both validated models and 437 in total. Unit tests cover the contract and the three fixes, whose tests fail on main. Pre-commit passes.

Limits

  • Nothing is calibrated, and label_mass for yes_no reads low because it counts lowercase yes and no only. Thresholds need labeled data from the workload.
  • Reasoning that neither the chat template nor the reasoning parser marks is not detected, and some templates of models that do not reason are refused.
  • On these hybrid attention models, answer probabilities moved by up to about 0.07 and label_mass by up to about 0.14 between cold and prefix-cached requests and across batch compositions, with the chosen option unchanged in our checks.
  • Leave any temperature in --preferred-sampling-params unset and --enable-mixed-chunk off for decisions, since the first can scale the probabilities and the second can reuse recurrent state never written for the prefix on these hybrid models.
  • The server does not check where the chat template puts an answer, and 8 accepted local checkpoints score one token early, so inspect prompt_token_ids before relying on a new model.

CI States

Latest PR Test (Base): ❌ Run #36054102942
Latest PR Test (Extra): ❌ Run #36054102642
Latest PR Test (AMD ROCm 10): ⏳ Run #36054102650

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 24, 2026
@rwang5203 rwang5203 added run-ci CI: run the baseline test suite on this PR highest-priority CI: all three control labels, plus never batch-cancelled or stale-closed labels Sep 24, 2026
@rwang5203 rwang5203 mentioned this pull request Sep 24, 2026
21 of 95 tasks
POST /v1/decisions answers typed questions about an input with the
probability of every answer, without generating. Questions are choice
(named options), score (ordered levels), and yes_no. Each question is
rendered as one user message with the served chat template, gets
one-token labels (A to Z, 0 to 9, yes and no) checked at the answer
position, and all questions of a request are scored in one score_prompts
call. Answers return the probabilities and label_mass, plus the chosen
option for choice or the expected level for score. A request temperature
scales the probabilities but not label_mass.

The server owns this prompt wording and versions it. Responses carry
prompt_format_version, a request can pin it, and return_prompt_token_ids
returns the scored ids so any answer can be replayed through /v1/score.

The route renders the way the chat route does. The reasoning toggle,
named by the chat template or, when detection finds none, by the
reasoning parser, is set off, server default chat template kwargs fill
other keys, and request kwargs apply last. A request toggle set to
anything but false gets a 400, as do templates that always reason or
start every answer with a reasoning block, a rendered prompt that leaves
a reasoning block open, and a model whose reasoning parser expects
answers to start inside a block the prompt does not close. Setups the
route cannot serve faithfully also get a 400: --enable-mis,
--dllm-algorithm, built-in named chat templates, Python chat encoders,
tokenizers whose chat text does not encode back to the same ids or that
split an answer label, and LoRA adapters. So do malformed requests.
Questions are encoded one at a time, and other requests run between
them.

Validating decisions also found two crashes and one wrong result, fixed
here:

- Keep token-ids logprobs tensors when only some requests in a batch ask
  for them, so move_logprobs_to_cpu no longer stops the scheduler when
  scoring is batched with plain generation.
- Treat absent logprob arrays as empty in convert_logprob_style, so a
  request whose grammar fails to compile returns its 400 instead of
  stopping the server.
- Give each token of a multi-token chat stream chunk its own
  top_logprobs row.

Docs: a Decision models page under Specialized Models with a card on the
Supported models overview, Semantic Decisions sections on the
Qwen3.8-27B and Qwen3.5 model pages, and native API entries for
/v1/decisions and the scoring fields it replays through.
@rwang5203

Copy link
Copy Markdown
Collaborator Author

Closing, as this change landed in main via #41208.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation highest-priority CI: all three control labels, plus never batch-cancelled or stale-closed run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant