Repository navigation
Conversation
rwang5203
requested review from
BBuf,
CatherineSue,
Edwardf0t1,
Fridge003,
HaiShaw,
JustinTong0323,
Ying1123,
ch-wan,
hnyls2002,
ispobock,
merrymercy,
slin1237,
sogalin,
sundar24295s,
wisclmy0611,
xiezhq-hermann and
zijiexia
as code owners
September 24, 2026 00:38
rwang5203
force-pushed
the
jev-decisions
branch
from
September 24, 2026 02:06
3fafddd to
db11f0d
Compare
rwang5203
force-pushed
the
jev-decisions
branch
from
September 24, 2026 05:29
ee07bff to
0adcd44
Compare
21 of 95 tasks
POST /v1/decisions answers typed questions about an input with the probability of every answer, without generating. Questions are choice (named options), score (ordered levels), and yes_no. Each question is rendered as one user message with the served chat template, gets one-token labels (A to Z, 0 to 9, yes and no) checked at the answer position, and all questions of a request are scored in one score_prompts call. Answers return the probabilities and label_mass, plus the chosen option for choice or the expected level for score. A request temperature scales the probabilities but not label_mass. The server owns this prompt wording and versions it. Responses carry prompt_format_version, a request can pin it, and return_prompt_token_ids returns the scored ids so any answer can be replayed through /v1/score. The route renders the way the chat route does. The reasoning toggle, named by the chat template or, when detection finds none, by the reasoning parser, is set off, server default chat template kwargs fill other keys, and request kwargs apply last. A request toggle set to anything but false gets a 400, as do templates that always reason or start every answer with a reasoning block, a rendered prompt that leaves a reasoning block open, and a model whose reasoning parser expects answers to start inside a block the prompt does not close. Setups the route cannot serve faithfully also get a 400: --enable-mis, --dllm-algorithm, built-in named chat templates, Python chat encoders, tokenizers whose chat text does not encode back to the same ids or that split an answer label, and LoRA adapters. So do malformed requests. Questions are encoded one at a time, and other requests run between them. Validating decisions also found two crashes and one wrong result, fixed here: - Keep token-ids logprobs tensors when only some requests in a batch ask for them, so move_logprobs_to_cpu no longer stops the scheduler when scoring is batched with plain generation. - Treat absent logprob arrays as empty in convert_logprob_style, so a request whose grammar fails to compile returns its 400 instead of stopping the server. - Give each token of a multi-token chat stream chunk its own top_logprobs row. Docs: a Decision models page under Specialized Models with a card on the Supported models overview, Semantic Decisions sections on the Qwen3.8-27B and Qwen3.5 model pages, and native API entries for /v1/decisions and the scoring fields it replays through.
rwang5203
force-pushed
the
jev-decisions
branch
from
September 24, 2026 20:19
0adcd44 to
2548920
Compare
Collaborator
Author
|
Closing, as this change landed in main via #41208. |
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Semantic decisions ask a model a fixed question about some input and need one answer plus a probability for every option, without parsing generated text. #40826 added the per-item candidate scoring that gives these probabilities, but each client still has to render the chat prompt, find answer labels that are single tokens, and keep thinking templates out of the way, and mistakes there change decisions without any error.
Changes
POST /v1/decisionsanswers typed questions (choice,score, andyes_no) about an input. The server renders each question with the served chat template and thinking off, checks that its answer labels (AtoZ,0to9, oryesandno) are single tokens at the answer position, and scores all questions in onescore_promptscall without generating. Each answer returns the probability of every option andlabel_mass, pluschoicefor a choice question orscorefor a score question.return_prompt_token_idslets any answer be replayed through/v1/score.--enable-mis,--dllm-algorithm, a LoRA adapter named inmodel, or labels that are not single tokens, return an actionable 400.top_logprobsinside a multi-token chunk.Validation
On main, Qwen3.8-27B and Qwen3.5-35B-A3B in BF16, one H200 each, with the documented launch command.
/v1/scorereplays of the returned ids, cold cache against cold cache, for 23 tested questions, and all 144 SemIf rows return well formed answers on both models (138 and 130 correct, a sanity check only)./v1/score, and/generatereturn 200 on both models, on both model pages' deploy commands, and on the 27B under NEXTN. Main stops the scheduler on the same load without decisions.top_logprobsmatch the non-streamed ones for every streamed token, where main repeats the first token's list.Limits
label_massforyes_noreads low because it counts lowercaseyesandnoonly. Thresholds need labeled data from the workload.label_massby up to about 0.14 between cold and prefix-cached requests and across batch compositions, with the chosen option unchanged in our checks.--preferred-sampling-paramsunset and--enable-mixed-chunkoff for decisions, since the first can scale the probabilities and the second can reuse recurrent state never written for the prefix on these hybrid models.prompt_token_idsbefore relying on a new model.CI States
Latest PR Test (Base): ❌ Run #36054102942
Latest PR Test (Extra): ❌ Run #36054102642
Latest PR Test (AMD ROCm 10): ⏳ Run #36054102650