Repository navigation
[Feature] System One: shared tasks, persistent frontend and replica pools - #4734
Merged
Merged
Conversation
An algorithm.decision selector had to name an explicit model_runtime deployment, so the Vela 2.0 decision model that already answers a request's signals could not also choose its model without a second copy of the same weights. A selector that names no deployment now asks the decision model's shared implicit deployment, as a routing.signals.decision question does: the Router resolves it with RouterConfig.DecisionSelectorDeployment, the model runtime serves it even when no signal uses it, and the selection reasoning names the deployment that chose. An explicit deployment keeps working; with decision_model: Vela-1.0 an omitted deployment is a load error that asks for one, in the Router and in vllm-sr config validate. A blank or padded deployment is rejected with the same hint. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
The Router publishes every option of a choice, set or span decision question as its own signal value (decision:<question>:<key>), but a projection score input could name only the question. A set question's bare value is its most probable label, so a score such as a reasoning effort could not weigh one label, like needs:deliberation, on its own. A projection input of type decision now takes <question>:<key> for one option of a choice, set or span question; validation rejects a key the question does not declare and an option of a noul or score question. Signal usage asks the question whenever a used projection reads one of its options, and the DSL validator accepts the same form. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…cleanly Three DSL findings from authoring a recipe that routes on Vela 2.0 decision questions over catalog-backed models; a maintained recipe's DSL must validate without diagnostics and round-trip byte for byte. - The route guard compared references without their labels, so routes on different options of one decision question or classifier (task:agentic and task:facts) were reported as overlapping. A labelled reference now names its label; the same option in two routes still warns. - A route that answers with fast_response, inline or through a template, calls no model, so it no longer warns that it has no MODEL. - The decompiler copied a catalog-backed model's parameter size from its built-in card into every route, which compiled back as operator metadata. Routes now repeat only an operator's own param_size. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
A routing.signals.decision set or span question that names no deployment asks the Router's decision model, but preparation checked the model card of the empty deployment name. The Router failed to start with "provider \"\" is not served by the model runtime" for any such question, so set and span questions only worked with an explicit deployment. Preparation now checks the card of the deployment the question asks: its own, or the decision model's shared deployment. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Routing Preview reported execution_required for every algorithm.decision route, so a routing-only evaluation could not show which model the decision model would choose. The decision selector keeps no state between requests: it asks the decision model one Choice question about the request. Preview now dry-runs it like static, multi_factor and latency_aware and reports the chosen model as selected; a failed or late answer falls back to the first modelRef as at request time. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Live CPU conformance derived hardware requirements only from explicit model_runtime deployment devices. A recipe whose global.model_catalog.system.decision_model is Vela-2.0-4B or Vela-2.0-9B has no explicit deployment, so the planner scheduled it on CPU, where vllm-sr serve refuses a GPU-only decision model. Such a recipe now lists the gpu requirement and stays out of the CPU matrix like any other hardware-bound recipe. Probe checks also accept the decision algorithm's Preview statuses, selected or execution_required, like the other base selectors. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Decision Balance serves vllm-sr/auto over GLM-5.3-Flash, Qwen3.8-Flash-Next and Qwen3.8-27B. Vela-2.0-4B answers five System One questions (task choice, difficulty score, precise-facts noul, needs set, correction noul) in the call that also answers the prompt guard and safety signals; heuristics cover images, tools, tool loops, earlier answers, input length and brief-answer requests. An effort projection over the difficulty score and the deliberation and verification labels sets the reasoning effort. Ten decisions route guard (fast_response), long_context, vision, recovery, agentic and facts to fixed models, frontier through algorithm: decision with the decision model itself, hard and standard through multi_factor, and the rest to Qwen3.8-27B with thinking off. Costs are relative GPU-seconds per token measured on MI325X; quality evidence reuses the catalog's third-party records and labels the operator ratings it adds for an effort without them. The CPU conformance plan excludes it as GPU-bound, like vela-amd. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
… policy Router Replay could be limited per decision or forbidden per recipe, but not by what a request contains. A recipe that keeps replay for its routing evidence therefore also stored the prompts and answers of requests with names, emails or phone numbers. routing.data_policy.replay_personal_data: false keeps the replay record of a request in which one of the recipe's PII signals matched, with its route, model, signals and detected PII types, but without the request or response body, prompt, tool definitions or tool trace. The recipe's PII signals are then evaluated for every request, even when no decision references them, and a PII classification that fails counts as personal data. The policy needs a routing.signals.pii rule; the Router and vllm-sr config validate reject it without one. The DSL round-trips it and the reference config sets it. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…in Playground A Playground reply showed the decision, algorithm and model but not the decision-model answers that chose them, nor how long routing took: the Router sends x-vsr-matched-decision-model, x-vsr-selected-recipe, x-vsr-selected-confidence and x-vsr-routing-latency-ms, but Playground did not collect or label them. It now shows them as System One Answers, Recipe, Decision Confidence and Routing Latency, the latency beside the other timings. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Decision Balance 0.2.0 asks the decision model's PII span head in the same call and sets routing.data_policy.replay_personal_data: false, so Router Replay keeps the routing evidence of a request with personal data but none of its content. The prompt guard threshold rises from 0.9 to 0.95: on public chat traffic, role-play and code requests scored between 0.9 and 0.95 while the attacks scored higher. A probe asserts the PII span on a request that names a person, an email address and a phone number. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…ewrites #4725 replaced nine built-in mom-v1 probe examples, which changes the materialized message text by 22 bytes, but left the Dashboard's receipt test pinned to the old text. The test fails on main; the new byte count and digest are those of the probes #4725 ships. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…h effort decision-balance 0.3.0 adds two lane rules, each from per-effort measurements on public benchmark samples: - code: a code task at medium or high effort goes to Qwen3.8-Flash-Next at medium effort. On LiveCodeBench it solved more problems at medium effort than either Qwen model at extra-high, with less than half the tokens. - hard: a STEM task the decision model rates at least multi-step (difficulty >= 2) runs at extra-high effort even when its effort score is lower. Asking for only the final answer lowers the deliberation answer, not the reasoning a GPQA-style problem needs; medium effort lost 8 to 15 points there. Probes move the two code examples from standard to the new code group and add collision variants for a letter-only chemistry question (hard), routine algebra (standard) and a one-line code fix (fast). The Model Card states the rules, their evidence and that long_output and creativity are reported but not routed on. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
A listener's api_keys decide who may call it, but any key holder could name a provider model and have it passed through, bypassing the recipes. listeners[].models lists the only request models a standalone listener accepts, by exact `model` value; empty or absent accepts every model. The gateway hands the allow-list to the routing core after the API key check, and the request-body phase checks it against the model request decoding parsed, the one routing uses, before any signal, cache or decision runs. Another model gets 403 model_not_allowed in the client's protocol, and /v1/models on the listener lists only the allowed names. Request-graph hops the Router makes itself are not restricted, so an auto model still reaches the provider models its decisions name, and a restricted listener ignores the skip-processing opt-out, which would bypass the check. --gateway extproc rejects a listener with models as unsupported (the Router's capability check and the CLI's Envoy generation), since the Envoy listener does not enforce it yet. The schema, the CLI model, the reference config, the Dashboard type and the docs follow. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…idge `vllm-sr serve` always passed `--add-host=host.docker.internal:host-gateway`. Docker derives `host-gateway` from its default bridge network, so a daemon configured with `"bridge": "none"` rejected every service container with `unable to derive the IP value for host-gateway`. - `VLLM_SR_HOST_GATEWAY_IP` maps `host.docker.internal` to an explicit address on Docker and Podman. - Without an override, Docker daemons that report no default bridge network skip the mapping with a warning that names the override; any other probe result keeps the previous `host-gateway` mapping. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
`vllm-sr config validate` and `vllm-sr serve` rejected a
`routing.signals.pii[]` rule without `threshold` ("Field required"), while
the Router accepts one and lets the rule take every span the PII model
reports. A Vela 2.0 model reports only spans above its size's calibrated
threshold, so a recipe for one size had to hard-code another size's
operating point to pass the CLI.
The CLI schema now makes the rule threshold optional, like the Router, and
the PII signal tutorial says what an omitted threshold means.
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
The selection factory initializes every algorithm at startup, so RouterDC embedded each model's description even when no decision routes with it. A generation prepares the description embedding model only for decisions that use router_dc or hybrid, so every start without one logged a warning per model: `embedding model "mmbert" was not prepared for this generation`. The embedding set now reports that case as ErrModelNotPrepared (same message, still a capability error), and RouterDC stops at the first such error with a debug line. Other embedding failures keep their warning. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…n threshold The CLI now accepts a PII rule without a threshold, so the recipe no longer hard-codes 0.05. Vela-2.0-4B reports only spans above its calibrated, length-aware span threshold, which is what the replay data policy should act on. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
decision-balance 0.3.1 moves the boundary between medium and high effort from 0.8 to 0.675. Repeated stratified cross-validation on the per-effort benchmark samples chose it together with the hard-STEM difficulty of 2 and the code lane in most folds; held-out accuracy rose by about 1.3 points over 0.8 for about 4% more GPU time per request. Probes near the new boundary are replaced by ones with clear margins, and the plan-without-tools negative moves from standard to hard. The Model Card lists the serving notes behind the measurements on AMD GPUs with vLLM 0.31: TRITON_ATTN for the 27B, whose default backend's decode cost grows with input length, and the GLM-5.3-Flash indexer fix of vllm-project/vllm#59412. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Xunzhuo
requested review from
AayushSaini101,
FAUST-BENCHOU,
Peterren,
WUKUNTAI-0211,
drivebyer,
ramkrishs,
shraderdm,
theohsiung and
wilsonwu
as code owners
October 7, 2026 21:18
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…decision-balance Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
1fanwang
added a commit
to 1fanwang/semantic-router
that referenced
this pull request
Oct 9, 2026
Resolves the conflict with vllm-project#4734 in e2e/testcases/signal_routing_helpers.go by sending main's vllm-sr/auto model and dropping the Model field this branch had added, which no signal routing test sets. Signed-off-by: 1fanwang <1fannnw@gmail.com>
5 tasks done
wilsonwu
added a commit
to wilsonwu/semantic-router
that referenced
this pull request
Oct 9, 2026
Since vllm-project#4734 removed routing.data_policy from the schema, the current parser rejects the immutable v0.4 mom-v1 snapshot before the test reaches any adaptations check. Round-trip only the maintained configs: config/config.yaml, the agent recipe and the latest built-in mom-v1, which carries the same 16 decision adaptations as the snapshot. Signed-off-by: Wilson Wu <iwilsonwu@gmail.com>
5 tasks done
This was referenced Oct 10, 2026
wilsonwu
added a commit
that referenced
this pull request
Oct 10, 2026
* [Bug] Keep decision adaptations when the Builder deploys The Builder decompiles the config to DSL, compiles it back, and deploys the compiled routing in replace mode. The DSL cannot express decision adaptations, and mergeDSLOwnedNodes replaced routing and recipes with the compiled nodes, so a Deploy without edits removed every decision's adaptations: 4 decisions in the default config, 2 in the agent recipe and all 16 in the built-in mom-v1 recipe. Those decisions then fall back to mode apply, and Router Learning acts on routes that were set to bypass or observe it. Before the DSL-owned sections are replaced or merged, copy each base decision's adaptations into the compiled decision with the same name, matching recipes by name first. A value already in the compiled fragment wins. This is the rule dsl.MergeRoutingIntoBase applies when sr-dsl compiles with --base. Signed-off-by: Wilson Wu <iwilsonwu@gmail.com> * Keep the v0.4 release snapshot out of the deploy round-trip test Since #4734 removed routing.data_policy from the schema, the current parser rejects the immutable v0.4 mom-v1 snapshot before the test reaches any adaptations check. Round-trip only the maintained configs: config/config.yaml, the agent recipe and the latest built-in mom-v1, which carries the same 16 decision adaptations as the snapshot. Signed-off-by: Wilson Wu <iwilsonwu@gmail.com> --------- Signed-off-by: Wilson Wu <iwilsonwu@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
System One and Router share decision tasks, model deployments and a persistent frontend. Users can serve native judgments without Chat backends, or use the same models for routing signals. Model workers can load, scale and recover while the Dashboard remains available.
Usage
servedefaults to Router.--engine/-eenables native serving; the optional positional model selects the default judgment model in either mode. Omitting the model preserves the configuration, or uses Vela 2.0 0.3B when none is configured. CPU, CUDA and ROCm placement share the same lifecycle.Changes
/v1/systemonesupportschoice,score,noul,spanandsetin either mode. Complete-input security checks preserve coverage, offsets and error policies. Managed and attached replicas have bounded admission, load-aware dispatch and graceful replacement.vllm-sr/autois the default entrypoint; explicit entrypoints name isolated recipes. Backend model discovery usesglobal.router.list_backend_modelsand listener grants. Replay defaults live underglobal.services.router_replay, with decision plugin overrides. Only active embedding consumers load their models.Validation
c08d8ae2, the full 12-case model-runtime Kubernetes profile passed, including offline attached services, worker recovery and replica isolation. Production Router, Dashboard and CLI use that same revision; HTTPS/auth, native questions, Chat SSE, 13 browser page loads, 10 UI checks and live task/replica telemetry passed. A separate Engine → Router → Engine lifecycle passed with Dashboard login, preserved configuration and real two-GPU native inference in every stage. All five native question types were also run through the actual Playground UI.Input limits are coverage bounds, not latency guarantees. Very long complete-input inference can exceed deadlines on a small CPU, and an active forward can continue after its caller times out. Readiness checks and runtime metrics do not substitute for model-quality evaluation.
System One automatic model routing, further same-GPU scaling and Decision Parallelism remain follow-up work. Additional task consumers are tracked in #4760.
Closes #4732.