Skip to content

[Feature] System One: shared tasks, persistent frontend and replica pools - #4734

Merged
Xunzhuo merged 76 commits into
mainfrom
xunzhuo/pytorchcon-decision-balance
Oct 9, 2026
Merged

Xunzhuo merged 76 commits into
mainfrom
xunzhuo/pytorchcon-decision-balance

Conversation

@Xunzhuo

@Xunzhuo Xunzhuo commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

System One and Router share decision tasks, model deployments and a persistent frontend. Users can serve native judgments without Chat backends, or use the same models for routing signals. Model workers can load, scale and recover while the Dashboard remains available.

Usage

# Route Chat requests using your configuration
vllm-sr serve vllm-sr/Vela-2.0-4B --config config.yaml

# Serve native judgment requests on two GPUs
vllm-sr serve vllm-sr/Vela-2.0-4B --engine --platform rocm -dp 2 --device-ids 0,1

serve defaults to Router. --engine/-e enables native serving; the optional positional model selects the default judgment model in either mode. Omitting the model preserves the configuration, or uses Vela 2.0 0.3B when none is configured. CPU, CUDA and ROCm placement share the same lifecycle.

Changes

  • Tasks and runtime: Decision 1.0/2.0 and Vela use shared task definitions and capability checks. Native /v1/systemone supports choice, score, noul, span and set in either mode. Complete-input security checks preserve coverage, offsets and error policies. Managed and attached replicas have bounded admission, load-aware dispatch and graceful replacement.
  • Dashboard: Build → System One contains Decision Models, Decision Playground and Decision Monitoring. Users can choose models, edit task bindings, deploy replicas, try complete examples and inspect readiness and task/runtime traffic. Catalog pagination, shared observations and independent loading states keep unrelated pages responsive. Mode is determined at startup.
  • Configuration: vllm-sr/auto is the default entrypoint; explicit entrypoints name isolated recipes. Backend model discovery uses global.router.list_backend_models and listener grants. Replay defaults live under global.services.router_replay, with decision plugin overrides. Only active embedding consumers load their models.
  • Compatibility and guidance: Authored complexity examples and reask rules retain their metrics; authored preference examples retain prototype scoring unless explicitly overridden. Unreleased flags and superseded configuration fields are removed, with replacements documented in the English and Chinese guides.

Validation

  • Canonical checks for the changed source passed, including affected Go suites, configuration/schema checks and focused race tests for attached-service recovery and generation retirement. English and Chinese documentation builds passed.
  • Existing live regression includes 631 CPU probes across eight recipes and 12 Kubernetes routing/error test cases, including the original 51-case fallback checks. Their measured revisions are retained in the acceptance records.
  • On c08d8ae2, the full 12-case model-runtime Kubernetes profile passed, including offline attached services, worker recovery and replica isolation. Production Router, Dashboard and CLI use that same revision; HTTPS/auth, native questions, Chat SSE, 13 browser page loads, 10 UI checks and live task/replica telemetry passed. A separate Engine → Router → Engine lifecycle passed with Dashboard login, preserved configuration and real two-GPU native inference in every stage. All five native question types were also run through the actual Playground UI.
  • Final-head GitHub CI is still running; this candidate is not yet fully accepted. The workshop archive will be published after those checks complete.

Input limits are coverage bounds, not latency guarantees. Very long complete-input inference can exceed deadlines on a small CPU, and an active forward can continue after its caller times out. Readiness checks and runtime metrics do not substitute for model-quality evaluation.

System One automatic model routing, further same-GPU scaling and Decision Parallelism remain follow-up work. Additional task consumers are tracked in #4760.

Closes #4732.

Xunzhuo added 18 commits October 8, 2026 05:07
An algorithm.decision selector had to name an explicit model_runtime
deployment, so the Vela 2.0 decision model that already answers a
request's signals could not also choose its model without a second copy
of the same weights.

A selector that names no deployment now asks the decision model's shared
implicit deployment, as a routing.signals.decision question does: the
Router resolves it with RouterConfig.DecisionSelectorDeployment, the model
runtime serves it even when no signal uses it, and the selection reasoning
names the deployment that chose. An explicit deployment keeps working;
with decision_model: Vela-1.0 an omitted deployment is a load error that
asks for one, in the Router and in vllm-sr config validate. A blank or
padded deployment is rejected with the same hint.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
The Router publishes every option of a choice, set or span decision
question as its own signal value (decision:<question>:<key>), but a
projection score input could name only the question. A set question's
bare value is its most probable label, so a score such as a reasoning
effort could not weigh one label, like needs:deliberation, on its own.

A projection input of type decision now takes <question>:<key> for one
option of a choice, set or span question; validation rejects a key the
question does not declare and an option of a noul or score question.
Signal usage asks the question whenever a used projection reads one of
its options, and the DSL validator accepts the same form.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…cleanly

Three DSL findings from authoring a recipe that routes on Vela 2.0
decision questions over catalog-backed models; a maintained recipe's DSL
must validate without diagnostics and round-trip byte for byte.

- The route guard compared references without their labels, so routes on
  different options of one decision question or classifier (task:agentic
  and task:facts) were reported as overlapping. A labelled reference now
  names its label; the same option in two routes still warns.
- A route that answers with fast_response, inline or through a template,
  calls no model, so it no longer warns that it has no MODEL.
- The decompiler copied a catalog-backed model's parameter size from its
  built-in card into every route, which compiled back as operator
  metadata. Routes now repeat only an operator's own param_size.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
A routing.signals.decision set or span question that names no deployment
asks the Router's decision model, but preparation checked the model card
of the empty deployment name. The Router failed to start with "provider
\"\" is not served by the model runtime" for any such question, so set
and span questions only worked with an explicit deployment.

Preparation now checks the card of the deployment the question asks: its
own, or the decision model's shared deployment.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Routing Preview reported execution_required for every algorithm.decision
route, so a routing-only evaluation could not show which model the
decision model would choose. The decision selector keeps no state between
requests: it asks the decision model one Choice question about the
request. Preview now dry-runs it like static, multi_factor and
latency_aware and reports the chosen model as selected; a failed or late
answer falls back to the first modelRef as at request time.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Live CPU conformance derived hardware requirements only from explicit
model_runtime deployment devices. A recipe whose
global.model_catalog.system.decision_model is Vela-2.0-4B or Vela-2.0-9B
has no explicit deployment, so the planner scheduled it on CPU, where
vllm-sr serve refuses a GPU-only decision model. Such a recipe now lists
the gpu requirement and stays out of the CPU matrix like any other
hardware-bound recipe.

Probe checks also accept the decision algorithm's Preview statuses,
selected or execution_required, like the other base selectors.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Decision Balance serves vllm-sr/auto over GLM-5.3-Flash,
Qwen3.8-Flash-Next and Qwen3.8-27B. Vela-2.0-4B answers five System One
questions (task choice, difficulty score, precise-facts noul, needs set,
correction noul) in the call that also answers the prompt guard and
safety signals; heuristics cover images, tools, tool loops, earlier
answers, input length and brief-answer requests.

An effort projection over the difficulty score and the deliberation and
verification labels sets the reasoning effort. Ten decisions route
guard (fast_response), long_context, vision, recovery, agentic and
facts to fixed models, frontier through algorithm: decision with the
decision model itself, hard and standard through multi_factor, and the
rest to Qwen3.8-27B with thinking off. Costs are relative GPU-seconds
per token measured on MI325X; quality evidence reuses the catalog's
third-party records and labels the operator ratings it adds for an
effort without them.

The CPU conformance plan excludes it as GPU-bound, like vela-amd.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
… policy

Router Replay could be limited per decision or forbidden per recipe, but
not by what a request contains. A recipe that keeps replay for its
routing evidence therefore also stored the prompts and answers of
requests with names, emails or phone numbers.

routing.data_policy.replay_personal_data: false keeps the replay record
of a request in which one of the recipe's PII signals matched, with its
route, model, signals and detected PII types, but without the request or
response body, prompt, tool definitions or tool trace. The recipe's PII
signals are then evaluated for every request, even when no decision
references them, and a PII classification that fails counts as personal
data. The policy needs a routing.signals.pii rule; the Router and
vllm-sr config validate reject it without one. The DSL round-trips it and
the reference config sets it.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…in Playground

A Playground reply showed the decision, algorithm and model but not the
decision-model answers that chose them, nor how long routing took: the
Router sends x-vsr-matched-decision-model, x-vsr-selected-recipe,
x-vsr-selected-confidence and x-vsr-routing-latency-ms, but Playground
did not collect or label them. It now shows them as System One Answers,
Recipe, Decision Confidence and Routing Latency, the latency beside the
other timings.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Decision Balance 0.2.0 asks the decision model's PII span head in the
same call and sets routing.data_policy.replay_personal_data: false, so
Router Replay keeps the routing evidence of a request with personal data
but none of its content. The prompt guard threshold rises from 0.9 to
0.95: on public chat traffic, role-play and code requests scored between
0.9 and 0.95 while the attacks scored higher. A probe asserts the PII span
on a request that names a person, an email address and a phone number.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…ewrites

#4725 replaced nine built-in mom-v1 probe examples, which changes the
materialized message text by 22 bytes, but left the Dashboard's receipt
test pinned to the old text. The test fails on main; the new byte count
and digest are those of the probes #4725 ships.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…h effort

decision-balance 0.3.0 adds two lane rules, each from per-effort
measurements on public benchmark samples:

- code: a code task at medium or high effort goes to Qwen3.8-Flash-Next at
  medium effort. On LiveCodeBench it solved more problems at medium effort
  than either Qwen model at extra-high, with less than half the tokens.
- hard: a STEM task the decision model rates at least multi-step
  (difficulty >= 2) runs at extra-high effort even when its effort score is
  lower. Asking for only the final answer lowers the deliberation answer,
  not the reasoning a GPQA-style problem needs; medium effort lost 8 to 15
  points there.

Probes move the two code examples from standard to the new code group and
add collision variants for a letter-only chemistry question (hard), routine
algebra (standard) and a one-line code fix (fast). The Model Card states
the rules, their evidence and that long_output and creativity are reported
but not routed on.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
A listener's api_keys decide who may call it, but any key holder could
name a provider model and have it passed through, bypassing the recipes.
listeners[].models lists the only request models a standalone listener
accepts, by exact `model` value; empty or absent accepts every model.

The gateway hands the allow-list to the routing core after the API key
check, and the request-body phase checks it against the model request
decoding parsed, the one routing uses, before any signal, cache or
decision runs. Another model gets 403 model_not_allowed in the client's
protocol, and /v1/models on the listener lists only the allowed names.
Request-graph hops the Router makes itself are not restricted, so an
auto model still reaches the provider models its decisions name, and a
restricted listener ignores the skip-processing opt-out, which would
bypass the check.

--gateway extproc rejects a listener with models as unsupported (the
Router's capability check and the CLI's Envoy generation), since the
Envoy listener does not enforce it yet. The schema, the CLI model, the
reference config, the Dashboard type and the docs follow.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…idge

`vllm-sr serve` always passed `--add-host=host.docker.internal:host-gateway`.
Docker derives `host-gateway` from its default bridge network, so a daemon
configured with `"bridge": "none"` rejected every service container with
`unable to derive the IP value for host-gateway`.

- `VLLM_SR_HOST_GATEWAY_IP` maps `host.docker.internal` to an explicit
  address on Docker and Podman.
- Without an override, Docker daemons that report no default bridge network
  skip the mapping with a warning that names the override; any other probe
  result keeps the previous `host-gateway` mapping.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
`vllm-sr config validate` and `vllm-sr serve` rejected a
`routing.signals.pii[]` rule without `threshold` ("Field required"), while
the Router accepts one and lets the rule take every span the PII model
reports. A Vela 2.0 model reports only spans above its size's calibrated
threshold, so a recipe for one size had to hard-code another size's
operating point to pass the CLI.

The CLI schema now makes the rule threshold optional, like the Router, and
the PII signal tutorial says what an omitted threshold means.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
The selection factory initializes every algorithm at startup, so RouterDC
embedded each model's description even when no decision routes with it.
A generation prepares the description embedding model only for decisions
that use router_dc or hybrid, so every start without one logged a warning
per model: `embedding model "mmbert" was not prepared for this generation`.

The embedding set now reports that case as ErrModelNotPrepared (same
message, still a capability error), and RouterDC stops at the first such
error with a debug line. Other embedding failures keep their warning.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…n threshold

The CLI now accepts a PII rule without a threshold, so the recipe no longer
hard-codes 0.05. Vela-2.0-4B reports only spans above its calibrated,
length-aware span threshold, which is what the replay data policy should
act on.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
decision-balance 0.3.1 moves the boundary between medium and high effort
from 0.8 to 0.675. Repeated stratified cross-validation on the per-effort
benchmark samples chose it together with the hard-STEM difficulty of 2
and the code lane in most folds; held-out accuracy rose by about 1.3
points over 0.8 for about 4% more GPU time per request.

Probes near the new boundary are replaced by ones with clear margins, and
the plan-without-tools negative moves from standard to hard. The Model
Card lists the serving notes behind the measurements on AMD GPUs with
vLLM 0.31: TRITON_ATTN for the 27B, whose default backend's decode cost
grows with input length, and the GLM-5.3-Flash indexer fix of
vllm-project/vllm#59412.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Copilot AI balanced review requested due to automatic review settings October 7, 2026 21:18

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added the pr/blocked Blocked on a named decision, dependency, or required check. label Oct 7, 2026
@github-actions github-actions Bot added pr/needs-rebase Needs rebase or conflict resolution. and removed pr/needs-review Ready for reviewer attention. labels Oct 8, 2026
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
…decision-balance

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
@github-actions github-actions Bot added pr/blocked Blocked on a named decision, dependency, or required check. and removed pr/needs-rebase Needs rebase or conflict resolution. labels Oct 9, 2026
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
@github-actions github-actions Bot added pr/needs-review Ready for reviewer attention. pr/needs-rebase Needs rebase or conflict resolution. and removed pr/blocked Blocked on a named decision, dependency, or required check. pr/needs-review Ready for reviewer attention. labels Oct 9, 2026
@Xunzhuo
Xunzhuo merged commit 9156d5b into main Oct 9, 2026
95 checks passed
@github-actions github-actions Bot removed the pr/needs-rebase Needs rebase or conflict resolution. label Oct 9, 2026
1fanwang added a commit to 1fanwang/semantic-router that referenced this pull request Oct 9, 2026
Resolves the conflict with vllm-project#4734 in e2e/testcases/signal_routing_helpers.go
by sending main's vllm-sr/auto model and dropping the Model field this
branch had added, which no signal routing test sets.

Signed-off-by: 1fanwang <1fannnw@gmail.com>
wilsonwu added a commit to wilsonwu/semantic-router that referenced this pull request Oct 9, 2026
Since vllm-project#4734 removed routing.data_policy from the schema, the current
parser rejects the immutable v0.4 mom-v1 snapshot before the test
reaches any adaptations check. Round-trip only the maintained configs:
config/config.yaml, the agent recipe and the latest built-in mom-v1,
which carries the same 16 decision adaptations as the snapshot.

Signed-off-by: Wilson Wu <iwilsonwu@gmail.com>
wilsonwu added a commit that referenced this pull request Oct 10, 2026
* [Bug] Keep decision adaptations when the Builder deploys

The Builder decompiles the config to DSL, compiles it back, and
deploys the compiled routing in replace mode. The DSL cannot express
decision adaptations, and mergeDSLOwnedNodes replaced routing and
recipes with the compiled nodes, so a Deploy without edits removed
every decision's adaptations: 4 decisions in the default config, 2
in the agent recipe and all 16 in the built-in mom-v1 recipe. Those
decisions then fall back to mode apply, and Router Learning acts on
routes that were set to bypass or observe it.

Before the DSL-owned sections are replaced or merged, copy each base
decision's adaptations into the compiled decision with the same name,
matching recipes by name first. A value already in the compiled
fragment wins. This is the rule dsl.MergeRoutingIntoBase applies when
sr-dsl compiles with --base.

Signed-off-by: Wilson Wu <iwilsonwu@gmail.com>

* Keep the v0.4 release snapshot out of the deploy round-trip test

Since #4734 removed routing.data_policy from the schema, the current
parser rejects the immutable v0.4 mom-v1 snapshot before the test
reaches any adaptations check. Round-trip only the maintained configs:
config/config.yaml, the agent recipe and the latest built-in mom-v1,
which carries the same 16 decision adaptations as the snapshot.

Signed-off-by: Wilson Wu <iwilsonwu@gmail.com>

---------

Signed-off-by: Wilson Wu <iwilsonwu@gmail.com>
@Xunzhuo
Xunzhuo deleted the xunzhuo/pytorchcon-decision-balance branch October 10, 2026 20:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

wg/mom-routing Owned by the MoM and Routing Workgroup.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] Recipes: route a reasoning fleet with one decision-model call (decision-balance)

3 participants