Skip to content

feat(policies): route least_load by token-work expected-wait - #1647

Merged
slin1237 merged 1 commit into
mainfrom
feat/least-load-token-work
Jun 10, 2026
Merged

slin1237 merged 1 commit into
mainfrom
feat/least-load-token-work

Conversation

@slin1237

Copy link
Copy Markdown
Member

Description

Problem

least_load scored workers by in-flight request count plus a KV-pressure
barrier. Under size-skewed traffic that spreads poorly: a 6k-token request and a
200-token request count the same, so a worker already holding a large request
still looks "light" and keeps drawing more large requests until it cliffs.

Solution

Score by token-work expected-wait — estimated time-to-drain (token-work over
throughput) plus the convex KV-pressure barrier — and route to the lowest:

score = (queued_tokens + inflight_tokens) / throughput + kv_pressure_weight * k/(1-k)

Both terms are in seconds. Token-work reflects the load a worker actually
carries; the in-flight term corrects for stale load polls; the M/M/1 barrier
keeps routing off near-full KV.

Changes

  • policies/least_load.rs — expected-wait score; per-worker in-flight
    token-work tracking, reset on each poll; graceful degradation (missing
    queued_tokens → 0; missing throughput → default_throughput; no fresh
    snapshot → drain-time estimate at the fleet-nominal rate; dark fleet → JSQ);
    10 unit tests; standalone algorithm doc + tuning-knob reference.
  • New num_waiting_uncached_tokens load signal — SchedulerLoadSnapshot
    (protocols), sglang_scheduler.proto field 17, the sglang servicer
    populate, and the gRPC From conversions (tokenspeed → 0).
  • Config — kv_pressure_weight, default_throughput, mean_prefill_tokens
    added to PolicyConfig::LeastLoad with serde defaults + validation.
  • Knobs exposed with matching defaults across the smg CLI (clap
    --least-load-* flags) and the Python router (router_args.py flags +
    PyO3 binding).

Test Plan

Pre-PR gate, run locally:

$ cargo +nightly fmt --all -- --check
(silent — clean)

$ cargo clippy -p smg -p smg-python --all-targets --all-features -- -D warnings
Finished — 0 warnings/errors (only the pre-existing amg/smg dual-bin-target note)

$ cargo test -p smg --lib
test result: ok. 1009 passed; 0 failed; 4 ignored

The 10 policies::least_load::tests cover: routing by queued token-work,
throughput normalization, the KV barrier steering off a near-full worker,
in-flight water-filling within a poll interval, update_loads resetting the
in-flight estimate, cold-start JSQ, the missing-snapshot drain-time estimate,
and the zero-throughput fallback.

grpc_servicer + router_args.py pass python3 -m py_compile.

make python-dev (maturin) was not run — maturin isn't available in my
environment — but the binding compiles under cargo clippy -p smg-python. A
maturin rebuild is needed to exercise the new Python flags end-to-end.

Notes

  • vLLM end-to-end depends on the vLLM GetLoads endpoint reporting
    token_usage. Refs feat(grpc): add GetLoads endpoint to the vLLM engine service #1630 — when that merges, its vLLM From<SchedulerLoad>
    needs num_waiting_uncached_tokens: 0 added (additive proto field).
  • ~520 lines / 13 files: one cohesive change (the policy plus the load signal it
    consumes and its config/CLI/binding surface). Happy to split if reviewers
    prefer.
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • Documentation updated (policy doc + tuning-knob reference)

Score each healthy worker by its estimated time-to-drain plus a convex
KV-pressure barrier, and route to the lowest (argmin):

    score = (queued_tokens + inflight_tokens) / throughput
            + kv_pressure_weight * k/(1-k)

- queued_tokens: new num_waiting_uncached_tokens load signal, wired through the
  sglang scheduler proto + servicer and the shared SchedulerLoadSnapshot;
  defaults to 0 for backends that do not report it.
- inflight_tokens: per-worker token-work this router dispatched since the last
  poll (stale-snapshot correction), reset on update_loads.
- throughput: the worker's gen_throughput when reported, else the configurable
  default_throughput (backends without a live generation rate).
- k/(1-k): M/M/1 KV-pressure barrier. Both terms are in seconds, so they add
  directly. Missing signals degrade to in-flight + barrier, then to JSQ.

Token-work, not request count, reflects the load a worker actually carries, so
size-skewed traffic is spread by work rather than by count.

Tuning knobs are exposed with matching defaults via PolicyConfig, the smg CLI,
and the Python router CLI/binding: kv_pressure_weight (0.15s),
default_throughput (2000 tok/s), mean_prefill_tokens (1024).

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Jun 10, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@slin1237, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 19 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 99507926-4a72-4a4c-af3d-9da3927185ce

📥 Commits

Reviewing files that changed from the base of the PR and between 27b528f and 71c29bc.

📒 Files selected for processing (13)
  • bindings/python/src/lib.rs
  • bindings/python/src/smg/router_args.py
  • crates/grpc_client/proto/sglang_scheduler.proto
  • crates/grpc_client/src/sglang_scheduler.rs
  • crates/grpc_client/src/tokenspeed_scheduler.rs
  • crates/protocols/src/worker.rs
  • grpc_servicer/smg_grpc_servicer/sglang/servicer.py
  • model_gateway/src/config/types.rs
  • model_gateway/src/config/validation.rs
  • model_gateway/src/main.rs
  • model_gateway/src/policies/factory.rs
  • model_gateway/src/policies/least_load.rs
  • model_gateway/src/policies/power_of_two.rs
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/least-load-token-work

Comment @coderabbitai help to get the list of available commands and usage tips.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 71c29bc596

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

utilization=result.utilization,
max_running_requests=result.max_running_requests,
# Queued token-work: waiting-queue tokens not served from cache.
num_waiting_uncached_tokens=result.num_waiting_uncached_tokens,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Guard the new SGLang load field

With the released sglang versions still allowed by grpc_servicer/pyproject.toml (sglang>=0.5.10), GetLoadsReqOutput does not expose num_waiting_uncached_tokens; when the load monitor calls GetLoads, this constructor raises AttributeError and the RPC returns no load snapshot, so the new least_load path never receives the token-work data it depends on. Please either use a safe fallback such as getattr(..., 0) or raise the dependency to a release that includes this field.

Useful? React with 👍 / 👎.

Comment on lines +783 to +785
least_load_kv_pressure_weight = 0.15,
least_load_default_throughput = 2000.0,
least_load_mean_prefill_tokens = 1024,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Append the new Python constructor args

These fields are inserted before max_idle_secs even though this struct explicitly says new parameters must be appended to avoid breaking positional _Router(...) callers. Any caller built against the previous signature that passes positional arguments after block_size now has max_idle_secs parsed as least_load_kv_pressure_weight and subsequent arguments shifted, which can produce type errors or silently wrong routing configuration.

Useful? React with 👍 / 👎.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request transitions the least-load routing policy to a least-token-work policy by incorporating queued token-work, in-flight token corrections, and throughput normalization. It updates protobuf schemas, Python bindings, CLI arguments, and the Rust model gateway to support these new parameters. The review feedback highlights several opportunities to improve robustness, specifically by defensively handling non-finite float values (like NaN or Infinity) in throughput and token usage calculations, and using getattr when accessing new fields on external SGLang objects to prevent crashes on older versions.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +157 to +165
let live_throughput = load.total_gen_throughput();
let throughput = if live_throughput > 0.0 {
live_throughput
} else {
self.default_throughput
};
let k = load.effective_token_usage().clamp(0.0, 0.999);
(queued_tokens + inflight_tokens) / throughput
+ self.kv_pressure_weight * k / (1.0 - k)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Defensively check that live_throughput and effective_token_usage are finite numbers. If effective_token_usage() returns NaN (which can happen if the backend reports invalid metrics or under division-by-zero scenarios), calling .clamp(0.0, 0.999) can propagate NaN to the score. This breaks the s < best_score comparison in the selection loop, potentially leading to silent routing failures or extreme load imbalance.

                let live_throughput = load.total_gen_throughput();
                let throughput = if live_throughput.is_finite() && live_throughput > 0.0 {
                    live_throughput
                } else {
                    self.default_throughput
                };
                let usage = load.effective_token_usage();
                let k = if usage.is_finite() {
                    usage.clamp(0.0, 0.999)
                } else {
                    0.0
                };
                (queued_tokens + inflight_tokens) / throughput
                    + self.kv_pressure_weight * k / (1.0 - k)

Comment on lines +205 to +210
let (tp_sum, tp_count) = healthy
.iter()
.filter_map(|&i| loads.and_then(|m| m.get(workers[i].url())))
.map(|l| l.total_gen_throughput())
.filter(|t| *t > 0.0)
.fold((0.0, 0u32), |(s, n), t| (s + t, n + 1));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Defensively filter out non-finite throughput values (such as NaN or Infinity) when calculating the fleet-nominal throughput. If a backend reports an invalid non-finite throughput, it could skew the average or result in an infinite nominal throughput, causing incorrect routing decisions for workers missing a fresh snapshot.

Suggested change
let (tp_sum, tp_count) = healthy
.iter()
.filter_map(|&i| loads.and_then(|m| m.get(workers[i].url())))
.map(|l| l.total_gen_throughput())
.filter(|t| *t > 0.0)
.fold((0.0, 0u32), |(s, n), t| (s + t, n + 1));
let (tp_sum, tp_count) = healthy
.iter()
.filter_map(|&i| loads.and_then(|m| m.get(workers[i].url())))
.map(|l| l.total_gen_throughput())
.filter(|t| t.is_finite() && *t > 0.0)
.fold((0.0, 0u32), |(s, n), t| (s + t, n + 1));

utilization=result.utilization,
max_running_requests=result.max_running_requests,
# Queued token-work: waiting-queue tokens not served from cache.
num_waiting_uncached_tokens=result.num_waiting_uncached_tokens,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Use getattr defensively when accessing num_waiting_uncached_tokens on the result object. Since SGLang is an external dependency, older or different versions of SGLang might not expose this field on GetLoadsReqOutput, which would cause an AttributeError and crash the load reporting path.

Suggested change
num_waiting_uncached_tokens=result.num_waiting_uncached_tokens,
num_waiting_uncached_tokens=getattr(result, "num_waiting_uncached_tokens", 0),

utilization=result.utilization,
max_running_requests=result.max_running_requests,
# Queued token-work: waiting-queue tokens not served from cache.
num_waiting_uncached_tokens=result.num_waiting_uncached_tokens,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: num_waiting_uncached_tokens is a new field on GetLoadsReqOutput (from sglang). If the deployed sglang version doesn't expose it yet, this AttributeError crashes the entire _convert_loads_to_protobuf call, which loses ALL load data for the worker — not just the new signal. The Rust scoring side was carefully designed for graceful degradation (queued_tokens = 0 when absent), but a crash here prevents that design from working.

A defensive getattr would keep existing load monitoring intact while the new signal degrades to 0:

Suggested change
num_waiting_uncached_tokens=result.num_waiting_uncached_tokens,
num_waiting_uncached_tokens=getattr(result, "num_waiting_uncached_tokens", 0),

@@ -402,7 +412,15 @@ fn default_least_load_interval() -> u64 {
}

fn default_least_load_kv_pressure_weight() -> f64 {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: The default changed from 1.5 to 0.15 and the units changed from "request-equivalents" to "seconds." Since kv_pressure_weight was just introduced (PR #1632), any configs or docs that hardcoded the old default of 1.5 would now produce a ~10× stronger barrier in the new formula than intended. Might be worth calling this out in the PR description or a changelog entry so operators know to re-evaluate explicit kv_pressure_weight values.

@github-actions github-actions Bot added python-bindings Python bindings changes grpc gRPC client and router changes protocols Protocols crate changes model-gateway Model gateway crate changes labels Jun 10, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well-designed change. The token-work scoring with in-flight correction is a solid improvement over raw request-count routing for size-skewed traffic. Lock ordering is consistent (no deadlock risk), validation and constructor defaults align across all 5 config surfaces, and the degradation paths (missing snapshot → time-estimate, dark fleet → JSQ, zero throughput → default) are clean.

Two minor nits posted inline — a defensive getattr on the servicer to match the Rust-side graceful degradation, and a migration note for the kv_pressure_weight unit change. Neither is blocking.

0 🔴 Important · 2 🟡 Nit · 0 🟣 Pre-existing

@slin1237
slin1237 merged commit 3de811b into main Jun 10, 2026
49 of 51 checks passed
@slin1237
slin1237 deleted the feat/least-load-token-work branch June 10, 2026 22:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes protocols Protocols crate changes python-bindings Python bindings changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant