Skip to content

feat(policies): add least_load load-balancing policy - #1629

Merged
slin1237 merged 1 commit into
mainfrom
feat/least-load-lb
Jun 10, 2026
Merged

slin1237 merged 1 commit into
mainfrom
feat/least-load-lb

Conversation

@slin1237

@slin1237 slin1237 commented Jun 10, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

SMG has no load-aware routing policy that accounts for KV-cache pressure. power_of_two compares two random workers on a single metric, and the count-based policies ignore how full each worker's KV cache is — so traffic can pile onto a worker that is about to hit the KV preemption/recompute cliff, spiking TTFT.

Solution

Add least_load: a policy that scores every healthy worker

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated — rustdoc on the policy and config variant
  • (Optional) Please join us on Slack #sig-smg

Summary by CodeRabbit

  • New Features
    • Added a "least_load" routing policy that prefers workers with the lowest combined score (in-flight + cache pressure). Exposed in CLI and Python bindings and selectable alongside existing policies.
  • Documentation
    • Router option docs updated to include the new policy.
  • Tests
    • Unit and integration tests added/updated for parsing, validation, and selection behavior.
  • Chores
    • Worker removal now clears per-worker load caches for load-aware policies.

@slin1237
slin1237 requested a review from CatherineSue as a code owner June 10, 2026 16:21
@github-actions github-actions Bot added the model-gateway Model gateway crate changes label Jun 10, 2026
@coderabbitai

coderabbitai Bot commented Jun 10, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 7d68a879-a1f7-446f-8f60-58fee89b1c6c

📥 Commits

Reviewing files that changed from the base of the PR and between fd663d0 and 4e61087.

📒 Files selected for processing (16)
  • bindings/python/src/lib.rs
  • bindings/python/src/smg/router.py
  • bindings/python/src/smg/router_args.py
  • bindings/python/tests/test_router_config.py
  • bindings/python/tests/test_startup_sequence.py
  • bindings/python/tests/test_validation.py
  • model_gateway/src/config/types.rs
  • model_gateway/src/config/validation.rs
  • model_gateway/src/main.rs
  • model_gateway/src/policies/factory.rs
  • model_gateway/src/policies/least_load.rs
  • model_gateway/src/policies/mod.rs
  • model_gateway/src/policies/power_of_two.rs
  • model_gateway/src/policies/registry.rs
  • model_gateway/src/worker/monitor.rs
  • model_gateway/src/workflow/steps/local/remove_from_policy_registry.rs

📝 Walkthrough

Walkthrough

Adds a new LeastLoad routing policy and integrates it across config, CLI/Python bindings, the policy factory, policy registry, worker monitor load polling, and worker-removal cache cleanup; unit tests verify selection and cache behavior.

Changes

Least-Load Balancing Policy

Layer / File(s) Summary
Configuration contract and validation
model_gateway/src/config/types.rs, model_gateway/src/config/validation.rs
PolicyConfig gains a LeastLoad variant with load_check_interval_secs and lambda fields (serde tag "least_load"), default helpers are provided, PolicyConfig::name() recognizes it, and validation enforces positive interval and finite non-negative lambda.
CLI arguments, Python bindings, and tests
model_gateway/src/main.rs, bindings/python/src/lib.rs, bindings/python/src/smg/router.py, bindings/python/src/smg/router_args.py, bindings/python/tests/*
CLI --policy, --prefill-policy, and --decode-policy accept least_load; CliArgs::parse_policy can construct a PolicyConfig::LeastLoad. Python exposes PolicyType::LeastLoad, maps string "least_load" to it, converts it into the config, updates CLI choices, and extends related tests.
Policy factory wiring
model_gateway/src/policies/factory.rs
Factory constructs LeastLoadPolicy from config (with_lambda) and recognizes "least_load"/"leastload" in dynamic name-based creation; imports updated.
LeastLoadPolicy worker selection and load caching
model_gateway/src/policies/least_load.rs
Adds LeastLoadPolicy with an RwLock-backed per-worker WorkerLoadResponse cache, scoring function in_flight + lambda * k/(1-k) with clamped k and lambda sanitization, argmin selection among healthy workers, selection logging, increment_processed(), update_loads/remove_worker, and unit tests validating selection and cache pruning.
Policy module export and load-aware registry
model_gateway/src/policies/mod.rs, model_gateway/src/policies/registry.rs
Adds least_load submodule and re-exports LeastLoadPolicy; LoadBalancingPolicy gains remove_worker; registry generalizes to get_all_load_aware_policies returning both power_of_two and least_load policy instances (deduplicated by Arc pointer equality) and adds remove_worker_from_load_aware.
Worker monitor load-fetch gating and policy updates
model_gateway/src/worker/monitor.rs
group_monitor_loop now computes load_aware_policies and only fetches/updates loads when that set is non-empty (or a DP-rank policy exists); fetched loads are applied to all load-aware policies before updating DP cache.
Worker removal cleanup
model_gateway/src/policies/power_of_two.rs, model_gateway/src/workflow/steps/local/remove_from_policy_registry.rs
PowerOfTwoPolicy::remove_worker removes cached load entries; removal step now calls remove_worker_from_load_aware(worker_url) to clear cached per-worker reports from load-aware policies during worker removal.

Sequence Diagram

sequenceDiagram
  participant WorkerMonitor
  participant LeastLoadPolicy
  participant Worker
  WorkerMonitor->>LeastLoadPolicy: update_loads(group reports)
  LeastLoadPolicy->>LeastLoadPolicy: compute score(worker) across healthy workers
  LeastLoadPolicy->>Worker: increment_processed() on selected worker
Loading

Estimated Code Review Effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly Related PRs

Suggested Labels

python-bindings

Suggested Reviewers

  • CatherineSue
  • key4ng
  • gongwei-130

Poem

🐰 I hopped through code with careful paws,

A least-load trail with gentle laws,
Lambdas tuned and caches trimmed,
Workers steered where loads are dimmed,
I nibble tests — the build applauds.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 70.59% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main change: adding a new least_load load-balancing policy, which is the primary focus of the changeset across all modified files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/least-load-lb

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new least_load load-balancing policy that routes requests based on a combination of active in-flight requests and KV-cache pressure. The changes integrate this policy into the configuration, validation, CLI, factory, registry, and worker monitor. The review feedback identifies several important improvements: robustly handling potential NaN values in token usage calculation to avoid routing lockups, using parking_lot::RwLock instead of std::sync::RwLock for codebase consistency and to avoid lock poisoning boilerplate, ensuring that processed request metrics are incremented in the single-worker fast path, and optimizing lock acquisition patterns.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread model_gateway/src/policies/least_load.rs
Comment thread model_gateway/src/policies/least_load.rs
Comment thread model_gateway/src/policies/least_load.rs
Comment thread model_gateway/src/policies/least_load.rs
Comment thread model_gateway/src/policies/least_load.rs

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 037ad40a15

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/policies/least_load.rs
Comment thread model_gateway/src/config/types.rs
Comment thread model_gateway/src/policies/least_load.rs
Comment thread model_gateway/src/policies/least_load.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/config/validation.rs`:
- Around line 326-345: The branch added validation for PolicyConfig::LeastLoad
but lacks direct regression tests for the edge cases; add unit tests that
exercise validation.rs's PolicyConfig::LeastLoad path to assert
ConfigError::InvalidValue is returned when load_check_interval_secs == 0 and
when lambda is non-finite (NaN/INFINITY) or negative. Create tests that
construct a PolicyConfig::LeastLoad with load_check_interval_secs = 0 and with
lambda = f64::NAN, f64::INFINITY, and lambda = -0.1, call the same validation
function used by the module (the validator that emits
ConfigError::InvalidValue), and assert the returned error contains field
"load_check_interval_secs" or "lambda" and the expected reason strings. Ensure
tests live alongside other config validation tests and use explicit assertions
(not panics) so these regression cases are covered in CI.

In `@model_gateway/src/main.rs`:
- Around line 931-934: The CLI branch that maps "least_load" to a PolicyConfig
currently hardcodes load_check_interval_secs: 5 which conflicts with
PolicyConfig::LeastLoad's default of 10 in config/types.rs; update parse_policy
(the code that returns PolicyConfig::LeastLoad for "least_load") to use the same
default (10) or to call Default::default() for the LeastLoad variant so CLI and
serde-config share the same defaults.

In `@model_gateway/src/policies/factory.rs`:
- Around line 22-24: Add unit tests that assert the factory returns a
LeastLoadPolicy when given a PolicyConfig::LeastLoad and when requesting the
policy by name; specifically, in the factory tests call create_from_config with
a PolicyConfig::LeastLoad (including a sample lambda) and assert the returned
Arc inner type corresponds to LeastLoadPolicy created via
LeastLoadPolicy::with_lambda, and similarly call create_by_name for the
least_load variant and assert the same; ensure the tests check the lambda value
is preserved (or that the concrete type is LeastLoadPolicy) to lock down the new
dispatch paths in create_from_config and create_by_name.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 2388af81-069a-425a-801a-f1e58177d91f

📥 Commits

Reviewing files that changed from the base of the PR and between 8e9e892 and 037ad40.

📒 Files selected for processing (8)
  • model_gateway/src/config/types.rs
  • model_gateway/src/config/validation.rs
  • model_gateway/src/main.rs
  • model_gateway/src/policies/factory.rs
  • model_gateway/src/policies/least_load.rs
  • model_gateway/src/policies/mod.rs
  • model_gateway/src/policies/registry.rs
  • model_gateway/src/worker/monitor.rs

Comment thread model_gateway/src/config/validation.rs
Comment thread model_gateway/src/main.rs
Comment thread model_gateway/src/policies/factory.rs
@slin1237
slin1237 force-pushed the feat/least-load-lb branch from 037ad40 to c98a059 Compare June 10, 2026 16:43
@github-actions github-actions Bot added python-bindings Python bindings changes tests Test changes labels Jun 10, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c98a059d9d

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/policies/least_load.rs
Comment thread bindings/python/src/lib.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/policies/least_load.rs`:
- Around line 129-133: The update_loads method in LeastLoadPolicy currently
extends cached_loads and never removes stale entries; modify the policy to prune
removed workers by either adding a remove_worker method to the
LoadBalancingPolicy trait and implement it in LeastLoadPolicy to delete entries
from cached_loads, or change update_loads (LeastLoadPolicy::update_loads) to
accept the full current worker set and replace the cache (write-locked
cached_loads = new_map) instead of extend so stale keys are dropped; ensure the
chosen approach integrates with PolicyRegistry::remove_worker_from_cache_aware
and WorkerMonitor eviction so LeastLoadPolicy entries are cleaned when workers
are removed.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: fa8c6695-a434-48ce-8567-cedd17a548c1

📥 Commits

Reviewing files that changed from the base of the PR and between 037ad40 and c98a059.

📒 Files selected for processing (14)
  • bindings/python/src/lib.rs
  • bindings/python/src/smg/router.py
  • bindings/python/src/smg/router_args.py
  • bindings/python/tests/test_router_config.py
  • bindings/python/tests/test_startup_sequence.py
  • bindings/python/tests/test_validation.py
  • model_gateway/src/config/types.rs
  • model_gateway/src/config/validation.rs
  • model_gateway/src/main.rs
  • model_gateway/src/policies/factory.rs
  • model_gateway/src/policies/least_load.rs
  • model_gateway/src/policies/mod.rs
  • model_gateway/src/policies/registry.rs
  • model_gateway/src/worker/monitor.rs

Comment thread model_gateway/src/policies/least_load.rs
@slin1237
slin1237 force-pushed the feat/least-load-lb branch from c98a059 to fd663d0 Compare June 10, 2026 17:20

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fd663d09ce

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/main.rs
// ==================== Routing Policy ====================
/// Load balancing policy to use
#[arg(long, default_value = "cache_aware", value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "prefix_hash", "consistent_hashing", "manual", "bucket"], help_heading = "Routing Policy")]
#[arg(long, default_value = "cache_aware", value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual", "bucket"], help_heading = "Routing Policy")]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Add least_load to Helm values schema

When this policy is selected via the Helm chart, chart values validation still rejects it because deploy/helm/smg/values.schema.json line 10 only enumerates cache_aware, round_robin, power_of_two, manual, random, and prefix_hash (I checked the Helm chart policy schema and values comments). This means Kubernetes/Helm users cannot enable the new least_load policy even though the gateway CLI and Python args now accept it; update the chart schema (and related values comment) alongside this new accepted policy.

Useful? React with 👍 / 👎.

Comment thread model_gateway/src/config/types.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

♻️ Duplicate comments (1)
model_gateway/src/policies/least_load.rs (1)

129-133: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

update_loads keeps scoring from stale load reports.

Lines 129-133 merge new reports into cached_loads, but they never clear keys that disappear from a poll. After a monitor miss, that worker keeps its old KV utilization indefinitely, so routing no longer matches the documented “missing report => k = 0” behavior.

Suggested fix
     fn update_loads(&self, loads: &HashMap<String, WorkerLoadResponse>) {
         if let Ok(mut cached) = self.cached_loads.write() {
-            cached.extend(loads.iter().map(|(k, v)| (k.clone(), v.clone())));
+            cached.clone_from(loads);
         }
     }

This follows the PR contract that workers without a current load report should be scored as if k = 0.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/policies/least_load.rs` around lines 129 - 133, The
update_loads function currently merges new reports into cached_loads but never
removes keys for workers that disappeared, causing stale scores; modify
update_loads (and the cached_loads write) to replace the cache contents with the
new loads set (or explicitly remove keys not present in the incoming loads) so
that workers missing from a poll are treated as absent (k = 0); locate the
update_loads method and change the write logic to clear or overwrite
cached_loads based on the keys of the provided loads HashMap instead of using
extend.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/config/types.rs`:
- Around line 376-387: Add unit tests that assert the PolicyConfig::LeastLoad
variant reports the correct name via its name() method, round-trips through
serde (serialize -> deserialize equals original), and that omitted fields
default to default_least_load_interval and default_least_load_lambda. Create
tests that construct LeastLoad with defaults omitted and with explicit values,
verify deserialized structs equal expected values, and assert the
load_check_interval_secs and lambda match the default_* functions when not
provided; reference symbols: PolicyConfig::LeastLoad, name(),
default_least_load_interval, default_least_load_lambda, and serde
(serde_json::to_string/from_str).

In `@model_gateway/src/main.rs`:
- Line 152: The CLI accepts "bucket" in the value_parser but parse_policy lacks
a "bucket" match arm so inputs silently fall back to "round_robin"; update the
parse_policy function to explicitly handle "bucket" (either map it to the
correct RoutingPolicy variant or return a parse error to fail closed), ensuring
the match in parse_policy covers the "bucket" string and returns a Result/Err
instead of defaulting to the fallback; reference the value_parser and
parse_policy identifiers and add the "bucket" branch (or explicit Err) to keep
CLI parsing consistent and visible to users.

In `@model_gateway/src/policies/least_load.rs`:
- Around line 98-100: The early return for the single-worker fast path skips the
shared post-selection bookkeeping, so change the block in least_load.rs that
currently does `if healthy.len() == 1 { return Some(healthy[0]); }` to instead
perform the same post-selection steps as the multi-worker path: call
`increment_processed()` for the selected worker (healthy[0]) and run any
remaining shared bookkeeping before returning the selection; ensure you
reference and invoke the same helper(s) used by the multi-worker path so the
single-worker path updates the processed counter identically.

---

Duplicate comments:
In `@model_gateway/src/policies/least_load.rs`:
- Around line 129-133: The update_loads function currently merges new reports
into cached_loads but never removes keys for workers that disappeared, causing
stale scores; modify update_loads (and the cached_loads write) to replace the
cache contents with the new loads set (or explicitly remove keys not present in
the incoming loads) so that workers missing from a poll are treated as absent (k
= 0); locate the update_loads method and change the write logic to clear or
overwrite cached_loads based on the keys of the provided loads HashMap instead
of using extend.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: e7ea00da-ab68-4a1a-917c-7187b0bc4ef7

📥 Commits

Reviewing files that changed from the base of the PR and between c98a059 and fd663d0.

📒 Files selected for processing (16)
  • bindings/python/src/lib.rs
  • bindings/python/src/smg/router.py
  • bindings/python/src/smg/router_args.py
  • bindings/python/tests/test_router_config.py
  • bindings/python/tests/test_startup_sequence.py
  • bindings/python/tests/test_validation.py
  • model_gateway/src/config/types.rs
  • model_gateway/src/config/validation.rs
  • model_gateway/src/main.rs
  • model_gateway/src/policies/factory.rs
  • model_gateway/src/policies/least_load.rs
  • model_gateway/src/policies/mod.rs
  • model_gateway/src/policies/power_of_two.rs
  • model_gateway/src/policies/registry.rs
  • model_gateway/src/worker/monitor.rs
  • model_gateway/src/workflow/steps/local/remove_from_policy_registry.rs

Comment thread model_gateway/src/config/types.rs
Comment thread model_gateway/src/main.rs
// ==================== Routing Policy ====================
/// Load balancing policy to use
#[arg(long, default_value = "cache_aware", value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "prefix_hash", "consistent_hashing", "manual", "bucket"], help_heading = "Routing Policy")]
#[arg(long, default_value = "cache_aware", value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual", "bucket"], help_heading = "Routing Policy")]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

bucket is accepted by CLI but silently parsed as round_robin.

value_parser allows "bucket" (Lines 152/217/221), but parse_policy has no "bucket" match arm, so Line 949 fallback applies. This causes incorrect routing policy selection without user-visible error.

Suggested fix (fail closed until bucket parsing is explicit)
-    #[arg(long, default_value = "cache_aware", value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual", "bucket"], help_heading = "Routing Policy")]
+    #[arg(long, default_value = "cache_aware", value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual"], help_heading = "Routing Policy")]
@@
-    #[arg(long, value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual", "bucket"], help_heading = "PD Disaggregation")]
+    #[arg(long, value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual"], help_heading = "PD Disaggregation")]
@@
-    #[arg(long, value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual", "bucket"], help_heading = "PD Disaggregation")]
+    #[arg(long, value_parser = ["random", "round_robin", "cache_aware", "power_of_two", "least_load", "prefix_hash", "consistent_hashing", "manual"], help_heading = "PD Disaggregation")]

Also applies to: 217-222, 916-950

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/main.rs` at line 152, The CLI accepts "bucket" in the
value_parser but parse_policy lacks a "bucket" match arm so inputs silently fall
back to "round_robin"; update the parse_policy function to explicitly handle
"bucket" (either map it to the correct RoutingPolicy variant or return a parse
error to fail closed), ensuring the match in parse_policy covers the "bucket"
string and returns a Result/Err instead of defaulting to the fallback; reference
the value_parser and parse_policy identifiers and add the "bucket" branch (or
explicit Err) to keep CLI parsing consistent and visible to users.

Comment thread model_gateway/src/policies/least_load.rs
@slin1237
slin1237 force-pushed the feat/least-load-lb branch from fd663d0 to 859a9c5 Compare June 10, 2026 18:42
Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237
slin1237 force-pushed the feat/least-load-lb branch from 859a9c5 to 4e61087 Compare June 10, 2026 18:43
@slin1237
slin1237 merged commit c768ba3 into main Jun 10, 2026
7 of 9 checks passed
@slin1237
slin1237 deleted the feat/least-load-lb branch June 10, 2026 18:43

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4e6108715e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +47 to +50
let k = loads
.and_then(|m| m.get(worker.url()))
.map(|l| l.effective_token_usage().clamp(0.0, 0.999))
.unwrap_or(0.0);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Ignore non-finite load samples before scoring

In gRPC mode the backend's token_usage is a proto double, so a worker can report NaN; f64::clamp leaves NaN unchanged. If the first healthy worker's cached sample is NaN, best_score becomes NaN and every later s < best_score comparison is false, causing least_load to keep selecting that first worker until a valid sample arrives. Treat non-finite usage as missing or clamp it only after checking is_finite().

Useful? React with 👍 / 👎.

@claude

claude Bot commented Jun 10, 2026

Copy link
Copy Markdown

👋 The PR description doesn't fully follow
PULL_REQUEST_TEMPLATE.md:

  • Missing header: ## Changes
  • Missing header: ## Test Plan

Please update the PR description so reviewers have the context they need.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes python-bindings Python bindings changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant