Skip to content

perf(least_load): sub-linear selection via power-of-two-choices - #1787

Closed
slin1237 wants to merge 1 commit into
mainfrom
perf/1692b-least-load-sublinear
Closed

slin1237 wants to merge 1 commit into
mainfrom
perf/1692b-least-load-sublinear

Conversation

@slin1237

@slin1237 slin1237 commented Jun 18, 2026 •

Copy link
Copy Markdown
Member

Part of #1692 (routing hot path is O(total workers) per request).

Problem

LeastLoadPolicy::select_worker was O(healthy) per request on two counts:

  1. It recomputed a fleet-mean nominal_throughput (mean of positive
    total_gen_throughput()) by scanning every healthy worker's cached load on
    each call.
  2. It then scored every healthy worker and took the argmin.

Both scale with fleet size and feed the O(workers) routing hot path #1692 targets.

Change (confined to model_gateway/src/policies/least_load.rs)

1. Cache the fleet-mean throughput off the hot path. The mean now lives in an
RwLock<f64>, recomputed in update_loads (and remove_worker) over the whole
load cache and read O(1) in select_worker. update_loads runs per model
group and extends the shared cache, so the mean is taken after the merge to
keep fleet-wide semantics across groups. The default_throughput fallback (no
positive reports) is unchanged.

2. Power-of-two-choices selection. For pools with <= 2 healthy workers we
still score the whole set (exact argmin); for larger pools we sample two distinct
random healthy workers and pick the lower-scored — the standard near-optimal
load-balancing approximation. The RNG mirrors power_of_two.rs (rand::rng() +
an offset that guarantees a distinct second pick in O(1)). The existing score(...)
fn is reused unchanged.

Preserved: empty/single-worker fast paths, healthy-only filtering,
fleet_has_loads semantics, the in-flight token credit to the chosen worker, and
increment_processed(). The only intended behavioral change is min-over-sample
instead of min-over-all
(power-of-two) plus the cached mean.

Load-balance tradeoff

Power-of-two is an approximation: a given request may not land on the global
argmin. But its expected maximum load is near-optimal (exponentially better than
random, close to full-scan), and the in-flight credit still water-fills load
across dispatches within a poll interval — so steady-state balance is preserved
while per-request cost drops from O(healthy) to O(1). The exact-scan path is kept
for tiny pools where sampling buys nothing.

Caveat

The cached mean is computed over the entire load cache rather than only the
currently-healthy subset (update_loads/remove_worker don't have the live
worker/health list). Stale entries are pruned by remove_worker, so the mean
tracks the reporting fleet; this is a negligible, off-hot-path approximation of
the previous per-call healthy-only mean.

Tests

New (all green): power-of-two picks the lower of the sampled pair (4-worker
sampling-branch distribution + deterministic pairwise rule); healthy-only
preserved; in-flight credit applied to the chosen worker; cached mean updates on
update_loads, recomputes on remove_worker, falls back with no positive rates,
and is used to score a missing-snapshot worker; degenerate sizes (0, 1, 2).
Existing least_load tests stay green.

Verification

  • cargo +nightly fmt --all — clean
  • cargo clippy -p smg --all-targets -- -D warnings — clean
  • cargo test -p smg least_load (lib) — 19 passed, 0 failed
  • cargo test -p smg --lib — 1085 passed, 0 failed, 4 ignored

Summary by CodeRabbit

  • Refactor

    • Improved worker selection algorithm now samples two candidates instead of exhaustively evaluating all, reducing computation overhead.
    • Enhanced throughput value handling with automatic fallback to defaults when invalid inputs are provided.
    • Optimized load balancing through caching mechanisms.
  • Tests

    • Extended test coverage for new selection behavior and edge cases.

Part of #1692. Per-request `select_worker` was O(healthy): it recomputed
the fleet-mean throughput by scanning every healthy worker's cached load,
then scored all healthy workers to take the argmin. Both scale with fleet
size and contribute to the routing hot path's O(workers) cost.

Two changes, both confined to `least_load.rs`:

1. Cache the fleet-mean throughput off the hot path. The mean of positive
   per-worker `total_gen_throughput()` now lives in an `RwLock<f64>`,
   recomputed in `update_loads` (and `remove_worker`) over the whole load
   cache and read O(1) in `select_worker`. `update_loads` runs per model
   group and extends the shared cache, so the mean is taken after the
   merge to preserve fleet-wide semantics. The `default_throughput`
   fallback (no positive reports) is unchanged.

2. Power-of-two-choices selection. For pools with <= 2 healthy workers we
   still score the whole set (exact argmin); for larger pools we sample two
   distinct random healthy workers and pick the lower-scored, the standard
   near-optimal load-balancing approximation. RNG mirrors
   `power_of_two.rs` (`rand::rng()` + offset for a distinct second pick).

Behavior preserved: empty/single-worker fast paths, healthy-only
filtering, `fleet_has_loads` semantics, in-flight token credit to the
chosen worker, and `increment_processed()`. The only intended behavioral
change is min-over-sample instead of min-over-all (power-of-two) plus the
cached mean.

Tradeoff: power-of-two is an approximation — a given request may not land
on the global argmin, but expected max-load is near-optimal and the
in-flight credit still water-fills across dispatches within a poll
interval, so steady-state balance is preserved while per-request cost
drops from O(healthy) to O(1).

Tests: power-of-two picks the lower of the sampled pair; healthy-only
preserved; in-flight credit applied to the chosen worker; cached mean
updates on `update_loads`/`remove_worker`, falls back with no positive
rates, and is used to score a missing-snapshot worker; degenerate sizes
(0,1,2). Existing least_load tests stay green.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237 slin1237 added enhancement New feature or request priority:high High priority labels Jun 18, 2026
@github-actions github-actions Bot added the model-gateway Model gateway crate changes label Jun 18, 2026
@coderabbitai

coderabbitai Bot commented Jun 18, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 972c028e-d352-45ad-be97-1703a5ae1901

📥 Commits

Reviewing files that changed from the base of the PR and between a18bbdd and a6b5318.

📒 Files selected for processing (1)
  • model_gateway/src/policies/least_load.rs

📝 Walkthrough

Walkthrough

LeastLoadPolicy gains a nominal_throughput: RwLock<f64> field caching the fleet-wide mean generation throughput. Worker selection is changed from exhaustive argmin to power-of-two choices (sampling two distinct healthy workers). The cache is recomputed after update_loads and remove_worker. Tests are extended to cover all new paths.

Changes

LeastLoadPolicy: Cached Nominal Throughput and Power-of-Two Selection

Layer / File(s) Summary
Struct field and initialization
model_gateway/src/policies/least_load.rs
Adds nominal_throughput: RwLock<f64> to LeastLoadPolicy, sanitizes default_throughput in with_params (rejecting non-finite/non-positive values), and initializes the cache to the sanitized fallback. Adds rand::Rng import for sampling.
Cache read/recompute helpers
model_gateway/src/policies/least_load.rs
Adds get_nominal_throughput to read the cached value and compute_nominal_throughput to derive it as the mean of positive total_gen_throughput values from the load cache, falling back to default_throughput when none exist.
select_worker: cache usage and power-of-two sampling
model_gateway/src/policies/least_load.rs
Reads nominal_throughput from cache (O(1)) instead of recomputing per call. Replaces full argmin scan with power-of-two choices: samples two distinct healthy workers and picks the lower score; scores all workers only when the healthy pool has two or fewer entries.
update_loads and remove_worker cache recomputation
model_gateway/src/policies/least_load.rs
After merging new loads in update_loads and after deleting a worker in remove_worker, recomputes and stores the fleet nominal throughput into the cache.
Tests
model_gateway/src/policies/least_load.rs
Adds WorkerStatus import; adds tests for empty fleets, probabilistic power-of-two sampling, deterministic two-worker pairwise winner, unhealthy-worker skipping, in-flight token credit, nominal throughput recomputation on update_loads, fallback when no positive rates exist, use of cached value when a snapshot is missing, and recomputation after remove_worker.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related issues

Possibly related PRs

  • lightseekorg/smg#1629: Implements the original LeastLoadPolicy argmin scoring and state handling that this PR replaces with cached nominal throughput and power-of-two sampling.

Suggested labels

tests

Suggested reviewers

  • CatherineSue
  • key4ng

Poem

🐇 Hop hop, no more scanning the whole fleet in line,
Two workers sampled — the lighter one's mine!
A cache holds the throughput, no math on each call,
The bunny picks fast, barely slowing at all.
Power of two choices, O(1) delight —
The least-loaded worker selected just right! 🥕

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically summarizes the main optimization: reducing worker selection from linear to sub-linear complexity via power-of-two-choices algorithm.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/1692b-least-load-sublinear

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request optimizes the LeastLoadPolicy by caching the fleet-nominal throughput off the hot path and implementing a power-of-two-choices selection strategy to reduce worker selection complexity to O(1). The reviewer feedback suggests further optimizing performance by replacing the RwLock<f64> used for nominal_throughput with a lock-free AtomicU64 (storing the float as bits), which eliminates lock contention and synchronization overhead during request routing.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread model_gateway/src/policies/least_load.rs
Comment on lines +140 to +141
// No loads yet -> start at the fallback; recomputed on first poll.
nominal_throughput: RwLock::new(default_throughput),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Initialize the nominal_throughput as a lock-free AtomicU64 using f64::to_bits.

Suggested change
// No loads yet -> start at the fallback; recomputed on first poll.
nominal_throughput: RwLock::new(default_throughput),
// No loads yet -> start at the fallback; recomputed on first poll.
nominal_throughput: std::sync::atomic::AtomicU64::new(default_throughput.to_bits()),

Comment on lines +197 to +201
fn nominal_throughput(&self) -> f64 {
self.nominal_throughput
.read()
.map(|t| *t)
.unwrap_or(self.default_throughput)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Load the cached throughput atomically using Ordering::Relaxed and convert it back to f64 with f64::from_bits. This is lock-free and extremely fast.

    fn nominal_throughput(&self) -> f64 {
        f64::from_bits(self.nominal_throughput.load(std::sync::atomic::Ordering::Relaxed))
    }

Comment on lines +313 to +316
let nominal = self.compute_nominal_throughput(&cached);
if let Ok(mut tp) = self.nominal_throughput.write() {
*tp = nominal;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Store the newly computed nominal throughput atomically using Ordering::Relaxed and f64::to_bits.

            let nominal = self.compute_nominal_throughput(&cached);
            self.nominal_throughput.store(nominal.to_bits(), std::sync::atomic::Ordering::Relaxed);

Comment on lines +331 to +334
let nominal = self.compute_nominal_throughput(&cached);
if let Ok(mut tp) = self.nominal_throughput.write() {
*tp = nominal;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Store the newly computed nominal throughput atomically using Ordering::Relaxed and f64::to_bits.

            let nominal = self.compute_nominal_throughput(&cached);
            self.nominal_throughput.store(nominal.to_bits(), std::sync::atomic::Ordering::Relaxed);

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a6b5318ba6

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +208 to +212
let (tp_sum, tp_count) = loads
.values()
.map(|l| l.total_gen_throughput())
.filter(|t| *t > 0.0)
.fold((0.0, 0u32), |(s, n), t| (s + t, n + 1));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exclude stale worker loads from nominal throughput

When a worker transitions away from Ready but is not removed, WorkerMonitor only evicts it from the watch/DP caches on the status-change path, and remove_worker_from_load_aware is only called by the removal workflow. This new cached nominal mean is computed from every entry left in cached_loads, so a NotReady worker's last throughput continues to affect scoring for healthy workers that are missing a fresh snapshot; the previous per-select calculation filtered through the current healthy worker slice, so those stale entries did not skew the fallback drain-time estimate.

Useful? React with 👍 / 👎.

@claude

claude Bot commented Jun 18, 2026

Copy link
Copy Markdown

👋 The PR description doesn't fully follow
PULL_REQUEST_TEMPLATE.md:

  • Missing header: ## Description
  • Missing header: ### Problem (found ## Problem — wrong heading level)
  • Missing header: ### Solution
  • Missing header: ## Changes (found ## Change — singular)
  • Missing header: ## Test Plan (found ## Tests — different name)

Please update the PR description so reviewers have the context they need.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, well-tested change. Power-of-two sampling matches the existing power_of_two.rs pattern, lock ordering is consistent (no deadlock risk), cached mean lifecycle is correct, and the test suite covers the key scenarios thoroughly. No issues found.

@github-actions

github-actions Bot commented Jul 3, 2026

Copy link
Copy Markdown

This pull request has been automatically marked as stale because it has not had any activity within 14 days. It will be automatically closed if no further activity occurs within 16 days. Leave a comment if you feel this pull request should remain open. Thank you!

@github-actions github-actions Bot added the stale PR has been inactive for 14+ days label Jul 3, 2026
@github-actions

github-actions Bot commented Jul 6, 2026

Copy link
Copy Markdown

This pull request has been automatically closed due to inactivity. Please feel free to reopen if you intend to continue working on it. Thank you!

@github-actions github-actions Bot closed this Jul 6, 2026
@lightseek-bot
lightseek-bot deleted the perf/1692b-least-load-sublinear branch July 6, 2026 20:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request model-gateway Model gateway crate changes priority:high High priority stale PR has been inactive for 14+ days

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant