Skip to content

perf(cache-aware): gather worker routing state in one pass - #1758

Merged
slin1237 merged 1 commit into
mainfrom
perf/cache-aware-routing-state
Jun 16, 2026
Merged

slin1237 merged 1 commit into
mainfrom
perf/cache-aware-routing-state

Conversation

@slin1237

@slin1237 slin1237 commented Jun 16, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

cache_aware worker selection made several O(workers) passes per request — the healthy filter (get_healthy_worker_indices), the is_imbalanced request-count fold, and the cache-hit tenant→worker / cache-miss min-load scans. Each per-worker access went through an ArcSwap guard: is_healthy()→status() and load() both load the worker's runtime guard, and circuit_breaker_can_execute() loads a second (circuit-breaker) guard. So a healthy worker was read ~3 times per request via separate guarded virtual calls.

In a no-GPU benchmark (4 reactor threads, saturating shared-prefix load, mock fleet) this per-worker guard traffic is what makes cache_aware's per-request CPU grow with worker count. gdb sampling at 2048 workers attributed the cost to arc_swap guard load/drop + per-worker dynamic dispatch in the selection loops — not the radix-tree insert, which is cheap and worker-count-independent.

Solution

Gather everything the balanced and imbalanced paths need in a single pass over the workers, reading each worker once via a new Worker::routing_state() that shares the runtime guard for status + load + processed_requests (the circuit breaker is a separate ArcSwap, so it keeps its own guard). The pass collects the healthy indices, the load min/max for the imbalance count-spread, and the min-load index — with the (load, processed_requests, idx) tie-break from #1714, free here since processed rides the same guard. is_imbalanced and the min-load fallback consume the gathered values instead of re-scanning; a cache hit resolves its tenant→worker with a hash-free linear scan over the gathered healthy indices (url() is a cheap field read — no per-request map).

Net: ~3 guarded virtual calls/worker across ~3 passes → 2 guards + 1 virtual call/worker in a single pass. Worker selection is unchanged except that a cache hit no longer routes onto a Ready-but-circuit-broken worker — it now falls through to min-load, consistent with the rest of selection.

Changes

  • worker/worker.rs: add RoutingState { healthy, can_execute, load, processed } and Worker::routing_state(). The default impl composes the existing accessors; BasicWorker overrides it to read status/load/processed under one runtime guard.
  • policies/cache_aware.rs: select_worker performs the single-pass gather; is_imbalanced takes the precomputed load bounds; the token-tree, string-tree, event-driven and min-load paths consume the gathered min-load index; cache-hit tenant→worker resolution is a hash-free scan over the gathered healthy indices.

Test Plan

No-GPU simulation: 4 reactor threads (saturating), mock fleet (--engine realistic), 16 shared prefixes at 0.8 fraction, 8k offered req/s, cache_aware with the imbalance valve held open so the balanced (tree-route) path is exercised. Metric is gateway cpu_ms/req (CPU-time per completed request; lower is better), reported as the marginal over round_robin at the same worker count so tokenization/proxy overhead cancels and the policy cost is isolated. Single-run, so figures carry measurement noise (~±0.07 at these windows); the marginal is the robust signal.

Routing marginal over round_robin, HTTP path (string tree — not masked by tokenization, so routing cost is visible). The shipped (hash-free scan) and the earlier (per-request map) revisions were A/B'd at 2048 workers back-to-back under identical conditions:

@2048 workers (same conditions) cpu_ms/req marginal vs round_robin
round_robin (baseline) 1.684 —
cache_aware — per-request url→index map (earlier revision) 1.865 +0.181
cache_aware — hash-free scan (shipped) 1.708 +0.024

The shipped version's cache_aware routing is within noise of round_robin at 2048 workers (+0.024 cpu_ms/req), down from +0.181 with the per-request map and ~+0.40 unoptimized — the per-worker ArcSwap guard traffic that scaled with worker count is gone. At 64 workers the marginal is within measurement noise for all variants. gRPC path (token tree) is unchanged — per-request tokenization dominates there either way. (Simulator numbers; absolute magnitudes need a live A/B.)

Gates on this branch (rebased on main):

cargo +nightly fmt --check                          clean
cargo clippy -p smg --all-targets -- -D warnings    clean (rc=0)
cargo test -p smg --lib policies                     107 passed
cargo test -p smg --lib worker::                     207 passed

Full cargo test -p smg --lib shows one unrelated failure, middleware::metrics::tests::distinct_ids_on_matched_route_do_not_grow_interner, which passes in isolation (a shared-interner test sensitive to parallel execution) and is untouched by this change. --all-features is not run locally (it pulls in the opencv-video feature absent in this environment, per CONTRIBUTING); CI runs the full matrix.

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Refactor
    • Optimized load-balancing algorithm to reduce computational overhead during request routing and worker selection.
    • Improved cache-affinity routing logic to better align worker selection with cached content placement.
    • Enhanced worker state management for more consistent and efficient request distribution across available resources.

@slin1237
slin1237 requested a review from CatherineSue as a code owner June 16, 2026 20:09
@coderabbitai

coderabbitai Bot commented Jun 16, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@slin1237, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 3 hours, 16 minutes, and 33 seconds. Learn how PR review limits work.

Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file).

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: dd5bae46-eca8-4682-95ee-ff318e67028c

📥 Commits

Reviewing files that changed from the base of the PR and between 4057120 and c3f0f09.

📒 Files selected for processing (2)
  • model_gateway/src/policies/cache_aware.rs
  • model_gateway/src/worker/worker.rs
📝 Walkthrough

Walkthrough

Introduces a RoutingState snapshot struct on the Worker trait so CacheAwarePolicy::select_worker can gather healthy indices, min/max load, and a tie-broken min-load index in a single O(workers) pass. All downstream routing helpers (select_worker_min_load, select_worker_event_driven, select_worker_with_tokens, select_worker_with_text) are updated to accept and reuse the precomputed min_load_idx rather than recomputing it independently. is_imbalanced is refactored to accept externally precomputed load bounds.

Changes

RoutingState contract and cache-aware single-pass routing

Layer / File(s) Summary
RoutingState struct and Worker trait method
model_gateway/src/worker/worker.rs
Defines pub struct RoutingState { healthy, can_execute, load, processed }; adds a default Worker::routing_state() implementation; BasicWorker overrides it with a single WorkerRuntime lock for status/load/processed and a separate circuit-breaker guard for can_execute.
select_worker single O(workers) pass and is_imbalanced refactor
model_gateway/src/policies/cache_aware.rs
select_worker calls routing_state() on each worker in one loop to build healthy_indices, min_load, max_load, and tie-broken min_load_idx; is_imbalanced now accepts precomputed min_load/max_load parameters; balanced dispatch threads healthy_indices and min_load_idx to tree-selection helpers.
Routing helper signatures updated to accept min_load_idx
model_gateway/src/policies/cache_aware.rs
select_worker_min_load, select_worker_event_driven, select_worker_with_tokens, and select_worker_with_text each gain min_load_idx: Option<usize>; cache-miss and no-overlap fallbacks consume the provided index; cache-hit branches resolve the tenant worker by searching within healthy_indices instead of scanning all workers.
KV imbalance tests updated for new is_imbalanced signature
model_gateway/src/policies/cache_aware.rs
Adds imbalanced(policy, workers) test helper that folds min/max load over healthy workers and forwards them to is_imbalanced; five KV-imbalance test cases are updated to use this helper.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • lightseekorg/smg#1714: Directly overlaps with this PR's cache-aware load-imbalance fix — both modify min_load tie-breaking logic using (load, processed_requests(), idx) in the same fallback selection paths of CacheAwarePolicy.
  • lightseekorg/smg#1621: Both PRs modify is_imbalanced and worker-selection logic in cache_aware.rs; #1621 extends is_imbalanced to add KV token-usage overload triggers at the same code paths this PR refactors.
  • lightseekorg/smg#1128: Makes WorkerRuntime status/load/processed counters publicly readable, which is the direct foundation for the BasicWorker::routing_state() snapshot introduced here.

Suggested reviewers

  • CatherineSue
  • key4ng
  • whybeyoung

Poem

🐇 One loop to rule them all, one pass to bind,
No more re-scanning workers — state's preassigned!
RoutingState leaps forward, healthy and bright,
min_load_idx cached for cache-miss delight.
The rabbit hops faster when data's aligned~ 🌿

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main optimization: consolidating worker state gathering into a single pass instead of multiple passes per request.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch perf/cache-aware-routing-state

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added the model-gateway Model gateway crate changes label Jun 16, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request optimizes worker selection in the cache-aware load balancing policy by introducing a single-pass routing_state() snapshot to retrieve worker status, load, and processed request counts, reducing ArcSwap guard overhead. The reviewer feedback identifies a performance concern regarding the creation of a HashMap (url_to_idx) on every request, which introduces heap allocation and hashing overhead on the hot path. To resolve this, the reviewer suggests replacing the HashMap lookup with an allocation-free linear scan over healthy_indices and using a stack-allocated SmallVec for healthy_indices to eliminate heap allocations entirely.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment thread model_gateway/src/policies/cache_aware.rs
Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment thread model_gateway/src/policies/cache_aware.rs
Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment on lines 1093 to 1099
workers: &[Arc<dyn Worker>],
text: &str,
healthy_indices: &[usize],
url_to_idx: &HashMap<&str, usize>,
min_load_idx: Option<usize>,
model_id: &str,
) -> Option<usize> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Remove the url_to_idx parameter from select_worker_with_text.

        workers: &[Arc<dyn Worker>],
        text: &str,
        healthy_indices: &[usize],
        min_load_idx: Option<usize>,
        model_id: &str,
    ) -> Option<usize> {

Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment thread model_gateway/src/policies/cache_aware.rs Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean perf optimization — the single-pass gather is well-structured, the tie-break logic from #1714 is correctly preserved, and the test helper mirrors the production gather faithfully. One 🟡 nit on a subtle (beneficial) behavioral tightening in the cache-hit tenant resolution path where the CB check is now included.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c1ea4012a4

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/policies/cache_aware.rs Outdated
@slin1237
slin1237 force-pushed the perf/cache-aware-routing-state branch from c1ea401 to 40775d3 Compare June 16, 2026 20:33
@slin1237

Copy link
Copy Markdown
Member Author

Addressed the review feedback (gemini-code-assist, chatgpt-codex, claude): dropped the per-request url→index HashMap — it allocated + hashed all worker URLs every request (incl. miss / imbalanced / event-driven paths that never used it). Cache-hit tenant resolution is now a hash-free linear scan over the already-gathered healthy indices (url() is a cheap field read). Also corrected the comment/PR note: a cache hit now (correctly) excludes circuit-broken workers and falls through to min-load, rather than chasing affinity into a CB-tripped worker. Force-pushed.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 40775d3615

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines 1029 to +1033
let tenant_url: &str = &result.tenant;
workers
.iter()
.position(|w| w.url() == tenant_url)
.filter(|&idx| workers[idx].is_healthy())
} else {
healthy_indices
.iter()
.min_by_key(|&&idx| {
(workers[idx].load(), workers[idx].processed_requests(), idx)
})
.copied()
.find(|&idx| workers[idx].url() == tenant_url)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use min-load when cached tenant is unavailable

When a cache hit points at a tenant that is no longer in healthy_indices (for example, the newly handled case where its circuit breaker is open), this branch leaves selected_idx as None; the post-closure fallback below still returns healthy_indices.first() without incrementing the processed counter, so all such requests are redirected to the first registry entry rather than the precomputed min_load_idx the comment says should be used. The same cache-hit fallback logic appears in the string path, so this should fall back to min_load_idx there as well.

Useful? React with 👍 / 👎.

@slin1237
slin1237 force-pushed the perf/cache-aware-routing-state branch from 40775d3 to 4057120 Compare June 16, 2026 21:09
@slin1237

Copy link
Copy Markdown
Member Author

Re-measured the shipped (scan) version vs the earlier (map) version back-to-back at 2048 workers, same conditions:

@2048w cpu_ms/req marginal vs round_robin
round_robin 1.684 —
cache_aware (per-request map) 1.865 +0.181
cache_aware (hash-free scan, shipped) 1.708 +0.024

Dropping the map cut the routing marginal +0.181 → +0.024 — cache_aware routing is now within noise of round_robin at 2048 workers (was ~+0.40 unoptimized). Test Plan updated.

select_worker made several O(workers) passes per request (the healthy filter, the is_imbalanced load fold, and the cache-hit/miss worker scans), and each per-worker access — status, circuit breaker, load — took its own arc_swap guard. At high worker counts that per-worker guard traffic dominated routing CPU.

Read each worker once via a new Worker::routing_state() that shares the runtime guard for status+load+processed, gathering the healthy set, load min/max and the min-load index in a single pass. is_imbalanced and the min-load fallback consume the gathered bounds; the cache-hit tenant lookup is a hash-free scan over the gathered healthy indices (url() is a cheap field read). The (load, processed_requests, idx) min-load tie-break from #1714 rides the same guard, so it costs nothing extra.

Selection is unchanged except that a cache hit no longer routes onto a Ready-but-circuit-broken worker (it falls through to min-load, like the rest of selection). No-GPU sim (4 threads, 2048 HTTP workers, shared-prefix load), A/B back-to-back: cache_aware routing is now within noise of round_robin at 2048 workers (+0.024 cpu_ms/req marginal, vs +0.18 with a per-request url map and ~+0.40 unoptimized). 26 cache_aware + 107 policy + 207 worker tests pass.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237
slin1237 force-pushed the perf/cache-aware-routing-state branch from 4057120 to c3f0f09 Compare June 16, 2026 21:42
@slin1237
slin1237 merged commit 1a6cbf7 into main Jun 16, 2026
12 of 15 checks passed
@slin1237
slin1237 deleted the perf/cache-aware-routing-state branch June 16, 2026 21:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant