Skip to content

fix(cache-aware): fix load imbalance when decode is faster than prefill - #1714

Merged
slin1237 merged 4 commits into
smg-project:mainfrom
SYChen123:load-fix
Jun 16, 2026
Merged

slin1237 merged 4 commits into
smg-project:mainfrom
SYChen123:load-fix

Conversation

@SYChen123

@SYChen123 SYChen123 commented Jun 13, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

In some extreme scenarios, default cache_aware policy will lead to severe load imbalance problem.

For example, 32k input and output only 1 token. In such case, all the requests will be routed to the same prefill worker and others are always left unused.

The reason is that in the select_worker of CacheAwarePolicy, only workers[idx].load() is used for sorting. After the worker is selected, worker.load() will only be updated till create_streaming_response is called (after stream response is returned from decode worker to the router). If prefill is very slow and decode is very fast, then at most of the time, worker.load() will be the same for all prefill workers. Then pd-router will always select the first prefill worker.

Solution

Changes

  1. Create WorkerLoadGuard before request is sent to the worker for inference.
  2. In select_worker, sort the workers according to (load, processed_requests, idx), instead of solely depending on load.

Test Plan

  1. Unit test test_streaming_load_tracking Passed
    cd ./model-gateway && TMPDIR=/private/tmp cargo test --lib test_streaming_load_tracking
image
  1. bench_serving
    benching random dataset with 32k input and 1 token output, all the prefill workers have requests received.
image
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

Summary of changes

  • Performance Improvements

    • Enhanced cache-aware worker selection by improving tie-breaking with additional workload metrics for more balanced routing.
  • Bug Fixes

    • Fixed streaming load tracking to consistently apply to both streaming success and streaming decode-error responses.
  • Refactor

    • Streamlined streaming response creation by passing a pre-built load-guard set through the streaming paths rather than constructing it internally.
  • Tests

    • Updated token-tree routing coverage to use the full page-sized token sequence in the empty-indexer scenario.

@github-actions github-actions Bot added the model-gateway Model gateway crate changes label Jun 13, 2026
@coderabbitai

coderabbitai Bot commented Jun 13, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 507d7cae-a2a8-4e02-baf5-75e67db64362

📥 Commits

Reviewing files that changed from the base of the PR and between 669b055 and 3228427.

📒 Files selected for processing (2)
  • model_gateway/src/policies/cache_aware.rs
  • model_gateway/src/routers/http/pd_router.rs

📝 Walkthrough

Walkthrough

Worker selection in the cache-aware policy now breaks ties using processed request counts alongside load. Load guard management in the PD router's streaming response path is refactored to construct guards upfront and pass them explicitly through internal helpers, rather than creating them conditionally inside response constructors.

Changes

Worker Selection and Load Guard Flow

Layer / File(s) Summary
Worker Selection Tie-Breaking
model_gateway/src/policies/cache_aware.rs
The imbalanced-mode fallback (select_worker_min_load) and both gRPC and HTTP low-match-rate fallbacks (select_worker_with_tokens, select_worker_with_text) now use (load, processed_requests, idx) as tie-breaking criteria instead of load alone. Event-driven selection's no-overlap fallback is also updated. A test is updated to send cache-eligible token sequences. This ensures deterministic and queue-sensitive worker selection across all fallback paths.
Load Guard Threading Through Streaming Response
model_gateway/src/routers/http/pd_router.rs
execute_dual_dispatch_internal now constructs load_guards upfront for both prefill and decode workers. This vector is passed explicitly to handle_decode_error_response (which now accepts only decode plus load_guards, dropping the prefill parameter) and create_streaming_response (which now accepts load_guards instead of worker parameters). Response wrapping uses AttachedBody::wrap_response() with the explicitly passed guards instead of creating them internally. The streaming tests are updated to construct guards externally and invoke the refactored signatures.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Suggested labels

tests

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237
  • gongwei-130

Poem

🐰 Through queues both short and long they hop,
With processed counts that never stop,
The guards now flow in orderly lines,
No more conditions, just designs—
Load-balanced hares in perfect time! 🎯

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly addresses the main change: fixing load imbalance by improving tie-breaking in load selection when decode is faster than prefill.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@mergify

mergify Bot commented Jun 13, 2026

Copy link
Copy Markdown
Contributor

Hi @SYChen123, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the worker selection logic in cache_aware.rs to tie-break using processed requests and worker indices, and refactors load management in pd_router.rs by eagerly creating and passing WorkerLoadGuards. The feedback suggests extracting the duplicated worker tie-breaking closure logic into a shared helper function to improve maintainability, and adding a trailing comma in the test file to match standard Rust formatting.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment thread model_gateway/src/policies/cache_aware.rs Outdated
Comment thread model_gateway/src/routers/http/pd_router.rs Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f81b8665ed

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/policies/cache_aware.rs Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/policies/cache_aware.rs`:
- Around line 486-487: The event-driven no-overlap fallback still chooses by
workers[idx].load() only; update that fallback to use the same tie-break key as
the other branches — (workers[idx].load(), workers[idx].processed_requests(),
idx) — so selection matches the min_by_key change, and add a regression test
(either test_event_driven_no_overlap_uses_min_load or
test_event_driven_short_request_uses_min_load) that exercises the gRPC+KV events
cold-start/no-overlap path to assert selection prefers lower processed_requests
on equal load.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 0d211090-1cf0-4d12-b659-e6399f37eab2

📥 Commits

Reviewing files that changed from the base of the PR and between 6deba3f and f81b866.

📒 Files selected for processing (2)
  • model_gateway/src/policies/cache_aware.rs
  • model_gateway/src/routers/http/pd_router.rs

Comment thread model_gateway/src/policies/cache_aware.rs
@slin1237 slin1237 changed the title [SMG] Fix load imbalance issue of cache_aware when decode is faster than prefill fix(cache-aware): fix load imbalance when decode is faster than prefill Jun 16, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1ad3d475b3

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/routers/http/pd_router.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/routers/http/pd_router.rs`:
- Around line 1578-1580: The call to handle_decode_error_response has arguments
in the wrong positions and types. The function expects the parameters in this
order: response, context reference, decode worker, and then load guards as a
Vec. The current code incorrectly passes prefill as the third argument instead
of decode, and passes decode (an Arc<dyn Worker>) as the fourth argument instead
of the required Vec<WorkerLoadGuard>. Fix this by passing decode as the third
argument and creating a Vec containing WorkerLoadGuard instances for both the
prefill and decode workers using WorkerLoadGuard::new() for each worker as the
fourth argument.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 06d1a442-209a-4753-a0a5-98352fae0697

📥 Commits

Reviewing files that changed from the base of the PR and between 4529235 and 1ad3d47.

📒 Files selected for processing (2)
  • model_gateway/src/policies/cache_aware.rs
  • model_gateway/src/routers/http/pd_router.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Inline review comments failed to post. This is likely due to GitHub's internal server error or limits when posting large numbers of comments. If you are seeing this consistently it is likely a permissions issue. Please check "Moderation" -> "Code review limits" under your organization settings.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/routers/http/pd_router.rs`:
- Around line 1578-1580: The call to handle_decode_error_response has arguments
in the wrong positions and types. The function expects the parameters in this
order: response, context reference, decode worker, and then load guards as a
Vec. The current code incorrectly passes prefill as the third argument instead
of decode, and passes decode (an Arc<dyn Worker>) as the fourth argument instead
of the required Vec<WorkerLoadGuard>. Fix this by passing decode as the third
argument and creating a Vec containing WorkerLoadGuard instances for both the
prefill and decode workers using WorkerLoadGuard::new() for each worker as the
fourth argument.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 06d1a442-209a-4753-a0a5-98352fae0697

📥 Commits

Reviewing files that changed from the base of the PR and between 4529235 and 1ad3d47.

📒 Files selected for processing (2)
  • model_gateway/src/policies/cache_aware.rs
  • model_gateway/src/routers/http/pd_router.rs
🛑 Comments failed to post (1)
model_gateway/src/routers/http/pd_router.rs (1)

1578-1580: ⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify if the test compiles with the current code
cd model_gateway && cargo test --no-run --test-threads=1 2>&1 | head -50

Repository: lightseekorg/smg

Length of output: 307


🏁 Script executed:

# Read the function signature around line 488-494
sed -n '488,494p' model_gateway/src/routers/http/pd_router.rs

Repository: lightseekorg/smg

Length of output: 286


🏁 Script executed:

# Read the test call around line 1578-1580
sed -n '1575,1585p' model_gateway/src/routers/http/pd_router.rs

Repository: lightseekorg/smg

Length of output: 409


🏁 Script executed:

# Find the variable declarations for prefill and decode in the test (around line 1550-1560)
sed -n '1545,1565p' model_gateway/src/routers/http/pd_router.rs

Repository: lightseekorg/smg

Length of output: 724


🏁 Script executed:

# Find WorkerLoadGuard definition to understand the type
rg "struct WorkerLoadGuard|enum WorkerLoadGuard" -A 5 model_gateway/src/

Repository: lightseekorg/smg

Length of output: 387


🏁 Script executed:

# Check if there are any imports or type aliases that might affect the types
sed -n '1,50p' model_gateway/src/routers/http/pd_router.rs

Repository: lightseekorg/smg

Length of output: 1535


Type mismatch: test passes Arc<dyn Worker> where Vec<WorkerLoadGuard> is expected.

The function signature at line 488-494 shows the parameters are res: reqwest::Response, context: &PDRequestContext<'_>, decode: Arc<dyn Worker>, and load_guards: Vec<WorkerLoadGuard>. However, the test at lines 1578-1580 passes decode_response, &context, prefill, and decode as arguments. The fourth argument is an Arc<dyn Worker> but the function expects a Vec<WorkerLoadGuard>. Additionally, passing prefill to the decode parameter is semantically incorrect—the decode worker should be passed instead.

Create load guards using WorkerLoadGuard::new() for both workers and pass them as the fourth argument, with decode as the third argument.

🐛 Proposed fix
-        let response = router
-            .handle_decode_error_response(decode_response, &context, prefill, decode)
-            .await;
+        let load_guards = vec![
+            WorkerLoadGuard::new(prefill.clone(), None),
+            WorkerLoadGuard::new(decode.clone(), None),
+        ];
+
+        let response = router
+            .handle_decode_error_response(decode_response, &context, decode, load_guards)
+            .await;
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

        let load_guards = vec![
            WorkerLoadGuard::new(prefill.clone(), None),
            WorkerLoadGuard::new(decode.clone(), None),
        ];

        let response = router
            .handle_decode_error_response(decode_response, &context, decode, load_guards)
            .await;
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/routers/http/pd_router.rs` around lines 1578 - 1580, The
call to handle_decode_error_response has arguments in the wrong positions and
types. The function expects the parameters in this order: response, context
reference, decode worker, and then load guards as a Vec. The current code
incorrectly passes prefill as the third argument instead of decode, and passes
decode (an Arc<dyn Worker>) as the fourth argument instead of the required
Vec<WorkerLoadGuard>. Fix this by passing decode as the third argument and
creating a Vec containing WorkerLoadGuard instances for both the prefill and
decode workers using WorkerLoadGuard::new() for each worker as the fourth
argument.

SYChen123 and others added 4 commits June 16, 2026 09:39
…h slower than decode

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
…drop dead decode param

Apply the (load, processed_requests, idx) tiebreak to the 4th least-load site (the event-driven KV-aware 'no overlap' fallback at cache_aware.rs) that the original change missed, so all least-load selections spread evenly when decode is faster than prefill. Remove the now-unused 'decode' parameter from create_streaming_response (Change B moved guard creation out but left it, failing clippy -D warnings) and its call sites. Run cargo +nightly fmt --all (fixes the mis-indented min_by_key closures + trailing whitespace).

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
… test

The empty-indexer fallthrough test routed a 4-token input, but PAGE_SIZE is
16, so the sequence was never cacheable: nothing was inserted and both
requests took the min-load branch. It only passed before because the
load-only tiebreak was stable by index. With the new
(load, processed_requests, idx) tiebreak, identical uncacheable requests
correctly spread across workers, so the test must use a >= PAGE_SIZE
sequence to exercise a genuine cache hit (which routes by tenant, not load).

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3228427ef3

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

.iter()
.min_by_key(|&&idx| workers[idx].load())
.min_by_key(|&&idx| {
(workers[idx].load(), workers[idx].processed_requests(), idx)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reserve worker before tree insertion

When several cold cache-miss requests are selected concurrently on different Tokio worker threads, this processed_requests() tie-breaker is not made visible until after match_and_insert_with finishes and increment_processed() runs below, while the WorkerLoadGuard is only created after select_pd_pair returns. For long prompts, multiple requests can therefore observe identical (load, processed_requests) values, all choose the lowest index, and insert for that same worker—the imbalance scenario this change is trying to avoid. Reserve or increment the chosen worker before the tree insertion, and apply the same ordering to the token/min-load paths.

Useful? React with 👍 / 👎.

@slin1237 slin1237 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm
pending ci

@slin1237
slin1237 merged commit 823cbbf into smg-project:main Jun 16, 2026
49 checks passed
slin1237 added a commit that referenced this pull request Jun 16, 2026
select_worker made several O(workers) passes per request (the healthy filter, the is_imbalanced load fold, and the cache-hit/miss worker scans), and each per-worker access — status, circuit breaker, load — took its own arc_swap guard. At high worker counts that per-worker guard traffic dominated routing CPU.

Read each worker once via a new Worker::routing_state() that shares the runtime guard for status+load+processed, gathering the healthy set, load min/max and the min-load index in a single pass. is_imbalanced and the min-load fallback consume the gathered bounds; the cache-hit tenant lookup is a hash-free scan over the gathered healthy indices (url() is a cheap field read). The (load, processed_requests, idx) min-load tie-break from #1714 rides the same guard, so it costs nothing extra.

Selection is unchanged except that a cache hit no longer routes onto a Ready-but-circuit-broken worker (it falls through to min-load, like the rest of the selection). No-GPU sim (4 threads, 2048 HTTP workers, shared-prefix load): cache_aware routing cost over round_robin drops ~2x at scale. 26 cache_aware + 107 policy + 207 worker tests pass.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
slin1237 added a commit that referenced this pull request Jun 16, 2026
select_worker made several O(workers) passes per request (the healthy filter, the is_imbalanced load fold, and the cache-hit/miss worker scans), and each per-worker access — status, circuit breaker, load — took its own arc_swap guard. At high worker counts that per-worker guard traffic dominated routing CPU.

Read each worker once via a new Worker::routing_state() that shares the runtime guard for status+load+processed, gathering the healthy set, load min/max and the min-load index in a single pass. is_imbalanced and the min-load fallback consume the gathered bounds; the cache-hit tenant lookup is a hash-free scan over the gathered healthy indices (url() is a cheap field read). The (load, processed_requests, idx) min-load tie-break from #1714 rides the same guard, so it costs nothing extra.

Selection is unchanged except that a cache hit no longer routes onto a Ready-but-circuit-broken worker (it falls through to min-load, like the rest of selection). No-GPU sim (4 threads, 2048 HTTP workers, shared-prefix load), A/B back-to-back: cache_aware routing is now within noise of round_robin at 2048 workers (+0.024 cpu_ms/req marginal, vs +0.18 with a per-request url map and ~+0.40 unoptimized). 26 cache_aware + 107 policy + 207 worker tests pass.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
slin1237 added a commit that referenced this pull request Jun 16, 2026
select_worker made several O(workers) passes per request (the healthy filter, the is_imbalanced load fold, and the cache-hit/miss worker scans), and each per-worker access — status, circuit breaker, load — took its own arc_swap guard. At high worker counts that per-worker guard traffic dominated routing CPU.

Read each worker once via a new Worker::routing_state() that shares the runtime guard for status+load+processed, gathering the healthy set, load min/max and the min-load index in a single pass. is_imbalanced and the min-load fallback consume the gathered bounds; the cache-hit tenant lookup is a hash-free scan over the gathered healthy indices (url() is a cheap field read). The (load, processed_requests, idx) min-load tie-break from #1714 rides the same guard, so it costs nothing extra.

Selection is unchanged except that a cache hit no longer routes onto a Ready-but-circuit-broken worker (it falls through to min-load, like the rest of selection). No-GPU sim (4 threads, 2048 HTTP workers, shared-prefix load), A/B back-to-back: cache_aware routing is now within noise of round_robin at 2048 workers (+0.024 cpu_ms/req marginal, vs +0.18 with a per-request url map and ~+0.40 unoptimized). 26 cache_aware + 107 policy + 207 worker tests pass.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@SYChen123

Copy link
Copy Markdown
Contributor Author

@slin1237 Thanks for your review and merge. Could you please merge the PR with the same modification in sglang as well? So that this fix can take effect in next version of sglang. Thanks a lot!
sgl-project/sglang#27547

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants