Skip to content

feat(gateway): Smooth transition from regular->PD - #1445

Merged
slin1237 merged 1 commit into
mainfrom
ekzhang/feature-pd-weighted-routing
May 6, 2026
Merged

slin1237 merged 1 commit into
mainfrom
ekzhang/feature-pd-weighted-routing

Conversation

@ekzhang

@ekzhang ekzhang commented May 4, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

It is not possible to transition an SMG deployment from regular routing to PD workers without downtime, or sending a storm of traffic to the first PD worker that starts up.

Solution

This adds support for smoothly transitioning from regular routing to PD routing and back, as well as between HTTP and gRPC worker engines.

Before if there was at least 1 P or D worker, all requests would be routed to PD mode, even if there were other regular workers. After this change, we weighted route requests between PD and regular workers based on the total worker count, and only send to prefill-decode disaggregated workers if min(P, D) >= 1.

Motivation here is to gradually transition a production deployment from regular to PD disaggregated mode without overloading a single worker with a storm of traffic as soon as it starts up. Instead, traffic will gradually be routed to the PD workers as they start up.

Also cleaned up the x-prefer-pd header which is only used when running in IGW mode without a model ID passed in, deprecated usage.

Changes

  • Weighted probabilistic router selection (pick_router_by_weights): instead of routing deterministically to the highest-priority router type, requests are now distributed proportionally across all eligible router types (gRPC-PD, HTTP-PD, gRPC-regular, HTTP-regular) by worker count. This applies to both per-model routing and the global no-model fallback path.
  • PD eligibility requires a complete pair: a prefill-decode router only receives weight if min(prefill_workers, decode_workers) >= 1 on that protocol. An incomplete PD deployment (e.g. only prefill workers registered so far) contributes zero weight and falls through to regular workers.
  • Unified routing function: replaced get_router_for_model with select_router_for_workers(&[Worker], model_id), which handles both the per-model and the no-model path with the same logic. select_router_for_request is now responsible for fetching the worker slice and both paths share the same default-router fallback.
  • Removed dead code: dropped the routers_snapshot field and its ArcSwap dependency, the x-prefer-pd request header (only took effect in the no-model IGW path and was otherwise silently ignored), the headers parameter from select_router_for_request, and the TODO scoring loop.

Test Plan

New unit test weighted_routing_splits_40_pd_60_regular registers 2 prefill + 2 decode + 6 regular workers against a model, runs 10,000 routing calls, and asserts that ~40% of requests are directed to the PD router (±5%). At n=10,000 the expected standard deviation of the ratio is ~0.005, so the ±5% band is ~10σ wide — false failure probability is negligible (~10⁻²⁴).

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Refactor

    • Improved router selection mechanism to dynamically distribute traffic based on available worker capacity instead of static preferences.
  • Tests

    • Added test coverage for multi-router weighted traffic distribution.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@github-actions github-actions Bot added the model-gateway Model gateway crate changes label May 4, 2026
@coderabbitai

coderabbitai Bot commented May 4, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@ekzhang has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 58 minutes and 5 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 88c1f735-020c-48e9-9c51-3e9d8b3f3bf0

📥 Commits

Reviewing files that changed from the base of the PR and between 95706ac and d5015ba.

📒 Files selected for processing (1)
  • model_gateway/src/routers/router_manager.rs
📝 Walkthrough

Walkthrough

RouterManager removes snapshot-based router selection and per-request header-driven PD preference, replacing it with worker-capability-driven weighted random routing. The new logic computes PD eligibility and weights from available prefill/decode/regular worker counts, then selects routers proportionally. All RouterTrait request handlers are updated to use the simplified select_router_for_request API.

Changes

Router Selection Refactor

Layer / File(s) Summary
Data Shape & API
model_gateway/src/routers/router_manager.rs (struct, signatures)
routers_snapshot field removed; get_router_for_model() method removed; select_router_for_request() signature drops headers parameter, retaining only model_id.
Core Weighted Routing Logic
model_gateway/src/routers/router_manager.rs (new helpers)
pick_router_by_weights implements weighted random selection among four router IDs; select_router_for_workers evaluates PD eligibility and derives weight distribution from prefill/decode/regular worker counts.
Request Handler Integration
model_gateway/src/routers/router_manager.rs (RouterTrait methods)
All ~17 RouterTrait implementations (health_generate, route_chat, route_completion, route_messages, route_responses, route_interactions, cancel_response, route_embeddings, route_classify, route_audio_transcriptions, route_rerank, route_realtime_*, etc.) updated to call select_router_for_request(model_id) without headers; external worker routing via provider_for_model remains unchanged.
Tests
model_gateway/src/routers/router_manager.rs (test suite)
PdStubRouter added to support testing; new weighted_routing_splits_40_pd_60_regular test registers PD and regular routers, creates a 2 prefill + 2 decode vs. 6 regular worker distribution, and asserts observed traffic split matches expected ~0.4 PD ratio.

Sequence Diagram

sequenceDiagram
    participant Client
    participant RouterManager
    participant WorkerRegistry
    participant WeightLogic
    participant Router
    
    Client->>RouterManager: route_request(model_id)
    RouterManager->>WorkerRegistry: get_workers(model_id or all)
    WorkerRegistry-->>RouterManager: [worker list]
    RouterManager->>WeightLogic: select_router_for_workers([workers])
    WeightLogic->>WeightLogic: count prefill/decode/regular workers
    WeightLogic->>WeightLogic: compute PD vs regular weights
    WeightLogic->>WeightLogic: pick_router_by_weights(weights)
    WeightLogic-->>RouterManager: selected_router_id
    RouterManager->>Router: execute_request()
    Router-->>Client: response
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

  • lightseekorg/smg#698: Modifies RouterManager public API and exported behavior in model_gateway/src/routers/router_manager.rs.
  • lightseekorg/smg#713: Directly related code-level changes to router-selection APIs and RouterManager signatures.

Suggested labels

model-gateway, tests, grpc

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐰 Router's New Path

No snapshots weighing down the way,
Just workers' counts to guide the day.
Prefill and decode dance in play,
While weights decide the traffic's sway—
Hop towards efficiency! 🎯

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 24.24% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title mentions a key aspect of the change (smooth transition from regular to PD routing), but it's somewhat vague about the core mechanism (weighted routing) that enables this transition.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ekzhang/feature-pd-weighted-routing

Tip

💬 Introducing Slack Agent: The best way for teams to turn conversations into code.

Slack Agent is built on CodeRabbit's deep understanding of your code, so your team can collaborate across the entire SDLC without losing context.

  • Generate code and open pull requests
  • Plan features and break down work
  • Investigate incidents and troubleshoot customer tickets together
  • Automate recurring tasks and respond to alerts with triggers
  • Summarize progress and report instantly

Built for teams:

  • Shared memory across your entire org—no repeating context
  • Per-thread sandboxes to safely plan and execute work
  • Governance built-in—scoped access, auditability, and budget controls

One agent for your entire SDLC. Right inside Slack.

👉 Get started


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@mergify

mergify Bot commented May 4, 2026

Copy link
Copy Markdown
Contributor

Hi @ekzhang, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 95706ac52b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +307 to +310
let workers = if let Some(model) = model_id {
self.worker_registry.get_by_model(model).to_vec()
} else {
// ZERO-ALLOCATION Snapshot Iteration (Hot Path Optimization)
// Atomic load avoids heap allocations and DashMap shard locks per-request
let routers_snapshot = self.routers_snapshot.load();
for router in routers_snapshot.iter() {
let mut score = 1.0;

let is_pd = router.is_pd_mode();
if prefer_pd && is_pd {
score += 2.0;
} else if !prefer_pd && !is_pd {
score += 1.0;
}
// TODO: Once routers expose worker stats, we can evaluate:
// - Average worker priority vs priority_threshold
// - Average worker cost vs max_cost
// - Current load and health status

if score > best_score && is_router_valid(is_pd) {
best_score = score;
best_router = Some(Arc::clone(router));
}
}
}
self.worker_registry.get_all()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep no-model requests off PD routers lacking cancel support

When model_id is absent, this now routes by global worker weights (get_all()), which means PD routers can be selected for endpoints like cancel_response that call select_router_for_request(None). Neither HTTP nor gRPC PD routers implement cancel_response, so they fall back to the RouterTrait default 501 Not Implemented path in routers/mod.rs; in mixed regular+PD deployments, cancel calls will fail intermittently depending on random selection. The previous no-model path preferred non-PD routing, so this introduces a user-visible regression specifically for cancellation flows.

Useful? React with 👍 / 👎.

Comment on lines +269 to +270
let grpc_pd = if grpc_prefill > 0 && grpc_decode > 0 {
grpc_prefill + grpc_decode

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Weight PD traffic by paired capacity, not total PD workers

PD routing weight is computed as prefill + decode once both are non-zero, but each PD request consumes one prefill and one decode worker, so effective capacity is bounded by the smaller side. In imbalanced rollouts (for example, many prefill workers but one decode worker), this formula over-allocates traffic to PD and can overload the scarce side, undermining the “smooth transition” goal of this change. Using paired capacity (for example based on min(prefill, decode)) would avoid this skew.

Useful? React with 👍 / 👎.

Comment on lines +219 to +227
let pick = ((rand::random::<f64>() * total as f64) as usize).min(total - 1);
let mut cum = 0usize;
for (weight, router_id) in &options {
cum += weight;
if pick < cum {
return self.routers.get(*router_id).map(|r| r.clone());
}
}
None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: If the randomly selected router ID has non-zero weight (workers exist) but the router itself isn't registered in self.routers, this returns None and falls through to the default router — even when other weighted router types are registered and would be valid.

For example: 3 gRPC-regular workers + 7 HTTP-regular workers, but only HTTP_REGULAR is registered. ~30% of requests would bypass HTTP_REGULAR and hit the default fallback instead of being routed proportionally.

In practice routers are all created at startup in from_config, so this should be rare, but a defensive approach would be to filter options to only routers present in self.routers before computing weights:

let options: Vec<(usize, &RouterId)> = [
    (grpc_pd, &router_ids::GRPC_PD),
    (http_pd, &router_ids::HTTP_PD),
    (grpc_regular, &router_ids::GRPC_REGULAR),
    (http_regular, &router_ids::HTTP_REGULAR),
]
.into_iter()
.filter(|(_, id)| self.routers.contains_key(*id))
.collect();

Comment on lines +269 to +278
let grpc_pd = if grpc_prefill > 0 && grpc_decode > 0 {
grpc_prefill + grpc_decode
} else {
None
}
0
};
let http_pd = if http_prefill > 0 && http_decode > 0 {
http_prefill + http_decode
} else {
0
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: The PD weight is prefill + decode, but the effective throughput of a PD pipeline is bottlenecked at min(prefill, decode). In an unbalanced deployment (e.g. 10 prefill + 1 decode + 6 regular), the PD weight would be 11 vs. 6 regular, routing ~65% of traffic to a pipeline bottlenecked at 1 decode worker.

Using 2 * min(prefill, decode) (or just min(prefill, decode)) would better reflect actual PD capacity relative to regular workers. For balanced deployments the two formulas are identical, so the change only affects unbalanced scaling.

Not blocking since the typical migration adds P and D workers in roughly equal numbers, but worth considering for robustness.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm I think it's more like prefill + decode in our case since each of the workers is around the same size

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean refactor. The weighted routing logic is correct and the test provides good statistical validation. Two minor nits posted inline (neither blocking):

  • 2 × 🟡 Nit: (1) pick_router_by_weights can return None when the randomly selected router isn't registered, even when other valid routers are available — filtering to registered routers before sampling would be more defensive. (2) PD weight uses prefill + decode rather than throughput-aware min(prefill, decode), which could over-route in unbalanced scaling scenarios.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/routers/router_manager.rs`:
- Around line 1021-1025: The test currently builds workers with
BasicWorkerBuilder::new(...) and relies on its default connection mode;
explicitly set the connection mode to HTTP by calling the builder method that
sets ConnectionMode::Http (e.g., add .connection_mode(ConnectionMode::Http) into
the chain before .build()) so BasicWorkerBuilder (and the created worker) cannot
switch to gRPC if defaults change; update usages where workers are constructed
(BasicWorkerBuilder::new, .worker_type, .labels, .circuit_breaker_config,
.build) to include the explicit connection mode.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: abd36245-0cfb-4d79-b9dd-b95c98d16c31

📥 Commits

Reviewing files that changed from the base of the PR and between 829604f and 95706ac.

📒 Files selected for processing (1)
  • model_gateway/src/routers/router_manager.rs

Comment thread model_gateway/src/routers/router_manager.rs Outdated
@slin1237 slin1237 self-assigned this May 5, 2026
@slin1237 slin1237 added the priority:high High priority label May 5, 2026
@slin1237 slin1237 assigned CatherineSue and unassigned slin1237 May 5, 2026
@ekzhang
ekzhang force-pushed the ekzhang/feature-pd-weighted-routing branch from 95706ac to 0fd0b52 Compare May 5, 2026 02:53
This adds support for smoothly transitioning from regular routing to PD routing and back, as well as between HTTP and gRPC worker engines.

Before if there was at least 1 P or D worker, all requests would be routed to PD mode, even if there were other regular workers. After this change, we weighted route requests between PD and regular workers based on the total worker count, and only send to prefill-decode disaggregated workers if min(P, D) >= 1.

Motivation here is to gradually transition a production deployment from regular to PD disaggregated mode without overloading a single worker with a storm of traffic as soon as it starts up. Instead, traffic will gradually be routed to the PD workers as they start up.

Also cleaned up the `x-prefer-pd` header which is only used when running in IGW mode without a model ID passed in, deprecated usage.

Signed-off-by: Eric Zhang <ekzhang1@gmail.com>
@ekzhang
ekzhang force-pushed the ekzhang/feature-pd-weighted-routing branch from 0fd0b52 to d5015ba Compare May 5, 2026 02:54
@ekzhang

ekzhang commented May 5, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai help

@coderabbitai

coderabbitai Bot commented May 5, 2026

Copy link
Copy Markdown

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.

CodeRabbit commands

These commands are invoked using PR/Issue comments.

  • @coderabbitai pause to pause the reviews on a PR.
  • @coderabbitai resume to resume the paused reviews.
  • @coderabbitai review to trigger an incremental review. This is useful when automatic reviews are disabled for the repository.
  • @coderabbitai full review to do a full review from scratch and review all the files again.
  • @coderabbitai rate limit to show your current review rate limit status.
  • @coderabbitai summary to regenerate the summary of the PR.
  • @coderabbitai evaluate custom pre-merge check --instructions <custom-checks-instructions> --name <custom-checks-title> [--mode <error|warning>] to validate the custom pre-merge checks instructions. Defaults to error when --mode is omitted.
  • @coderabbitai ignore pre-merge checks to override pre-merge checks and get an approval on PR.
  • @coderabbitai run pre-merge checks to run pre-merge checks on the pull request.
  • @coderabbitai generate docstrings to generate docstrings for this PR.
  • @coderabbitai generate sequence diagram to generate a sequence diagram of the changes in this PR.
  • @coderabbitai generate unit tests to generate unit tests for this PR.
  • @coderabbitai resolve merge conflict to automatically resolve merge conflicts.
  • @coderabbitai autofix to automatically fix issues identified in unresolved review comments.
  • @coderabbitai resolve to resolve all the CodeRabbit review comments.
  • @coderabbitai configuration to show the current CodeRabbit configuration for the repository.
  • @coderabbitai generate configuration to create a PR that adds the current resolved configuration as .coderabbit.yaml (or show it if already present).
  • @coderabbitai help to get help.

Other keywords and placeholders

  • Add @coderabbitai ignore or @coderabbit ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit configuration file (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • You can also validate your configuration using the online YAML validator.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Status, support, documentation and community

  • Visit our status page to check the current availability of CodeRabbit.
  • Create a ticket on our support page for assistance with any issues or questions.
  • Visit our documentation site for detailed information on how to use CodeRabbit.
  • Join our Discord community to connect with other users and get help from the community.
  • Follow us on X/Twitter for updates and announcements.

@ekzhang

ekzhang commented May 5, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented May 5, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@slin1237
slin1237 merged commit 845539e into main May 6, 2026
101 of 107 checks passed
@slin1237
slin1237 deleted the ekzhang/feature-pd-weighted-routing branch May 6, 2026 04:23
@coderabbitai coderabbitai Bot mentioned this pull request Jun 11, 2026
2 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes priority:high High priority

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants