Skip to content

fix(router): answer 503 on the gRPC path when a model's workers are all unavailable - #2465

Merged
smg-project-bot merged 1 commit into
mainfrom
fix/grpc-unavailable-workers-503
Sep 8, 2026
Merged

smg-project-bot merged 1 commit into
mainfrom
fix/grpc-unavailable-workers-503

Conversation

@hello-alexmcc

Copy link
Copy Markdown
Collaborator

Description

Problem

The gRPC selection stage answered 404 model_not_found whenever no worker could take a request, no matter why. An unknown model and a model whose only worker is restarting produced the same answer. The HTTP router already distinguishes the two (routers/http/router.rs): NoCandidates is a 404, Unavailable / PolicyDeclined is a 503 no_available_workers.

For a client of the gRPC path this reads as "the model does not exist" during any worker outage, and for a disaggregated model it happens whenever one leg loses its last worker. It surfaced while writing the PD topology suite (#2464), whose sole-leg outage test had to accept both answers, and it is what the worker-restart test on #2462 hit.

Solution

  • selection_failure now collects each demanded leg's verdict: a shed is returned as built (unchanged), Unavailable / PolicyDeclined on any leg yields the 503, and only when every leg has no candidates at all does the 404 remain.
  • pair_failure maps the same way for the pair verdict.
  • disaggregated_leg_shed becomes disaggregated_leg_verdict and returns the full PlacementFailure via placement::failure_from, so the per-leg fallback can tell "nobody serves this leg" from "the leg's workers are down".
  • Message: All workers for model '<id>' are unavailable (unhealthy or circuit breaker open), the HTTP router's code no_available_workers.

Changes

  • model_gateway/src/routers/grpc/common/stages/worker_selection.rs: mapping above; two unit tests (an_unavailable_regular_worker_answers_503_not_404, an_unavailable_decode_leg_answers_503_not_404); the existing 404 tests for absent models still pass.

Test Plan

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes (-p smg --all-targets; the all-features build needs OpenCV locally)
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

…ll unavailable

The gRPC selection stage answered 404 model_not_found whenever no worker
could take a request, whether the model was unknown or its workers were
merely unhealthy or circuit-broken. The HTTP router already separates the
two: an unknown model is a 404, workers that exist but cannot serve are a
503 no_available_workers. A client of the gRPC path was told the model did
not exist while its workers restarted, and for a disaggregated model the
same happened whenever one leg lost its last worker.

Both the per-leg fallback and the pair verdict now map Unavailable and
PolicyDeclined to the 503 with the HTTP router's code and keep 404 for a
model nobody serves. Unit tests pin the regular and the decode-leg cases.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Sep 8, 2026
hello-alexmcc added a commit that referenced this pull request Sep 8, 2026
The first CI run of the worker-restart test caught the gateway answering
404 model_not_found while the only worker was restarting, on SGLang and on
vLLM. That answer is fixed in #2465; the test now asserts the 503 with
no_available_workers instead of any 5xx. The queue-depth helper raises
instead of calling pytest.fail so mypy sees the missing return without
pytest stubs, which is how the lint job runs it.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Team

Run ID: eb9540bf-dbb9-4b40-9d95-56fdde2652eb

📥 Commits

Reviewing files that changed from the base of the PR and between a4c6987 and 982189f.

📒 Files selected for processing (1)
  • model_gateway/src/routers/grpc/common/stages/worker_selection.rs

Included review availability: 4 reviews are currently available. Your included PR review attempts over the past 7 days set your current allowance at 8 reviews per hour.


📝 Summary

Summary by CodeRabbit

  • Bug Fixes
    • Improved error handling when processing requests across worker pools.
    • Requests now return a service-unavailable response when no eligible workers are available or policy restrictions prevent selection.
    • Requests continue to return a not-found response when no matching model is available.
    • Improved failure reporting for multi-stage processing, including cases where decoding workers are unavailable.
    • Added coverage for worker availability scenarios to improve reliability.

Walkthrough

Worker selection now distinguishes unavailable or policy-declined workers from absent models. Regular, disaggregated, and paired selections return 503 no_available_workers or 404 model_not_found as appropriate. Tests cover these outcomes.

Changes

Worker selection failure classification

Layer / File(s) Summary
Failure classification and coverage
model_gateway/src/routers/grpc/common/stages/worker_selection.rs
Regular, disaggregated, and paired selections now evaluate complete placement failures. Unavailable or policy-declined workers return 503 no_available_workers. Missing candidates return 404 model_not_found. Tests cover regular and decode-worker unavailability. The unused overload import was removed.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Merge Risk: ⚪ Minimal · up to 98218

Worker-selection failures now distinguish unavailable workers from unknown models, returning 503 for temporary worker unavailability and retaining 404 for absent models. Covered regular and disaggregated failure paths have no identified merge-blocking risk.

Sequence Diagram(s)

sequenceDiagram
  participant WorkerSelection
  participant PlacementFailure
  participant gRPCResponse
  WorkerSelection->>PlacementFailure: evaluate candidate placement
  PlacementFailure-->>WorkerSelection: return failure verdict
  WorkerSelection->>gRPCResponse: return 503 for unavailable workers
  WorkerSelection->>gRPCResponse: return 404 for absent models
Loading
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the main change: the gRPC path now returns 503 when all workers for a model are unavailable.
Description check ✅ Passed The description directly explains the worker-availability problem, the 503 versus 404 behavior, the affected regular and disaggregated paths, and the validation performed.
Docstring Coverage ✅ Passed Docstring coverage is 87.50% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 8 functions across 1 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/grpc-unavailable-workers-503

Comment @coderabbitai help to get the list of available commands.

Comment on lines +359 to +369
match verdict {
PlacementFailure::AllOverloaded(shed) => return shed,
PlacementFailure::Unavailable | PlacementFailure::PolicyDeclined(_) => {
unavailable = true;
}
PlacementFailure::NoCandidates => {}
}
}
if unavailable {
return self.workers_unavailable(model_id);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Important: A healthy leg reports Unavailable, so the 404 this PR intends to keep for "nobody serves this leg" is unreachable whenever any other leg has workers.

placement::failure_from is a three-way classifier that only knows empty / all-overloaded / everything else — it returns Unavailable for any non-empty pool that isn't fully vetoed, including a pool whose workers are all Ready and available. That is sound in the single-leg callers (single_failure runs only after selection over that exact pool failed), but here the loop asks the question of every demanded leg, including legs that were perfectly fine.

Concrete failure: EPD mode, a model with healthy prefill+decode workers and zero encode workers registered, request carries encode items. select_encode_prefill_decode_workers returns None and legs = [Prefill, Decode, Encode]:

  • Prefill → pool non-empty, workers available → Unavailable → unavailable = true
  • Decode → same
  • Encode → NoCandidates

Result: 503 no_available_workers / "All workers for model 'X' are unavailable (unhealthy or circuit breaker open)" — for a permanent deployment gap, in which no worker is unhealthy and no breaker is open. Before this PR that was a 404. The 503 is retryable, so clients and proxies now retry a misconfiguration forever instead of failing fast. The same happens for the mixed-runtime EPD case (no shared runtime across legs) with legs = [Prefill, Decode].

The verdict for a leg is only meaningful if that leg is the one that had nothing to give. Consider deriving unavailable from the availability of each leg's own pool, e.g. treat a leg as blocking only when it has candidates but none passes is_available():

PlacementFailure::Unavailable | PlacementFailure::PolicyDeclined(_) => {
    // `failure_from` cannot distinguish "this leg is drained" from
    // "this leg was fine and another leg is why we are here", so ask
    // the pool directly.
    unavailable |= !leg_has_available_worker(...);
}

or, since the shared-runtime/encode-leg cause is already known at the call site, pass the failing leg in rather than re-deriving it from all of them.

Comment on lines +1296 to +1301
WorkerSelectionMode::PrefillDecode,
);

let fallback =
stage.selection_failure(model_id, &[WorkerType::Prefill, WorkerType::Decode], None);
let pair = stage.pair_failure(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: an_unavailable_decode_leg_answers_503_not_404 passes without exercising the decode leg at all — the fallback half of it would be green even if decode.set_status(WorkerStatus::NotReady) were deleted.

register_pd_workers builds Ready workers with health checks disabled, so when selection_failure walks [Prefill, Decode] the prefill leg is the one that sets unavailable = true: its pool is non-empty and not all-overloaded, so failure_from says Unavailable even though that worker is fine (see the comment on selection_failure). The NotReady decode worker never influences the outcome.

The pair and absent halves do test real mappings, since they hand pair_failure a verdict directly. To make the fallback half meaningful, assert the negative too — with every leg healthy, selection_failure should not answer 503:

// Sanity: the 503 must come from the downed leg, not from a healthy
// leg that `failure_from` cannot distinguish from a drained one.
let healthy = WorkerRegistry::new();
register_pd_workers(&healthy, model_id, 1);
let healthy_stage = WorkerSelectionStage::new(
    Arc::new(healthy),
    Arc::new(PolicyRegistry::new(PolicyConfig::RoundRobin)),
    WorkerSelectionMode::PrefillDecode,
);
assert_ne!(
    healthy_stage
        .selection_failure(model_id, &[WorkerType::Prefill, WorkerType::Decode], None)
        .status(),
    StatusCode::SERVICE_UNAVAILABLE
);

That assertion fails today, which is exactly the bug in the loop.

Comment on lines +384 to +395
fn workers_unavailable(&self, model_id: &str) -> Response {
error!(
function = "WorkerSelectionStage::execute",
mode = ?self.mode,
model_id = %model_id,
"No available workers for model"
);
error::service_unavailable(
"no_available_workers",
format!("All workers for model '{model_id}' are unavailable (unhealthy or circuit breaker open)"),
)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: On the reselect path this 503 is retryable, so it converts a fail-fast into max_retries reselect attempts with backoff — worth an explicit decision rather than an accident of the status code.

reselect reaches selection_failure from inside the attempt loop (pipeline.rs:457), and the loop gates on is_retryable_response(&failure) (pipeline.rs:610). A 404 was terminal there; this 503 is not marked non-retryable, so a request whose workers are down now burns every remaining attempt: each one calls record_error and the retry counters, so one client request inflates the error/retry metrics N-fold and pays the full backoff before returning the same 503.

overload.rs faced exactly this and chose mark_non_retryable, with the reasoning that "the veto clears at the poll interval, which no backoff window outlives". Health and circuit-breaker state can flip faster than the overload flag, so retrying here is arguably right — but if that's the intent, please say so in the doc comment (it currently says "the client should retry", which reads as advice to the client, not a statement about the internal retry layer) and consider a test pinning the choice, mirroring "a shed must be terminal for the retry layer" at line 924.

Separately, the doc on selection_failure (lines 321-322) is now stale — "a 503 shed when a leg's whole candidate pool is vetoed, the existing 404 otherwise" no longer describes the three outcomes this function has.

Comment on lines +404 to +406
PlacementFailure::Unavailable | PlacementFailure::PolicyDeclined(_) => {
self.workers_unavailable(model_id)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: Two of the three verdicts routed here produce a factually wrong message.

workers_unavailable says "All workers for model 'X' are unavailable (unhealthy or circuit breaker open)", but:

  • PolicyDeclined(policy) is raised by select_pair after both legs were filtered for availability (placement.rs: "Both legs were filtered for availability above, so a miss here is the policy's own decision, never an overloaded pool"). Every worker is healthy and no breaker is open — the policy just picked none. The HTTP router distinguishes these (router.rs:829-833 picks "Policy returned no eligible worker" when non_dp_workers.iter().any(|w| w.is_available())); this path does not, so the claimed parity is only partial.
  • The homogeneous-runtime narrowing in select_pair hardcodes PlacementFailure::Unavailable for "No available PD pair for runtime {runtime}". A PD model deployed with an SGLang prefill and a vLLM decode is a permanent configuration error with every worker healthy, and it now reports a transient health failure instead of the previous 404.

Since pair_failure has failure.leg in hand, it could pick the message the way the HTTP router does, or workers_unavailable could take the message so PolicyDeclined names the policy.

@smg-project-bot
smg-project-bot merged commit fb508f8 into main Sep 8, 2026
49 of 55 checks passed
@smg-project-bot
smg-project-bot deleted the fix/grpc-unavailable-workers-503 branch September 8, 2026 01:51
hello-alexmcc added a commit that referenced this pull request Sep 8, 2026
The first CI run of the worker-restart test caught the gateway answering
404 model_not_found while the only worker was restarting, on SGLang and on
vLLM. That answer is fixed in #2465; the test now asserts the 503 with
no_available_workers instead of any 5xx. The queue-depth helper raises
instead of calling pytest.fail so mypy sees the missing return without
pytest stubs, which is how the lint job runs it.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 8, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 8, 2026
The first CI run of the worker-restart test caught the gateway answering
404 model_not_found while the only worker was restarting, on SGLang and on
vLLM. That answer is fixed in #2465; the test now asserts the 503 with
no_available_workers instead of any 5xx. The queue-depth helper raises
instead of calling pytest.fail so mypy sees the missing return without
pytest stubs, which is how the lint job runs it.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 9, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 9, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 9, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 9, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 9, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 9, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 9, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 10, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
hello-alexmcc added a commit that referenced this pull request Sep 10, 2026
…ow bursts, prefill aborts, batched choices

The nightly sweep covered one transport per engine and left the two
failure modes we just fixed without a regression test. It now also runs:

- SGLang's HTTP PD router at 1p1d and 2p2d (its own pairing, KV handoff
  and error answers); the other engines skip those rows
- vLLM over Mooncake next to its NIXL default, as the PR lane does
- a decode pinned to a four-request window against a burst six times
  wider: every request is served or shed with the overload code and a
  Retry-After, the burst finishes in bounded time, and no engine leg logs
  a bootstrap timeout (the gateway's admission gate at work)
- clients that drop the connection before the first token, which must
  leave both legs idle like an abandoned stream does
- a completion with four choices, which fans out one room per choice

The sole-leg outage now asserts the 503 the gateway answers since #2465.
The PR lane keeps one gRPC topology plus the runtime-assembled fleet and
adds the over-window burst; HTTP PD runs nightly only.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants