Skip to content

feat(worker): add Draining state with workflow-driven drain - #1491

Merged
slin1237 merged 3 commits into
mainfrom
feat/worker-draining-state
May 14, 2026
Merged

slin1237 merged 3 commits into
mainfrom
feat/worker-draining-state

Conversation

@slin1237

@slin1237 slin1237 commented May 14, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Adds WorkerStatus::Draining = 4 so policies stop selecting a worker the moment RemoveWorker starts, closing the window where a request could be routed to a worker that's about to be torn down.
  • New DrainWorkersStep workflow step (inserted between find_workers_to_remove and remove_from_policy_registry) transitions each Ready worker in the removal batch to Draining (revision-checked, so a same-URL replace isn't clobbered) and sleeps the max drain_settle_secs across the batch. Non-Ready workers (Pending/NotReady/Failed) skip both transition and sleep — keeps --remove-unhealthy-workers fast.
  • Drain semantics live in the workflow, not service discovery: every Job::RemoveWorker (K8s pod deletion, --remove-unhealthy-workers, manual API) gets uniform drain behaviour.
  • Configurable via --drain-settle-secs (default 5s) on HealthCheckConfig::drain_settle_secs, with per-worker overrides through WorkerSpec::health.drain_settle_secs (mirrors the existing HealthCheckUpdate pattern). Set to 0 to skip draining entirely.
  • WorkerRegistry::get_id_by_url helper added so the step can call transition_status_if_revision without scanning all workers.
  • Job-queue caller timeout bumped to 30 + drain_settle_secs to accommodate the drain window.

Test plan

  • cargo +nightly fmt --all (silent)
  • cargo clippy --all-targets --all-features -- -D warnings (clean)
  • cargo test — 3518 passed / 0 failed across the workspace.
  • WorkerStatus tests (crates/protocols): try_from_u8(4), from_u8(4), Display → "draining", serde round-trip Draining ↔ "draining", is_routable() false for Draining, parametrized routability check.
  • HealthCheckConfig tests (crates/protocols): default drain_settle_secs == 5; HealthCheckUpdate override; is_empty includes the new field; existing JSON without the field deserializes (serde default).
  • policies test: get_healthy_worker_indices parametrized — Ready included, Pending/NotReady/Failed/Draining excluded.
  • registry tests: get_id_by_url (registered + missing); transition_status to Draining emits WorkerEvent::StatusChanged (proves mesh sync forwards Draining the same way as Ready/NotReady).
  • DrainWorkersStep tests (paused-time, tokio::time::advance): empty list → no-op; only-non-Ready → no transition + no sleep; single Ready → immediate Draining + sleep until window elapses; multiple workers → uses max drain_settle_secs; drain_settle_secs == 0 → still transitions to Draining (so policies stop) but skips the sleep.
  • Manual smoke against a live K8s cluster: deploy a worker, scale to zero, observe Draining → wait → removed (left to reviewer).

Summary by CodeRabbit

  • New Features

    • New CLI option --drain-settle-secs (default: 5) to control graceful drain duration.
    • Workers gain a Draining status to support graceful shutdowns.
    • Worker lookup by URL added.
    • New drain step in the removal workflow that transitions Ready workers to Draining and waits the configured settle window.
  • Behavior Changes

    • Draining workers are excluded from routing/probing and treated as terminal for health transitions.
    • Toggle-health flows now include the drain-settle parameter.

Review Change Stack

Selecting a worker that has just received a `RemoveWorker` job opens a
narrow window where in-flight requests can be sent to a worker that is
about to be torn down. Add a transitional `WorkerStatus::Draining` so
policies stop selecting the worker the moment removal starts, and a
new `DrainWorkersStep` in the worker_removal workflow that holds the
worker in `Draining` for a configurable settle window before the
existing remove-from-* steps run.

Drain semantics live in the workflow, not service discovery: any
`Job::RemoveWorker` (K8s pod deletion, --remove-unhealthy-workers,
manual API) gets uniform draining. The step skips the sleep entirely
when no worker in the batch is Ready, so cleanup of broken workers
stays fast.

Per-worker overrides follow the existing `HealthCheckUpdate` pattern:
default lives on `HealthCheckConfig::drain_settle_secs` (5s, exposed
as --drain-settle-secs), and any worker may override via its
`WorkerSpec::health.drain_settle_secs`. The step takes the max
across the batch for a single sleep.

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@github-actions github-actions Bot added python-bindings Python bindings changes dependencies Dependency updates protocols Protocols crate changes model-gateway Model gateway crate changes labels May 14, 2026
@coderabbitai

coderabbitai Bot commented May 14, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 43048147-574b-4b10-8526-d5b87b9f7926

📥 Commits

Reviewing files that changed from the base of the PR and between 5d15749 and 250f74b.

📒 Files selected for processing (2)
  • model_gateway/src/workflow/steps/local/drain_workers.rs
  • model_gateway/src/workflow/steps/local/mod.rs

📝 Walkthrough

Walkthrough

This PR implements a graceful worker draining mechanism by introducing a Draining worker status, a configurable settle period, and a new workflow step. Workers transition to Draining when removal is requested, remain unavailable for routing, and complete in-flight requests before final removal from the registry. The settle duration is configurable via CLI and wired through the full configuration stack.

Changes

Worker Draining with Configurable Settle Period

Layer / File(s) Summary
Protocol types and drain configuration
crates/protocols/src/worker.rs
Adds WorkerStatus::Draining = 4 as a non-routable state with u8 conversions and Display impl. HealthCheckConfig gains drain_settle_secs: u64 with serde defaults, and HealthCheckUpdate adds an optional drain_settle_secs override with merge semantics. Comprehensive tests cover conversions, serialization round-trips, routability, and backward compatibility.
Configuration propagation and bindings
model_gateway/src/main.rs, model_gateway/src/config/types.rs, bindings/python/src/lib.rs, model_gateway/Cargo.toml
CLI argument --drain-settle-secs flows through CliArgs → HealthCheckConfig → ProtocolHealthCheckConfig and RouterConfig. Python bindings expose the parameter with default 5. Tokio test-util feature added for time control in tests.
Worker state machine draining support
model_gateway/src/worker/manager.rs, model_gateway/src/policies/mod.rs
Health-check state machine treats Draining as terminal alongside Failed, preventing transitions on probe failure. Test ensures get_healthy_worker_indices excludes all non-Ready states including Draining. Test fixture configs updated.
Worker registry drain transition support
model_gateway/src/worker/registry.rs
Adds WorkerRegistry::get_id_by_url() helper for reverse URL lookup. Tests verify lookup returns correct IDs and that transitioning to Draining emits WorkerEvent::StatusChanged with expected old/new status.
DrainWorkersStep workflow implementation
model_gateway/src/workflow/steps/local/drain_workers.rs
New step transitions only Ready workers to Draining, computes maximum drain_settle_secs across transitioned workers, and sleeps for that duration before succeeding. Includes test helpers and five test cases covering empty lists, non-ready workers, settle delays, maximum settle window calculation, and zero-duration drains.
Workflow DAG wiring and job queue timeout
model_gateway/src/workflow/steps/local/mod.rs, model_gateway/src/workflow/job_queue.rs, model_gateway/src/worker/builder.rs, model_gateway/src/worker/worker.rs
Integrates DrainWorkersStep between find_workers_to_remove and downstream removal, making remove_from_policy_registry depend on drain_workers. Job queue RemoveWorker timeout extended from fixed 30s to 30 + drain_settle_secs. All test fixtures updated with the new configuration field.
TUI command integration
tui/src/app.rs
Updates :toggle-health command and action-menu ToggleHealthCheck to include drain_settle_secs: None in HealthCheckUpdate payload.

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant Workflow
  participant DrainStep
  participant Registry
  participant Sleep

  Client->>Workflow: request RemoveWorker(workers)
  Workflow->>DrainStep: execute(context)
  DrainStep->>Registry: transition_ready_to_draining(id, rev)
  Registry-->>DrainStep: transition result
  DrainStep->>Sleep: sleep(max_drain_settle_secs)
  Sleep-->>DrainStep: wake
  DrainStep-->>Workflow: return Success
  Workflow->>Client: completion / result
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested labels

tests

Suggested reviewers

  • CatherineSue
  • key4ng
  • gongwei-130

Poem

🐰 I nudged the workers, soft and kind,
They finished tasks they left behind,
A settle pause — a gentle rest,
Then off they hop to seek the nest,
Hooray — the drain was simply blessed.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat(worker): add Draining state with workflow-driven drain' accurately describes the main changes: introducing a new WorkerStatus::Draining state and implementing workflow-driven drain behavior via the new DrainWorkersStep.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/worker-draining-state

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.


Comment @coderabbitai help to get the list of available commands and usage tips.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f89b15d64d

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/workflow/job_queue.rs Outdated
Comment thread model_gateway/src/workflow/job_queue.rs Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@bindings/python/src/lib.rs`:
- Around line 789-790: The new parameter drain_settle_secs was inserted
mid-signature of _Router which breaks positional-arg callers; update the API by
either moving drain_settle_secs to the end of the _Router/__init__ parameter
list or make new parameters keyword-only (add a * before it in the signature) so
existing positional usage (e.g., Router/Robot wrapping) remains compatible;
adjust the _Router constructor signature and any internal callers (e.g.,
Robot(_Router(...))) accordingly and add a brief note/test ensuring external
positional calls continue to work.

In `@model_gateway/src/workflow/job_queue.rs`:
- Around line 397-402: The timeout for the RemoveWorker path is currently built
using only the global default (router_config.health_check.drain_settle_secs)
which can undercut longer per-worker overrides; update the timeout calculation
where timeout_duration is set in job_queue.rs (the RemoveWorker handling) to use
the worker's effective drain_settle_secs when present (e.g., take max(global
router_config.health_check.drain_settle_secs, worker.override.drain_settle_secs)
or otherwise use the worker-config value) so the Duration::from_secs uses the
per-worker override if larger; ensure you reference the worker config / job
payload that contains the override when computing the final timeout.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 3e73c7b6-5833-4754-a1f6-5ba4b7e094a9

📥 Commits

Reviewing files that changed from the base of the PR and between 253c79d and f89b15d.

📒 Files selected for processing (14)
  • bindings/python/src/lib.rs
  • crates/protocols/src/worker.rs
  • model_gateway/Cargo.toml
  • model_gateway/src/config/types.rs
  • model_gateway/src/main.rs
  • model_gateway/src/policies/mod.rs
  • model_gateway/src/worker/builder.rs
  • model_gateway/src/worker/manager.rs
  • model_gateway/src/worker/registry.rs
  • model_gateway/src/worker/worker.rs
  • model_gateway/src/workflow/job_queue.rs
  • model_gateway/src/workflow/steps/local/drain_workers.rs
  • model_gateway/src/workflow/steps/local/mod.rs
  • tui/src/app.rs

Comment thread bindings/python/src/lib.rs Outdated
Comment thread model_gateway/src/workflow/job_queue.rs Outdated
/// callers that need to invoke `transition_status_if_revision` with
/// the current worker revision.
pub fn get_id_by_url(&self, url: &str) -> Option<WorkerId> {
self.url_to_id.get(url).map(|id| id.clone())

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: Clippy preference — .map(|id| id.clone()) can be .cloned().

Suggested change
self.url_to_id.get(url).map(|id| id.clone())
self.url_to_id.get(url).map(|id| id.clone())

(Actually, DashMap's Ref doesn't impl the right trait for .cloned() on the Option<Ref> — disregard if that's the case here.)

Comment on lines +638 to 642
WorkerStatus::Failed | WorkerStatus::Draining => {
// Terminal for the health-state machine. Failed is removed
// by `--remove-unhealthy-workers`; Draining is removed by
// the discovery drain timer once in-flight requests settle.
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: Draining is correctly terminal in the state machine, but launch_due_probes (line ~399) only short-circuits for Failed — Draining workers still get health-probed every interval even though compute_next_status always returns None for them.

Consider adding Draining alongside Failed in the early-exit check in launch_due_probes to skip pointless probes during the drain window:

if launched_status == WorkerStatus::Failed || launched_status == WorkerStatus::Draining {
    next_check.remove(&worker_id);
    ...

(For Draining you'd skip the removal push since drain-initiated removal handles that.)

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, well-structured addition of the Draining worker state. The workflow-driven drain step, config plumbing across CLI/Python/protocol layers, and comprehensive tests all look solid.

Summary: 0 🔴 Important · 3 🟡 Nit · 0 🟣 Pre-existing

Nits:

  • job_queue.rs: Timeout for wait_for_completion uses only the global drain_settle_secs — per-worker overrides could cause spurious timeout errors.
  • registry.rs: Minor style — .map(|id| id.clone()) → .cloned().
  • manager.rs: Draining workers still get health-probed even though compute_next_status treats them as terminal. Consider skipping probes like Failed does.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a Draining status for workers and a configurable drain_settle_secs period to allow in-flight requests to finish before a worker is removed from the registry. Key changes include the addition of the DrainWorkersStep to the worker removal workflow, updates to the health check configuration, and new CLI parameters. Review feedback highlights a potential timeout issue in the job queue when per-worker overrides exceed the global default and suggests refactoring the use of '0' as a special value for disabling the settle period to align with project conventions.

Comment thread model_gateway/src/workflow/job_queue.rs Outdated
Comment thread model_gateway/src/workflow/steps/local/drain_workers.rs Outdated
- bindings/python: move `drain_settle_secs` to the end of the
  `_Router` parameter list so external positional callers aren't
  broken by the mid-signature insertion (CodeRabbit, Codex).
- workflow/job_queue: cap the `RemoveWorker` caller wait at
  `30 + max(global_drain_settle_secs, 600)` so a per-worker
  override larger than the global default no longer produces a
  spurious timeout error from the queue while the workflow is
  still draining (Codex, CodeRabbit, Claude, Gemini).
- worker/manager: short-circuit `launch_due_probes` for `Draining`
  workers the same way `Failed` is handled — `compute_next_status`
  is already a no-op for them, and the drain workflow owns
  removal so the probe pass should not push another `RemoveWorker`
  candidate (Claude).
- workflow/drain_workers: drop the `max_drain_secs == 0`
  short-circuit; `sleep(Duration::ZERO)` is a no-op, and removing
  the special-case keeps the step from inventing a new "0 disables"
  convention (Gemini).

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c5f0987402

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

// Timeout is generous so the per-worker drain_settle_secs sleep
// can complete even when several workers stack their settle
// windows. The step itself caps by taking max() across them.
.with_timeout(Duration::from_secs(600))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid timing out valid drain windows

When any worker (or the global CLI/config) sets drain_settle_secs above 600, DrainWorkersStep sleeps that full value but the workflow engine wraps this step in this fixed 600s timeout (tokio::time::timeout in the workflow engine marks it failed). In that configuration the worker has already been moved to Draining, the downstream removal steps never run, and the worker can remain stuck out of service instead of being removed. Even with the new caller wait timeout, the step timeout itself still needs to be derived from or capped against the configured drain window.

Useful? React with 👍 / 👎.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/workflow/steps/local/drain_workers.rs`:
- Around line 46-64: The loop currently uses the snapshot `worker` (from
`workers_to_remove`) to decide drain eligibility and revision; instead, fetch
the live registry entry before deciding and transitioning: use
`app_context.worker_registry.get_id_by_url(url)` (as already present) then
retrieve the current registry entry (e.g. via the registry's
get/get_by_id/get_entry method) and use that entry's status and revision for the
`WorkerStatus::Ready` check and for the call to
`transition_status_if_revision(&worker_id, revision, WorkerStatus::Draining)`;
only proceed if the registry's status is `Ready`, and read
`health_config.drain_settle_secs` from the registry entry's metadata for
updating `max_drain_secs`, so decisions are based on current registry state
rather than the snapshot `worker`.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 22a44e95-5caf-4636-9087-c86909a656db

📥 Commits

Reviewing files that changed from the base of the PR and between f89b15d and c5f0987.

📒 Files selected for processing (4)
  • bindings/python/src/lib.rs
  • model_gateway/src/worker/manager.rs
  • model_gateway/src/workflow/job_queue.rs
  • model_gateway/src/workflow/steps/local/drain_workers.rs

Comment thread model_gateway/src/workflow/steps/local/drain_workers.rs Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5d15749ad9

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread model_gateway/src/workflow/steps/local/drain_workers.rs
Comment on lines +100 to +106
// No `with_timeout`: the step's purpose is to sleep for the
// resolved `drain_settle_secs` (which can be set per-worker
// via `WorkerSpec::health.drain_settle_secs`). A static
// workflow-level timeout would be set at definition time
// without visibility into runtime config and would
// pre-emptively fail the workflow for legitimately long
// drain windows, leaving workers stuck in `Draining`.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: This is now the only step in the entire codebase without a with_timeout. Every other step — including variable-duration ones — sets an upper bound (e.g., mcp_registration uses 7200s, tokenizer_registration uses 300s). Removing the timeout entirely means a misconfigured drain_settle_secs (or a future bug in the sleep path) would hang the workflow indefinitely, leaving workers stuck in Draining.

The previous 600s cap was too rigid, but the fix could be a generous safety-net timeout rather than none at all — e.g., with_timeout(Duration::from_secs(3600)) gives a 1-hour ceiling that no realistic drain window should hit, while still preventing unbounded hangs.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@model_gateway/src/workflow/steps/local/drain_workers.rs`:
- Around line 297-305: The test incorrectly shares the same Arc-backed worker
between snapshot and live registry so calling
worker.set_status(WorkerStatus::Ready) mutates both; update the test to make the
snapshot contain a distinct worker instance (e.g. create a second Arc via
build_worker or clone a new worker with the same id/status) instead of reusing
the live Arc so the snapshot stays Pending while the live worker is set to
Ready; adjust the snapshot creation (the snapshot variable) while leaving
make_app_context and the live worker unchanged, and ensure any other occurrences
(around lines with make_context) use the independent snapshot instance.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 69cdb44d-8ebb-4ff8-9293-03d386a6e987

📥 Commits

Reviewing files that changed from the base of the PR and between c5f0987 and 5d15749.

📒 Files selected for processing (2)
  • model_gateway/src/workflow/steps/local/drain_workers.rs
  • model_gateway/src/workflow/steps/local/mod.rs

Comment thread model_gateway/src/workflow/steps/local/drain_workers.rs
- workflow/drain_workers: read status / revision / drain_settle_secs
  from the live registry entry, not the snapshot captured by
  find_workers_to_remove. A worker that became `Ready` between the
  snapshot and this step (eg. a probe just succeeded) would
  otherwise bypass the transition and start serving traffic up
  until the registry-removal step ran. Adds a regression test that
  flips a Pending → Ready transition between snapshot and execute
  and confirms the step still drains it (CodeRabbit).
- workflow/local/mod: drop the static `with_timeout(600s)` on the
  drain step. Workflow-level timeouts are set at definition time
  so they cannot track the runtime `drain_settle_secs`, and a
  per-worker override above 600s would otherwise pre-emptively
  fail the workflow with the worker stuck in `Draining` (Codex).

Signed-off-by: Simo Lin <25425177+slin1237@users.noreply.github.com>
@slin1237
slin1237 force-pushed the feat/worker-draining-state branch from 5d15749 to 250f74b Compare May 14, 2026 15:35

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 250f74bfc2

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

// `MAX_DRAIN_WAIT_SECS` (large enough for realistic
// overrides), and at least the global default if it's
// higher. 30s on top covers the other removal steps.
const MAX_DRAIN_WAIT_SECS: u64 = 600;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Derive removal wait from the actual drain window

Fresh evidence after the earlier timeout review is that this version still uses a fixed 600s floor here instead of the per-worker value. When a worker sets health.drain_settle_secs above 600 while the global setting remains lower, DrainWorkersStep sleeps the full per-worker value, but wait_for_completion returns Workflow timeout after 630s and the RemoveWorker job is reported failed even though the drain is valid and still in progress. The wait budget needs to be computed from the workers being removed, or the configured per-worker drain window must be capped to match this timeout.

Useful? React with 👍 / 👎.

@slin1237
slin1237 merged commit 271e424 into main May 14, 2026
36 of 38 checks passed
@slin1237
slin1237 deleted the feat/worker-draining-state branch May 14, 2026 15:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates model-gateway Model gateway crate changes protocols Protocols crate changes python-bindings Python bindings changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant