Skip to content

feat(gateway): add --remove-unhealthy-workers - #714

Merged
slin1237 merged 3 commits into
smg-project:mainfrom
ekzhang:ekzhang/remove-unhealthy
Mar 14, 2026
Merged

slin1237 merged 3 commits into
smg-project:mainfrom
ekzhang:ekzhang/remove-unhealthy

Conversation

@ekzhang

@ekzhang ekzhang commented Mar 10, 2026 •

Copy link
Copy Markdown
Contributor

Description

This option automatically removes unhealthy workers from the gateway registry after failing the configured health check threshold.

Especially useful for some setups in inference gateway mode, and you periodically re-register workers.

cc @slin1237

Problem

Solution

Changes

CLI option was added

Note: I updated the CLI options and documentation as well, it's documented in two places right now

Test Plan

Tests in worker_registry.rs

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • New Features

    • Added --remove-unhealthy-workers option (default: disabled) to remove workers when marked unhealthy and propagate setting to health checks.
  • Documentation

    • Updated health checks and configuration reference to document the new option.
  • Tests

    • Added tests validating removal behavior when enabled and preservation when disabled.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@github-actions github-actions Bot added documentation Improvements or additions to documentation python-bindings Python bindings changes model-gateway Model gateway crate changes labels Mar 10, 2026
@coderabbitai

coderabbitai Bot commented Mar 10, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

A new boolean configuration, remove_unhealthy_workers, was added and propagated through CLI, config types, Python bindings, server startup, and the worker registry health-checker so unhealthy workers can be removed automatically when enabled. Documentation was updated to document the option.

Changes

Cohort / File(s) Summary
Config types & CLI
model_gateway/src/config/types.rs, model_gateway/src/main.rs
Added remove_unhealthy_workers: bool to HealthCheckConfig (serde default false) and to CLI args; default false.
Worker registry & health-check logic
model_gateway/src/core/worker_registry.rs
Added shallow_clone(); extended start_health_checker(..., remove_unhealthy: bool) to optionally remove unhealthy workers during checks; added tests for both behaviors.
Server startup
model_gateway/src/server.rs
Passes remove_unhealthy_workers flag into start_health_checker() call at startup.
Python bindings & CLI args (bindings)
bindings/python/src/lib.rs, bindings/python/src/smg/router_args.py
Added remove_unhealthy_workers field to Python Router and RouterArgs; PyO3 constructor and CLI flag --remove-unhealthy-workers exposed.
Documentation
docs/concepts/reliability/health-checks.md, docs/reference/configuration.md
Documented new --remove-unhealthy-workers option and its default (false).

Sequence Diagram

sequenceDiagram
    participant Server
    participant HealthChecker
    participant WorkerRegistry
    participant Storage

    Server->>HealthChecker: start_health_checker(interval, remove_unhealthy=true)
    HealthChecker->>WorkerRegistry: shallow_clone()  (if removal enabled)

    loop periodic
        HealthChecker->>WorkerRegistry: perform health checks
        WorkerRegistry-->>HealthChecker: report unhealthy workers
        alt remove_unhealthy == true
            HealthChecker->>WorkerRegistry: remove(unhealthy_worker)
            WorkerRegistry->>Storage: persist removal / sync state
            WorkerRegistry->>HealthChecker: confirm removal
        else remove_unhealthy == false
            HealthChecker->>WorkerRegistry: mark unhealthy (no removal)
        end
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • key4ng

Poem

🐰 I hopped through configs, flags in tow,

remove_unhealthy_workers now helps things go,
A sniff, a poke — if a worker's unwell,
I nudge it out gently, no tale to tell,
Registry hops tidy, on with the show!

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main feature added: a new CLI flag --remove-unhealthy-workers for the gateway.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the gateway's worker health check mechanism by providing an option to automatically deregister workers that are deemed unhealthy. This feature improves the reliability and self-healing capabilities of the gateway, ensuring that only responsive and functional workers are part of the routing pool, which is especially beneficial for setups with frequently changing or transient worker instances.

Highlights

  • New CLI Option: Introduced a new CLI option, --remove-unhealthy-workers, allowing the gateway to automatically remove workers from its registry if they consistently fail health checks.
  • Automated Worker Management: The health checker now supports an optional mode to actively remove unhealthy workers, which is particularly useful for dynamic environments like inference gateways where workers are frequently re-registered.
  • Configuration and Documentation: The new functionality is fully integrated into the gateway's configuration system and documented in the relevant health-checks.md and configuration.md files.
Changelog
  • bindings/python/src/lib.rs
    • Added remove_unhealthy_workers field to the Router struct.
    • Integrated remove_unhealthy_workers into the Router's configuration and builder methods.
  • docs/concepts/reliability/health-checks.md
    • Updated the health checks documentation to include the new --remove-unhealthy-workers CLI option.
  • docs/reference/configuration.md
    • Added the --remove-unhealthy-workers option to the main configuration reference documentation.
  • model_gateway/src/config/types.rs
    • Introduced remove_unhealthy_workers field to the HealthCheckConfig struct, defaulting to false.
  • model_gateway/src/core/worker_registry.rs
    • Implemented a shallow_clone method for WorkerRegistry to create cheap handles.
    • Modified start_health_checker to accept a remove_unhealthy boolean parameter.
    • Added logic within the health checker task to remove workers from the registry if they become unhealthy and the remove_unhealthy flag is set.
    • Included new unit tests to verify the functionality of removing and retaining unhealthy workers based on the configuration.
  • model_gateway/src/main.rs
    • Added --remove-unhealthy-workers as a new command-line argument for the gateway.
    • Mapped the new CLI argument to the HealthCheckConfig during server initialization.
  • model_gateway/src/server.rs
    • Passed the remove_unhealthy_workers configuration from the router config to the start_health_checker function.
Activity
  • A new CLI option --remove-unhealthy-workers was introduced to control the removal of unhealthy workers.
  • The pull request includes updates to the CLI options and documentation in two separate files.
  • Tests for the new worker removal logic were added to worker_registry.rs to ensure correct behavior.
  • The code passed cargo +nightly fmt and cargo clippy checks, indicating adherence to code style and best practices.
  • Documentation was explicitly updated as part of the changes.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@mergify

mergify Bot commented Mar 10, 2026

Copy link
Copy Markdown
Contributor

Hi @ekzhang, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

@ekzhang
ekzhang force-pushed the ekzhang/remove-unhealthy branch from 73140eb to 91aa03d Compare March 10, 2026 23:05

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new --remove-unhealthy-workers option that automatically removes workers from the registry after they fail health checks. However, the current implementation poses a significant Denial of Service (DoS) risk as it can lead to the removal of all workers during transient failures, lacking built-in recovery or safeguards to maintain a minimum pool of workers. Additionally, while the implementation is well-structured and consistently applied, there is a suggestion to improve logging by handling the Result from the health check function.

Comment thread model_gateway/src/core/worker_registry.rs
Comment thread model_gateway/src/core/worker_registry.rs
coderabbitai[bot]
coderabbitai Bot previously requested changes Mar 10, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
model_gateway/src/core/worker_registry.rs (1)

746-766: ⚠️ Potential issue | 🔴 Critical

Remove unhealthy workers by snapshot identity, not by URL.

This races with same-URL re-registration. remove_by_url() clears url_to_id before it removes the worker, so a replacement that registers between the check and the delete can be removed or lose its URL mapping. That breaks the exact IGW re-registration flow this flag is meant for. Carry (WorkerId, Arc<dyn Worker>) through the health-check future, then remove only if the current registry entry is still the same Arc. Please also add a regression test for same-URL replacement during an in-flight health check.

Possible fix sketch
-                let workers: Vec<Arc<dyn Worker>> = workers_ref
+                let workers: Vec<(WorkerId, Arc<dyn Worker>)> = workers_ref
                     .iter()
-                    .map(|entry| entry.value().clone())
+                    .map(|entry| (entry.key().clone(), entry.value().clone()))
                     .collect();

                 // Collect workers whose deadline has passed
                 let due_workers: Vec<_> = workers
                     .iter()
-                    .filter(|w| !w.metadata().health_config.disable_health_check)
+                    .filter(|(_, w)| !w.metadata().health_config.disable_health_check)
                     .filter(|w| {
+                        let (_, worker) = w;
                         next_check
-                            .get(w.url())
+                            .get(worker.url())
                             .is_some_and(|deadline| now >= *deadline)
                     })
                     .cloned()
                     .collect();

                 // Run due health checks in parallel and schedule the next deadline
                 if !due_workers.is_empty() {
                     for worker in &due_workers {
+                        let (_, worker) = worker;
                         let secs = worker.metadata().health_config.check_interval_secs;
                         let secs = if secs > 0 {
                             secs
@@
                     let futs: Vec<_> = due_workers
                         .into_iter()
-                        .map(|w| async move {
+                        .map(|(worker_id, w)| async move {
                             let _ = w.check_health_async().await;
-                            w
+                            (worker_id, w)
                         })
                         .collect();
                     let checked_workers = futures::future::join_all(futs).await;

                     // Remove workers that transitioned to unhealthy
                     if let Some(ref registry) = registry {
-                        for worker in &checked_workers {
-                            if !worker.is_healthy() {
+                        for (worker_id, worker) in &checked_workers {
+                            if !worker.is_healthy()
+                                && registry
+                                    .get(worker_id)
+                                    .is_some_and(|current| Arc::ptr_eq(&current, worker))
+                            {
                                 let url = worker.url().to_string();
                                 tracing::warn!(
                                     worker_url = %url,
                                     "Removing unhealthy worker from registry"
                                 );
                                 next_check.remove(&url);
-                                registry.remove_by_url(&url);
+                                registry.remove(worker_id);
                             }
                         }
                     }
                 }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/core/worker_registry.rs` around lines 746 - 766, The
health-check currently maps over due_workers and only carries Arc<dyn Worker>,
then calls registry.remove_by_url(&url) which races with a same-URL
re-registration and can remove a new worker; change the future mapping to carry
(WorkerId, Arc<dyn Worker>) through (e.g. map each w -> (w.id(), w) or similar)
and after awaiting checked_workers, when a worker is unhealthy verify the
registry still points to the exact same Arc (by fetching the current entry by
URL or by ID and comparing pointer/ID equality) before removing; use
registry.remove_by_id/remove_by_identity (or only call remove_by_url if the
registry lookup confirms the stored Arc matches the captured Arc) and update
next_check removal to operate on the WorkerId snapshot rather than
unconditionally removing by URL; add a regression test that registers a
replacement worker for the same URL mid-health-check to assert the new worker is
not removed.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@bindings/python/src/lib.rs`:
- Line 782: The Python high-level API never exposes the Rust flag
remove_unhealthy_workers because RouterArgs (in the Python wrapper) lacks that
field and the router wrapper doesn't forward it; add a boolean field
remove_unhealthy_workers (default False) to RouterArgs and ensure the router
wrapper (the function that builds kwargs from RouterArgs) includes that key so
it gets passed through to the Rust binding (keep naming consistent with the Rust
binding and update any dict/serialization code that builds kwargs for the Rust
call).

In `@docs/concepts/reliability/health-checks.md`:
- Line 112: The docs entry for the flag `--remove-unhealthy-workers` is missing
that it disables automatic rejoin/self-healing; update the table row and the
duplicate reference entry to explicitly state that when
`--remove-unhealthy-workers` is true the worker is removed from the registry and
cannot be marked healthy again by later health checks (it will only return if it
re-registers), and ensure the wording references the service's "Self-Healing"
behavior to avoid contradiction.

---

Outside diff comments:
In `@model_gateway/src/core/worker_registry.rs`:
- Around line 746-766: The health-check currently maps over due_workers and only
carries Arc<dyn Worker>, then calls registry.remove_by_url(&url) which races
with a same-URL re-registration and can remove a new worker; change the future
mapping to carry (WorkerId, Arc<dyn Worker>) through (e.g. map each w ->
(w.id(), w) or similar) and after awaiting checked_workers, when a worker is
unhealthy verify the registry still points to the exact same Arc (by fetching
the current entry by URL or by ID and comparing pointer/ID equality) before
removing; use registry.remove_by_id/remove_by_identity (or only call
remove_by_url if the registry lookup confirms the stored Arc matches the
captured Arc) and update next_check removal to operate on the WorkerId snapshot
rather than unconditionally removing by URL; add a regression test that
registers a replacement worker for the same URL mid-health-check to assert the
new worker is not removed.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: d3f6d95c-ace0-4910-9c66-ca030005eb6a

📥 Commits

Reviewing files that changed from the base of the PR and between 5bc32eb and 91aa03d.

📒 Files selected for processing (7)
  • bindings/python/src/lib.rs
  • docs/concepts/reliability/health-checks.md
  • docs/reference/configuration.md
  • model_gateway/src/config/types.rs
  • model_gateway/src/core/worker_registry.rs
  • model_gateway/src/main.rs
  • model_gateway/src/server.rs

Comment thread bindings/python/src/lib.rs
Comment thread docs/concepts/reliability/health-checks.md
@ekzhang
ekzhang force-pushed the ekzhang/remove-unhealthy branch from cb2d5a6 to 034d6a1 Compare March 10, 2026 23:26
@ekzhang

ekzhang commented Mar 11, 2026 •

Copy link
Copy Markdown
Contributor Author

Hey @slin1237, let me know if you have any feedback on this PR

@slin1237

Copy link
Copy Markdown
Member

@ekzhang
PR LGTM from the get go. just some minor format issue
Can you address it, then we can merge this one

@ekzhang

ekzhang commented Mar 13, 2026

Copy link
Copy Markdown
Contributor Author

Hi @slin1237, I double-checked that format passes

➜  smg git:(main) cargo +nightly fmt --check && echo "passed"
passed

ekzhang added 3 commits March 14, 2026 16:10
This option automatically removes unhealthy workers from the gateway registry after failing the configured health check threshold.

Especially useful for some setups in inference gateway mode, and you periodically re-register workers.

cc @slin1237

Signed-off-by: Eric Zhang <ekzhang1@gmail.com>
Signed-off-by: Eric Zhang <ekzhang1@gmail.com>
Signed-off-by: Eric Zhang <ekzhang1@gmail.com>
@slin1237
slin1237 force-pushed the ekzhang/remove-unhealthy branch from 034d6a1 to 4a8586e Compare March 14, 2026 23:11
@slin1237
slin1237 requested a review from gongwei-130 as a code owner March 14, 2026 23:11

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
model_gateway/src/core/worker_registry.rs (1)

703-730: ⚠️ Potential issue | 🔴 Critical

Make unhealthy-worker removal identity-safe.

register() reuses the existing WorkerId for a URL and overwrites workers[worker_id]. If a worker re-registers while its old Arc is still inside check_health_async(), remove_by_url() here will delete the fresh registration instead of the stale unhealthy instance. That race hits the exact “periodically re-registered workers” deployment this flag is meant for. Carry the WorkerId through the snapshot and only remove when the registry still points to the same Arc that was checked.

🐛 Proposed fix
-                let workers: Vec<Arc<dyn Worker>> = workers_ref
-                    .iter()
-                    .map(|entry| entry.value().clone())
-                    .collect();
+                let workers: Vec<(WorkerId, Arc<dyn Worker>)> = workers_ref
+                    .iter()
+                    .map(|entry| (entry.key().clone(), entry.value().clone()))
+                    .collect();

                 // Sync schedule with registry: add new workers, prune removed
                 // and disabled ones so stale deadlines don't cause wakeups.
                 let checkable_urls: std::collections::HashSet<String> = workers
                     .iter()
-                    .filter(|w| !w.metadata().health_config.disable_health_check)
-                    .map(|w| w.url().to_string())
+                    .filter(|(_, w)| !w.metadata().health_config.disable_health_check)
+                    .map(|(_, w)| w.url().to_string())
                     .collect();
                 next_check.retain(|url, _| checkable_urls.contains(url));
                 for url in &checkable_urls {
                     next_check.entry(url.clone()).or_insert(now);
                 }

                 // Collect workers whose deadline has passed
                 let due_workers: Vec<_> = workers
                     .iter()
-                    .filter(|w| !w.metadata().health_config.disable_health_check)
-                    .filter(|w| {
+                    .filter(|(_, w)| !w.metadata().health_config.disable_health_check)
+                    .filter(|(_, w)| {
                         next_check
                             .get(w.url())
                             .is_some_and(|deadline| now >= *deadline)
                     })
                     .cloned()
                     .collect();

                 // Run due health checks in parallel and schedule the next deadline
                 if !due_workers.is_empty() {
-                    for worker in &due_workers {
+                    for (_, worker) in &due_workers {
                         let secs = worker.metadata().health_config.check_interval_secs;
                         let secs = if secs > 0 {
                             secs
                         } else {
                             default_interval_secs
@@
                     }
                     let futs: Vec<_> = due_workers
                         .into_iter()
-                        .map(|w| async move {
+                        .map(|(worker_id, w)| async move {
                             let _ = w.check_health_async().await;
-                            w
+                            (worker_id, w)
                         })
                         .collect();
                     let checked_workers = futures::future::join_all(futs).await;

                     // Remove workers that transitioned to unhealthy
                     if let Some(ref registry) = registry {
-                        for worker in &checked_workers {
+                        for (worker_id, worker) in &checked_workers {
                             if !worker.is_healthy() {
                                 let url = worker.url().to_string();
+                                let Some(current) = registry.get(worker_id) else {
+                                    continue;
+                                };
+                                if !Arc::ptr_eq(&current, worker) {
+                                    tracing::debug!(
+                                        worker_url = %url,
+                                        "Skipping removal because worker was replaced during health check"
+                                    );
+                                    continue;
+                                }
                                 tracing::warn!(
                                     worker_url = %url,
                                     "Removing unhealthy worker from registry"
                                 );
                                 next_check.remove(&url);
-                                registry.remove_by_url(&url);
+                                registry.remove(worker_id);
                             }
                         }
                     }
                 }

Also applies to: 746-766

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@model_gateway/src/core/worker_registry.rs` around lines 703 - 730, The
snapshot currently clones only worker Arcs and uses url-based removal which
races with re-registration; instead capture and carry the WorkerId alongside
each Arc in the snapshot (e.g. produce Vec<(WorkerId, Arc<dyn Worker>)>) in the
check_health_async() snapshot/collection logic used around the
next_check/due_workers code and the similar block at the other site, and when a
worker is found unhealthy verify the registry still points to the same Arc
before removing: either use a remove_by_id or fetch registry entry by WorkerId
and compare with Arc::ptr_eq to the snapshot Arc, only then call
remove_by_url/remove_by_id. This ensures re-registered workers with the same
WorkerId are not incorrectly removed.
♻️ Duplicate comments (1)
docs/concepts/reliability/health-checks.md (1)

112-112: ⚠️ Potential issue | 🟡 Minor

Document that this opts out of self-healing.

With this flag enabled the worker is removed from the registry, so later health checks cannot mark it healthy again; it only returns if something re-registers it. Please spell that out here, and mirror the same wording in the duplicate reference / CLI descriptions, so this page does not contradict its own “Self-Healing” section.

📝 Suggested wording
-| `--remove-unhealthy-workers` | `false` | Remove workers after being marked unhealthy |
+| `--remove-unhealthy-workers` | `false` | Remove workers from the registry after they are marked unhealthy. Use this only when workers are periodically re-registered, because removed workers do not automatically rejoin through later health checks |
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@docs/concepts/reliability/health-checks.md` at line 112, Clarify that the
`--remove-unhealthy-workers` flag opts out of self-healing by explicitly stating
that when enabled the worker is removed from the registry and cannot be marked
healthy again by subsequent health checks (it will only return if
re-registered); update the sentence for `--remove-unhealthy-workers` to include
this wording and then mirror that exact phrasing in the duplicate reference and
any CLI description strings to ensure consistency with the “Self-Healing”
section and avoid contradiction.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@model_gateway/src/core/worker_registry.rs`:
- Around line 703-730: The snapshot currently clones only worker Arcs and uses
url-based removal which races with re-registration; instead capture and carry
the WorkerId alongside each Arc in the snapshot (e.g. produce Vec<(WorkerId,
Arc<dyn Worker>)>) in the check_health_async() snapshot/collection logic used
around the next_check/due_workers code and the similar block at the other site,
and when a worker is found unhealthy verify the registry still points to the
same Arc before removing: either use a remove_by_id or fetch registry entry by
WorkerId and compare with Arc::ptr_eq to the snapshot Arc, only then call
remove_by_url/remove_by_id. This ensures re-registered workers with the same
WorkerId are not incorrectly removed.

---

Duplicate comments:
In `@docs/concepts/reliability/health-checks.md`:
- Line 112: Clarify that the `--remove-unhealthy-workers` flag opts out of
self-healing by explicitly stating that when enabled the worker is removed from
the registry and cannot be marked healthy again by subsequent health checks (it
will only return if re-registered); update the sentence for
`--remove-unhealthy-workers` to include this wording and then mirror that exact
phrasing in the duplicate reference and any CLI description strings to ensure
consistency with the “Self-Healing” section and avoid contradiction.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: d1fd0da5-a0d4-4972-9898-e55407911414

📥 Commits

Reviewing files that changed from the base of the PR and between 91aa03d and 4a8586e.

📒 Files selected for processing (8)
  • bindings/python/src/lib.rs
  • bindings/python/src/smg/router_args.py
  • docs/concepts/reliability/health-checks.md
  • docs/reference/configuration.md
  • model_gateway/src/config/types.rs
  • model_gateway/src/core/worker_registry.rs
  • model_gateway/src/main.rs
  • model_gateway/src/server.rs

@slin1237
slin1237 merged commit 6d1b268 into smg-project:main Mar 14, 2026
39 of 52 checks passed
smfirmin pushed a commit to smfirmin/smg that referenced this pull request Mar 20, 2026
Signed-off-by: Eric Zhang <ekzhang1@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation model-gateway Model gateway crate changes python-bindings Python bindings changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants