Skip to content

test(discovery): end-to-end reconcile tests against a scripted fake API server - #2062

Merged
slin1237 merged 1 commit into
mainfrom
test/discovery-fake-apiserver-it
Aug 6, 2026
Merged

slin1237 merged 1 commit into
mainfrom
test/discovery-fake-apiserver-it

Conversation

@slin1237

@slin1237 slin1237 commented Aug 5, 2026

Copy link
Copy Markdown
Member

Stacked on #2061.

Description

Adds model_gateway/tests/k8s_discovery_test.rs: the real watcher → reflector Store → reconciler → JobQueue → registration workflow → worker registry pipeline runs end to end; only two things are doubles — a scripted in-process pods endpoint (LIST + line-delimited WATCH, with 410-Gone expiry to force re-LISTs) and tests/common mock engine servers on ephemeral ports.

Scenarios:

  • Multi-port lifecycle — a pod with smg.ai/worker-ports: "p1,p2" registers one worker per engine, each stamped with the pod-uid ownership label; deleting the pod removes both.
  • Graceful termination — setting deletion_timestamp on a still-Ready pod removes its worker while the pod object still exists (drain at grace-period start).
  • Zombie cleanup / manual isolation — a pre-registered worker labeled with a nonexistent pod uid (the SMG Roadmap 2026 H2 #2031 @neverCase scenario: registration outliving its pod) is removed; an unlabeled manually-added worker survives every pass.
  • Desync convergence — a pod deleted while the watch is expired (410 "too old resource version") is only visible via re-LIST; the reconciler converges after reconnect.

The client-injection seam (start_service_discovery_with_client) is gated behind a new test-util cargo feature that a self dev-dependency enables for this crate's own test targets only — production builds compile no test surface (cargo build -p smg verified without the feature).

Test Plan

  • cargo test -p smg --test k8s_discovery_test: 13 passed, 0 failed (~5s)
  • cargo test -p smg: 1943 passed, 0 failed (22 binaries)
  • cargo clippy --all-targets -- -D warnings exit 0; cargo +nightly fmt --all --check clean
  • cargo build -p smg (no test-util) green — seam absent from production builds

@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • Reliability

    • Improved Kubernetes service discovery across pod additions, readiness changes, replacements, removals, and termination.
    • Service discovery now recovers from interrupted watch streams and refreshes discovered services.
    • Verified multi-port registration, annotation-based port updates, stale-service cleanup, worker preservation, registration retries, and configurable draining behavior.
    • Improved synchronization of router mesh-cluster state.
  • Tests

    • Added comprehensive integration coverage for service discovery lifecycle, recovery, and failure scenarios.

Walkthrough

The change adds test-only Kubernetes client injection and end-to-end service-discovery tests. The tests cover pod LIST/WATCH handling, worker registration, reconciliation, watch recovery, deletion draining, registration retry, and router mesh-cluster state updates.

Changes

Kubernetes service discovery testing

Layer / File(s) Summary
Test client injection and runner wiring
model_gateway/Cargo.toml, model_gateway/src/service_discovery.rs
Adds the test-util feature, enables it for test builds, and adds a client-injection entry point backed by a shared discovery runner.
Scripted Kubernetes API fixtures
model_gateway/tests/k8s_discovery_test.rs
Adds an in-memory API with pod LIST/WATCH responses, mutations, silent removals, reconnections, and expired-watch events.
Worker registration and reconciliation
model_gateway/tests/k8s_discovery_test.rs
Tests multi-port registration, readiness changes, pod replacement, annotation port changes, stale-worker cleanup, registration retry, and deletion draining.
Watch recovery and router state
model_gateway/tests/k8s_discovery_test.rs
Tests watch recovery, cleanup after silent pod removal, and router mesh-cluster state transitions.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant FakeKubernetesApi
  participant ServiceDiscovery
  participant WorkerRegistry
  participant RouterMeshState
  FakeKubernetesApi->>ServiceDiscovery: Serve pod LIST and WATCH responses
  ServiceDiscovery->>WorkerRegistry: Register or reconcile pod-owned workers
  FakeKubernetesApi->>ServiceDiscovery: Emit pod changes or watch failures
  ServiceDiscovery->>FakeKubernetesApi: Reconnect or re-LIST
  ServiceDiscovery->>WorkerRegistry: Drain or remove stale workers
  ServiceDiscovery->>RouterMeshState: Set mesh cluster Alive or Down
Loading

Possibly related PRs

  • smg-project/smg#2061: Covers the Kubernetes service-discovery implementation tested by this change.

Suggested reviewers: catherinesue, key4ng

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the end-to-end discovery reconciliation tests added by this pull request.
Description check ✅ Passed The description directly explains the new integration tests, covered scenarios, test seam, and validation results.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch test/discovery-fake-apiserver-it

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added dependencies Dependency updates tests Test changes model-gateway Model gateway crate changes labels Aug 5, 2026
@claude

claude Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

👋 The PR description doesn't fully follow
PULL_REQUEST_TEMPLATE.md:

  • Missing header: ### Problem
  • Missing header: ### Solution
  • Missing header: ## Changes

Please update the PR description so reviewers have the context they need.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thorough review complete — no issues found. Clean, well-structured PR:

  • Production refactoring: The run_service_discovery extraction is a clean mechanical split. The enabled guard stays in the public entry point, return types are correct, and ring initialization is properly placed.
  • Feature gating: test-util is only activated by the self dev-dependency in [dev-dependencies] — production builds compile no test surface.
  • Test quality: The fake K8s API server faithfully models LIST + streaming WATCH with 410-Gone expiry. All four scenarios (multi-port lifecycle, graceful termination, zombie cleanup, desync convergence) are well-motivated and cover the key reconciler behaviors. Proper reuse of existing common::mock_worker and common::create_test_context infrastructure.

0 🔴 Important · 0 🟡 Nit · 0 🟣 Pre-existing

@slin1237
slin1237 force-pushed the feat/discovery-informer-multiport branch from a1a1b44 to 2acb222 Compare August 6, 2026 02:48
@slin1237
slin1237 requested a review from gongwei-130 as a code owner August 6, 2026 02:48
@slin1237
slin1237 force-pushed the test/discovery-fake-apiserver-it branch from 9747100 to 1a898a3 Compare August 6, 2026 03:10
Base automatically changed from feat/discovery-informer-multiport to main August 6, 2026 04:05
@slin1237
slin1237 force-pushed the test/discovery-fake-apiserver-it branch from 1a898a3 to 763101b Compare August 6, 2026 04:12
Comment on lines +134 to +140
let _ = self
.state
.watch_tx
.lock()
.unwrap()
.send(format!("{status}\n"));
*self.state.watch_tx.lock().unwrap() = broadcast::channel(64).0;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: Two separate lock acquisitions on watch_tx create a window where a concurrent watch request could subscribe to the old channel (between the send and the replace), then immediately get a Closed error instead of the 410. Unlikely to cause a flake in practice since the kube watcher must process the 410 before reconnecting, but holding a single guard is trivially safer:

Suggested change
let _ = self
.state
.watch_tx
.lock()
.unwrap()
.send(format!("{status}\n"));
*self.state.watch_tx.lock().unwrap() = broadcast::channel(64).0;
let mut guard = self.state.watch_tx.lock().unwrap();
let _ = guard.send(format!("{status}\n"));
*guard = broadcast::channel(64).0;

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
model_gateway/src/service_discovery.rs (1)

466-478: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

🟡 Nit: Document the Tokio runtime requirement.

start_service_discovery_with_client is synchronous, but run_service_discovery calls task::spawn. If a caller invokes this function outside a Tokio runtime, task::spawn panics. Every current caller is a #[tokio::test] function, so the panic is not reachable today. Add a # Panics note so the requirement is visible at the call site.

♻️ Proposed doc addition
 /// Run discovery against an injected client. Compiled only for this crate's
 /// own integration tests, which point it at a scripted API server.
+///
+/// # Panics
+///
+/// Panics if called outside a Tokio runtime.
 #[cfg(feature = "test-util")]
 pub fn start_service_discovery_with_client(
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/src/service_discovery.rs` around lines 466 - 478, Add a Rustdoc
# Panics section to start_service_discovery_with_client documenting that it
panics when called outside an active Tokio runtime because run_service_discovery
uses task::spawn. Keep the existing function behavior unchanged.
model_gateway/tests/k8s_discovery_test.rs (1)

156-164: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

🟡 Nit: Make dropped watch events fail loudly.

The Lagged arm discards the number of skipped events and continues. If the broadcast buffer ever overflows, the watcher silently misses pod events. The test then fails with a wait_for timeout that gives no indication that the fixture dropped the events.

The current tests emit only a few events, so overflow is not reachable today. Surfacing the condition keeps future high-volume scenarios diagnosable.

♻️ Proposed change to surface dropped events
             loop {
                 match rx.recv().await {
                     Ok(line) => return Some((Ok::<Bytes, Infallible>(Bytes::from(line)), rx)),
-                    Err(broadcast::error::RecvError::Lagged(_)) => continue,
+                    Err(broadcast::error::RecvError::Lagged(skipped)) => {
+                        panic!("fake k8s watch stream dropped {skipped} events; increase the broadcast capacity");
+                    }
                     Err(broadcast::error::RecvError::Closed) => return None,
                 }
             }

As per coding guidelines: "Run the silent-failure-hunter agent on changed files to detect swallowed errors, inappropriate fallbacks, and missing error propagation."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/tests/k8s_discovery_test.rs` around lines 156 - 164, Update the
futures::stream::unfold receiver handling so broadcast::error::RecvError::Lagged
no longer silently continues; propagate or surface the skipped-event error
through the stream using the existing error type, while preserving normal Ok
line delivery and Closed termination behavior.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@model_gateway/src/service_discovery.rs`:
- Around line 466-478: Add a Rustdoc # Panics section to
start_service_discovery_with_client documenting that it panics when called
outside an active Tokio runtime because run_service_discovery uses task::spawn.
Keep the existing function behavior unchanged.

In `@model_gateway/tests/k8s_discovery_test.rs`:
- Around line 156-164: Update the futures::stream::unfold receiver handling so
broadcast::error::RecvError::Lagged no longer silently continues; propagate or
surface the skipped-event error through the stream using the existing error
type, while preserving normal Ok line delivery and Closed termination behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 2c508773-dc31-44ba-97d3-9a87865242d3

📥 Commits

Reviewing files that changed from the base of the PR and between 3c4e508 and 763101b.

📒 Files selected for processing (3)
  • model_gateway/Cargo.toml
  • model_gateway/src/service_discovery.rs
  • model_gateway/tests/k8s_discovery_test.rs

…PI server

Add an integration binary that runs the real watcher, reflector store,
reconciler, JobQueue, registration workflow, and worker registry against
a scripted in-process pods endpoint (LIST + line-delimited WATCH with
per-object resourceVersions, 410 expiry, and stream-sever support) and
mock engine servers.

Steady-state scenarios: multi-port pod registers one labeled worker per
engine and removes all on pod deletion; a terminating pod drains at
grace-period start while its object still exists; a zombie worker whose
pod is gone is cleaned up while a manually added worker is untouched;
live annotation edits grow and shrink the worker set; a router pod
drives mesh cluster state Alive/Down through the injected ClusterState.

Failure and recovery scenarios: a registration that keeps failing while
the engine port is closed leaves no phantom worker and succeeds once the
engine appears; a pod created after startup registers only when it turns
Ready (pure watch-event path); a same-URL pod replacement converges to
the new pod uid through the revision-guarded remove + re-register; a
silent deletion during a 410 desync converges through the re-LIST; a
severed watch stream reconnects and processes subsequent events; a 2s
drain window is honored across reconcile passes ticking every 300ms
(mutation-checked: disabling the reconciler's in-flight gate fails it).

The client-injection seam (start_service_discovery_with_client) is
compiled only under the new test-util feature, which a self
dev-dependency enables for this crate's own test targets — production
builds carry no test surface.

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@slin1237
slin1237 force-pushed the test/discovery-fake-apiserver-it branch from 763101b to 596fde7 Compare August 6, 2026 04:19
@coderabbitai

coderabbitai Bot commented Aug 6, 2026

Copy link
Copy Markdown

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (3)
model_gateway/tests/k8s_discovery_test.rs (3)

173-181: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

🟡 Nit: The watch stream drops lagged events silently.

RecvError::Lagged is skipped with continue. A lagged receiver means the fixture dropped scripted pod events, and the affected test then fails as an opaque 15-20s timeout instead of reporting the cause. Capacity is 64 and each test sends a few events, so lag is unlikely, but a diagnostic here is cheap.

Consider emitting the count, for example eprintln!("fake k8s watch lagged, dropped {n} events"), before you continue.

As per coding guidelines: "Run the silent-failure-hunter agent on changed files to detect swallowed errors, inappropriate fallbacks, and missing error propagation."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/tests/k8s_discovery_test.rs` around lines 173 - 181, Update the
watch stream’s RecvError::Lagged branch inside futures::stream::unfold to emit a
diagnostic containing the number of dropped events before continuing; keep the
existing retry behavior and Closed handling unchanged.

Source: Coding guidelines


136-158: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

🟡 Nit: expire_watch can drop the 410 event if no watch is subscribed yet.

sever_watch_and_await_reconnect waits for receiver_count() > 0 before it returns. expire_watch has no equivalent guard. If the reflector has finished the initial LIST but has not yet issued the WATCH request, the 410 line goes to a sender with zero receivers and is lost. The subsequent sender swap then leaves the watcher on a stream that never re-LISTs, and watch_interruption_relists_and_converges fails only after its 20s timeout with an unclear message.

The window is small because the test first waits for registration, so this is a flake-hardening suggestion, not a current failure.

♻️ Wait for a subscriber before sending the expiry event
-    fn expire_watch(&self) {
+    async fn expire_watch(&self) {
+        let deadline = std::time::Instant::now() + Duration::from_secs(10);
+        while self.state.watch_tx.lock().unwrap().receiver_count() == 0 {
+            assert!(
+                std::time::Instant::now() < deadline,
+                "no watch stream to expire"
+            );
+            tokio::time::sleep(Duration::from_millis(25)).await;
+        }
         let status = serde_json::json!({

Update the single call site at line 461 to fake.expire_watch().await;.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/tests/k8s_discovery_test.rs` around lines 136 - 158, Make
expire_watch asynchronous and wait until the watch_tx sender has at least one
receiver before sending the 410 event; then replace the call site in
watch_interruption_relists_and_converges with an awaited invocation. Preserve
the existing event send and sender replacement behavior after the subscriber
guard.

588-591: 🩺 Stability & Availability | 🔵 Trivial | 💤 Low value

🟡 Nit: The reserve-then-release port pattern can race with concurrent tests.

The test binds an ephemeral port, releases it, and rebinds it more than four seconds later at line 618. Cargo runs the tests in this binary concurrently, and start_mock_engine also binds port 0. If the kernel assigns the released port to another test in that window, two failures become possible: the negative assertion at line 605 sees a healthy foreign listener, or engine.start().unwrap() at line 618 panics with an address-in-use error.

The kernel normally cycles the ephemeral range before reuse, so this is unlikely. Consider starting MockWorker with port: 0 first, recording its port, stopping it, and reusing that port — or accept the risk and document it.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@model_gateway/tests/k8s_discovery_test.rs` around lines 588 - 591, Eliminate
the reserve-then-release port race in the test around MockWorker and
start_mock_engine by obtain­ing an assigned port from a started MockWorker
configured with port 0, then stop it and reuse that recorded port for the later
assertions and engine startup. Preserve the test’s intended free-port and
negative-listener checks without relying on an unprotected four-second gap.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@model_gateway/tests/k8s_discovery_test.rs`:
- Around line 173-181: Update the watch stream’s RecvError::Lagged branch inside
futures::stream::unfold to emit a diagnostic containing the number of dropped
events before continuing; keep the existing retry behavior and Closed handling
unchanged.
- Around line 136-158: Make expire_watch asynchronous and wait until the
watch_tx sender has at least one receiver before sending the 410 event; then
replace the call site in watch_interruption_relists_and_converges with an
awaited invocation. Preserve the existing event send and sender replacement
behavior after the subscriber guard.
- Around line 588-591: Eliminate the reserve-then-release port race in the test
around MockWorker and start_mock_engine by obtain­ing an assigned port from a
started MockWorker configured with port 0, then stop it and reuse that recorded
port for the later assertions and engine startup. Preserve the test’s intended
free-port and negative-listener checks without relying on an unprotected
four-second gap.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 1503a637-49da-4712-85f6-cc489dbaceca

📥 Commits

Reviewing files that changed from the base of the PR and between 3c4e508 and 596fde7.

📒 Files selected for processing (3)
  • model_gateway/Cargo.toml
  • model_gateway/src/service_discovery.rs
  • model_gateway/tests/k8s_discovery_test.rs
🚧 Files skipped from review as they are similar to previous changes (2)
  • model_gateway/Cargo.toml
  • model_gateway/src/service_discovery.rs

@slin1237
slin1237 merged commit 711ee70 into main Aug 6, 2026
11 of 13 checks passed
@slin1237
slin1237 deleted the test/discovery-fake-apiserver-it branch August 6, 2026 04:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates model-gateway Model gateway crate changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant