Skip to content

perf(smg): mitigate the worker-sync lag-resync burst - #1667

Merged
CatherineSue merged 4 commits into
mainfrom
chang/mesh-resync-burst
Jun 13, 2026
Merged

CatherineSue merged 4 commits into
mainfrom
chang/mesh-resync-burst

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Jun 11, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

Two related burst amplifiers in the worker mesh sync, found by the #1661 perf review: (1) the registry's WorkerEvent broadcast had 64 slots — bulk startup registration or a probe storm lags it, and every lag forces a full mesh resync; (2) that resync (resync_local) re-put every local worker unconditionally. Each re-put mints a fresh Lamport op, which invalidates every peer's per-key send watermark — so one lag event burst W ops to N peers (~16× amplification at 1,000 workers) even when nothing had changed, and a sustained status storm could re-trigger it repeatedly.

Solution

  • Channel sized for fleets: 64 → 1024 slots (~98 KB fixed; verified the per-slot math, and tokio clears a consumed slot's payload on last receive, so retention is transient).
  • Resync skips unchanged workers: store_matches consults the store before re-putting. Cheap scalar fields gate the comparison; load is not compared (volatile, unread by importers); specs compare as JSON values rather than bytes, since WorkerSpec holds maps and two encodings of an identical spec can differ byte-wise. A skipped put is a watermark delta that never ships.

Correctness containment: a lost publish or foreign tombstone leaves the store without (or with different) state, so store_matches returns false and the resync still re-asserts — the lag-recovery role is preserved. False negatives merely re-put (the old behavior).

Changes

  • model_gateway/src/worker/registry.rs — event channel capacity + sizing rationale
  • model_gateway/src/mesh/adapters/worker_sync.rs — store_matches + resync gating; test via the local-subscriber observable (put ⇒ notify, skip ⇒ empty channel)

Test Plan

  • cargo test -p smg — suite green; new resync_skips_republish_when_state_unchanged asserts both directions (unchanged ⇒ skipped; health flip ⇒ republished).
  • cargo clippy --all-targets --all-features -- -D warnings — clean.
  • Adversarially reviewed (correctness + memory/perf lenses): zero confirmed must-fix; the comparison-ordering suggestion is folded in.
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes

    • Optimized worker state synchronization to avoid re-publishing unchanged local worker state and reduce unnecessary updates.
  • Performance Improvements

    • Increased internal event queue capacity to better handle fleet-scale bursts of worker events.
  • Tests

    • Added coverage to ensure resync skips redundant publishes and emits updates when state actually changes.

The 64-slot broadcast lagged under bulk startup registration and probe
storms; every lag forces a full mesh resync. 1024 slots covers
realistic worker counts at ~100 KB fixed cost.

Signed-off-by: Chang Su <8605658+CatherineSue@users.noreply.github.com>
resync_local re-put every local worker on broadcast lag; each re-put
mints a fresh Lamport op that invalidates every peer's send watermark
for the key, bursting W ops to N peers even when nothing changed. The
store is now consulted first: volatile load is ignored and specs
compare as JSON values (WorkerSpec holds maps, so equal specs can
differ byte-wise).

Signed-off-by: Chang Su <8605658+CatherineSue@users.noreply.github.com>
Changed workers now skip the two JSON parses, and the WorkerState clone
is gone — field-wise comparison replaces the zero-and-compare.

Signed-off-by: Chang Su <8605658+CatherineSue@users.noreply.github.com>
@CatherineSue
CatherineSue requested a review from slin1237 as a code owner June 11, 2026 17:31
@coderabbitai

coderabbitai Bot commented Jun 11, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 28458e23-8ea4-4cc2-9df3-74d0536712ee

📥 Commits

Reviewing files that changed from the base of the PR and between 2eb4ea5 and 857b747.

📒 Files selected for processing (1)
  • model_gateway/src/mesh/adapters/worker_sync.rs

📝 Walkthrough

Walkthrough

WorkerSyncAdapter's local resync now avoids republishing worker state when the mesh store already holds an equivalent WorkerState. A store_matches helper compares stored and in-memory states via identity fields and JSON-semantic spec equality. Event broadcast capacity increased from 64 to 1024 to handle fleet-scale bursts.

Changes

Worker State Management

Layer / File(s) Summary
Resync store-matching optimization
model_gateway/src/mesh/adapters/worker_sync.rs
Added store_matches helper to compare stored WorkerState by identity/health/version and JSON-semantic spec equality. Updated resync_local to conditionally republish only when state differs, while maintaining tombstoning of stale keys. Tests updated and new async test resync_skips_republish_when_state_unchanged verifies unchanged resyncs emit no CRDT writes and status flips still trigger republish.
Event buffer capacity increase
model_gateway/src/worker/registry.rs
Broadcast channel capacity for WorkerEvent increased from 64 to 1024 in WorkerRegistry::new() and documentation/comments updated accordingly.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • lightseekorg/smg#1313: Both PRs touch the same model_gateway/src/mesh/adapters/worker_sync.rs implementation of WorkerSyncAdapter, with foundational CRDT namespace plumbing related to this change.
  • lightseekorg/smg#1661: Related earlier work on worker_sync resync logic and mesh worker sync plumbing used by this optimization.
  • lightseekorg/smg#912: Changes mesh worker-state handling (spec, health) that intersect with the stored-vs-in-memory equality checks added here.

Suggested labels

mesh, tests

Suggested reviewers

  • slin1237
  • claude

Poem

🐰 I nibble bytes and prune the noise,
I only shout when state's a change,
Specs parsed gently, idle writes avoid,
Buffers grown to ride the range. 🥕

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'perf(smg): mitigate the worker-sync lag-resync burst' directly and specifically summarizes the main change: optimizing worker synchronization to reduce lag and resync traffic amplification.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch chang/mesh-resync-burst

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added the model-gateway Model gateway crate changes label Jun 11, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request optimizes the local worker resync process by skipping republishing when the stored state is already equivalent, and increases the worker registry's event broadcast channel capacity to 1024 to handle fleet-scale bursts. The review feedback suggests two key performance improvements: deferring the expensive worker_state_of serialization until after checking if the store matches, and refactoring store_matches to compare the worker's version/revision directly instead of performing expensive JSON deserialization and comparison of the specs.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +298 to +301
let state = worker_state_of(&id, &worker);
if !self.store_matches(&id, &state) {
self.on_worker_changed(id.as_str(), &state);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Defer the expensive worker_state_of call (which serializes the WorkerSpec to JSON) until after we have determined that the store does not match. This avoids unnecessary JSON serialization for all unchanged workers during a resync.

Suggested change
let state = worker_state_of(&id, &worker);
if !self.store_matches(&id, &state) {
self.on_worker_changed(id.as_str(), &state);
}
if !self.store_matches(&id, &worker) {
let state = worker_state_of(&id, &worker);
self.on_worker_changed(id.as_str(), &state);
}
References
  1. Avoid heap allocations (like Vec<u8>) within performance-critical comparison logic, especially during bulk processing like log compaction. Use allocation-free helpers or compare decoded values directly to reduce memory pressure.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deferring worker_state_of only pays if store_matches can decide "unchanged" without the spec bytes — i.e. under the version-only comparison in the sibling thread, which is unsound across restarts (revision() is an in-process counter that resets per boot). With semantic spec comparison kept, the common case (everything matches) is exactly the case that must compare specs, so building the state up front isn't wasted work in the steady state.

The cost worth cutting was the two JSON parses per matched worker: 857b747 short-circuits on spec byte equality, which holds in-process since re-serialising the same spec instance is byte-stable. The parse now only runs when bytes drift while scalars match (effectively once, after a restart).

Comment on lines +318 to +340
fn store_matches(&self, id: &WorkerId, state: &WorkerState) -> bool {
let Some(bytes) = self.workers.get(&format!("{PREFIX}{}", id.as_str())) else {
return false;
};
let Ok(stored) = bincode::deserialize::<WorkerState>(&bytes) else {
return false;
};
if stored.worker_id != state.worker_id
|| stored.model_id != state.model_id
|| stored.url != state.url
|| stored.health != state.health
|| stored.version != state.version
{
return false;
}
match (
serde_json::from_slice::<serde_json::Value>(&stored.spec),
serde_json::from_slice::<serde_json::Value>(&state.spec),
) {
(Ok(a), Ok(b)) => a == b,
_ => stored.spec == state.spec,
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Since WorkerSpec is immutable for a given BasicWorker instance, and any update to a worker's specification (via replace or register_or_replace) always increments its monotonic version (revision) counter, comparing the version field is entirely sufficient to guarantee that the underlying spec has not changed.

By changing store_matches to take &Arc<dyn Worker> instead of &WorkerState, we can perform all necessary comparisons using cheap scalar fields and completely eliminate the expensive, CPU-bound JSON deserialization and semantic comparison of stored.spec and state.spec.

    fn store_matches(&self, id: &WorkerId, worker: &Arc<dyn Worker>) -> bool {
        let Some(bytes) = self.workers.get(&format!("{PREFIX}{}", id.as_str())) else {
            return false;
        };
        let Ok(stored) = bincode::deserialize::<WorkerState>(&bytes) else {
            return false;
        };
        stored.worker_id == id.as_str()
            && stored.model_id == worker.model_id()
            && stored.url == worker.url()
            && stored.health == worker.is_healthy()
            && stored.version == worker.revision()
    }
References
  1. Avoid heap allocations (like Vec<u8>) within performance-critical comparison logic, especially during bulk processing like log compaction. Use allocation-free helpers or compare decoded values directly to reduce memory pressure.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The version-only comparison is unsound across restarts: revision() is an in-process atomic that resets to 0 when the gateway restarts, and with deterministic ids (#1668) a restarted node re-registers the same worker_id. If the spec changed across the restart (config edit) while the republished revision number matches what the store holds from the prior incarnation, version-equality would false-match and suppress the re-assert — exactly the prior-incarnation shadow the reconcile pass exists to fix. So the semantic spec comparison stays as ground truth.

Adopted the spirit of this in 857b747: a byte-equality fast path now gates the JSON parses. Re-serialising the same in-process spec instance is byte-stable, so steady-state resync/reconcile short-circuits on memcmp and only falls through to the JSON comparison on key-order drift (i.e. across restarts, where the semantic check is the point).

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean, well-contained perf fix. store_matches correctly skips load (volatile/unread), uses JSON value comparison for spec to handle non-deterministic map ordering, and gates the expensive JSON parses behind cheap scalar checks. Deserialization failures conservatively fall through to re-put (the old behavior). Channel capacity bump (64→1024, ~100 KB fixed) is well-justified. Test covers both directions. No issues found.

0 🔴 Important · 0 🟡 Nit · 0 🟣 Pre-existing

…re_matches

Re-serialising the same in-process spec instance is byte-stable, so the
steady-state resync/reconcile comparison short-circuits on bytes and only
falls through to the semantic JSON comparison on key-order drift (e.g.
across restarts). Suggested by review on #1667.

Signed-off-by: Chang Su <8605658+CatherineSue@users.noreply.github.com>
@CatherineSue
CatherineSue merged commit 174f0c5 into main Jun 13, 2026
43 of 47 checks passed
@CatherineSue
CatherineSue deleted the chang/mesh-resync-burst branch June 13, 2026 09:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant