Skip to content

fix(mesh): prevent stale snapshot chunks from mixing across retries - #837

Merged
slin1237 merged 1 commit into
mainfrom
slin/mesh-snapshot-cleanup
Mar 21, 2026
Merged

slin1237 merged 1 commit into
mainfrom
slin/mesh-snapshot-cleanup

Conversation

@slin1237

@slin1237 slin1237 commented Mar 20, 2026 •

Copy link
Copy Markdown
Member

Summary

Fixes a snapshot corruption bug where partial chunks from a failed snapshot transfer could mix with chunks from a new attempt.

Problem

The snapshot_state HashMap was keyed by (store_type, total_chunks). If a peer disconnected mid-snapshot and reconnected, and the new snapshot had the same total_chunks value, old partial chunks from the previous attempt would be mixed with new ones — producing corrupted state on the receiving node.

What changed

  • ping_server.rs: Key snapshot_state by store_type only (not total_chunks). Track expected_total alongside the chunk vector. When total_chunks changes for a store, discard old partial chunks and start fresh.

Test plan

  • cargo test -p smg-mesh — 152 pass
  • cargo clippy -p smg-mesh --all-targets -- -D warnings — clean

Summary by CodeRabbit

  • Refactor
    • More robust snapshot synchronization: transfers that restart or deliver chunks out of order are handled reliably, reducing failed or partial syncs.
    • Incoming transfer state is reset when a new transfer starts; chunk sequences are verified for continuity before applying snapshots, preventing mixed or corrupted restores and adding clearer discard logging.

@slin1237
slin1237 requested review from llfl and tonyluj as code owners March 20, 2026 19:46
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request resolves a critical snapshot corruption bug where partial data from a failed snapshot transfer could mix with chunks from a new attempt if the total_chunks value remained the same. The changes refactor the snapshot reception state tracking mechanism to key by store_type only and explicitly manage the expected total chunks. This ensures that any new snapshot request for a store type automatically discards incomplete data from prior attempts, thereby preventing data mixing and guaranteeing the integrity of received snapshots.

Highlights

  • Snapshot State Keying Refactor: The snapshot_state HashMap's key was changed from (store_type, total_chunks) to store_type only, to prevent mixing of stale snapshot chunks across retries.
  • Tracking Expected Total Chunks: The snapshot_state now explicitly tracks the expected_total number of chunks alongside the received chunks for each store type, allowing for better validation and state management.
  • Discarding Stale Chunks: Logic was added to detect when total_chunks changes for a given store type during a snapshot transfer. If a change is detected, any previously received partial chunks are automatically discarded, ensuring data integrity for new attempts.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 20, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: b36b14d5-7902-4dcc-a4cf-7f0390ce36fc

📥 Commits

Reviewing files that changed from the base of the PR and between 69a3089 and aaea31d.

📒 Files selected for processing (1)
  • crates/mesh/src/ping_server.rs

📝 Walkthrough

Walkthrough

Snapshot reception buffering was restructured to key by store_type mapping to (Vec<SnapshotChunk>, expected_total). Receiving a chunk with chunk_index == 0 clears prior buffered chunks for that store_type and updates expected_total; completion compares buffer length to expected_total, verifies contiguous indices, applies, then removes the store_type entry.

Changes

Cohort / File(s) Summary
Snapshot Chunk State Accumulation
crates/mesh/src/ping_server.rs
Refactored buffering from a composite (store_type, total_chunks) -> Vec<SnapshotChunk> key to store_type -> (Vec<SnapshotChunk>, expected_total). On chunk_index == 0 the prior buffer for that store_type is discarded and expected_total updated; completion uses the stored expected_total, verifies contiguous indices before apply, and clears the store_type entry after successful apply. Logs warnings and clears state when indices are non-contiguous.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Suggested reviewers

  • tonyluj
  • llfl

Poem

🐰 I hop and stash each tiny part,

When a zero comes I make a fresh start,
One burrow holds the pieces whole,
I count the hops to reach the goal,
Then stitch the snapshot — carrot-roll.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: fixing a bug where stale snapshot chunks from failed transfers could mix with subsequent attempts. This directly aligns with the root cause and solution implemented in the PR.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch slin/mesh-snapshot-cleanup

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request effectively addresses a critical snapshot corruption bug by refining the snapshot_state tracking mechanism. By keying the snapshot_state HashMap solely by LocalStoreType and storing the expected_total chunks alongside the received_chunks, the system can now correctly identify and discard stale partial snapshot data when a new snapshot attempt for the same store type is initiated. This prevents the mixing of chunks from different attempts, which was the root cause of the corruption. The changes are well-reasoned and directly solve the problem.

chunks.clear();
*expected = chunk.total_chunks;
}
chunks.push(chunk.clone());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The chunk.clone() operation here, and subsequently received_chunks.to_vec() on line 926, involves cloning SnapshotChunks. A SnapshotChunk contains a Vec<StateUpdate>, and each StateUpdate contains a Vec<u8> for its value. Cloning Vec<u8> results in a full memory copy, which can be inefficient for large snapshots. While necessary to ensure ownership for storage and sorting, consider the performance implications if snapshot sizes are expected to be very large. If StateUpdate.value could be bytes::Bytes (which is reference-counted), these clones would be much cheaper.

References
  1. Cloning Vec<u8> results in a full memory copy, which can be inefficient for large data. Using reference-counted types like bytes::Bytes can make cloning cheaper.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 990ca17ea9

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread crates/mesh/src/ping_server.rs Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@crates/mesh/src/ping_server.rs`:
- Around line 902-915: The snapshot assembly logic in snapshot_state (variables:
snapshot_state, store_type, chunks, expected) only clears partial state when
chunk.total_chunks changes, so if a sender restarts with the same total_chunks
stale chunks can mix; modify the block that handles inserting into
snapshot_state to also treat a received chunk with chunk.chunk_index == 0 as the
start of a fresh transfer: if chunk.chunk_index == 0 and !chunks.is_empty() then
clear chunks and set *expected = chunk.total_chunks before pushing the new
chunk, ensuring retries that reuse the same total_chunks don't corrupt the
assembly.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 04b32461-d7ec-4fa3-a3cc-59a588899377

📥 Commits

Reviewing files that changed from the base of the PR and between ccf54bf and 990ca17.

📒 Files selected for processing (1)
  • crates/mesh/src/ping_server.rs

Comment thread crates/mesh/src/ping_server.rs Outdated
@slin1237
slin1237 force-pushed the slin/mesh-snapshot-cleanup branch from 990ca17 to a4a7b7c Compare March 20, 2026 20:12

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@crates/mesh/src/ping_server.rs`:
- Around line 906-917: The snapshot handling currently rewrites expected total
on every frame and treats a snapshot complete based only on buffer length;
modify the logic in the snapshot_state handling (where
snapshot_state.entry(store_type) returns (chunks, expected) and
chunk.total_chunks / chunk.chunk_index are used) so that expected is set only
when starting a new transfer (e.g., when chunks.is_empty() or when
chunk.chunk_index == 0) and not overwritten on subsequent frames, and before
applying a snapshot verify that the collected chunks contain each index
0..expected-1 exactly once (check for duplicate or missing chunk.chunk_index
values) rather than relying on Vec length; reject or restart transfers if
indexes are out of range, duplicated, or missing.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: b4cf9b53-0d34-4209-9fcb-0b98c5375007

📥 Commits

Reviewing files that changed from the base of the PR and between 990ca17 and a4a7b7c.

📒 Files selected for processing (1)
  • crates/mesh/src/ping_server.rs

Comment thread crates/mesh/src/ping_server.rs
@slin1237
slin1237 force-pushed the slin/mesh-snapshot-cleanup branch from a4a7b7c to 2b4e9f3 Compare March 20, 2026 21:30

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
crates/mesh/src/ping_server.rs (1)

906-931: ⚠️ Potential issue | 🟠 Major

Keep expected_total stable and validate chunk identity before applying.

Line 916 still rewrites the tracked total on every frame, and Line 923 still treats len() == total as completion. A retried or inconsistent stream like [0,1,1] or [0(total=3),1(total=2)] can therefore still apply a truncated snapshot. Set expected_total only when starting/resetting a transfer, reject duplicate/out-of-range chunk_index values, and only apply once 0..expected_total-1 is present exactly once.

🛠️ Suggested hardening
 let (chunks, expected) = snapshot_state
     .entry(store_type)
     .or_insert_with(|| (Vec::new(), chunk.total_chunks));
 if chunk.chunk_index == 0 && !chunks.is_empty() {
     log::info!(
         "New snapshot transfer for {:?}, discarding {} partial chunks",
         store_type, chunks.len()
     );
     chunks.clear();
-}
-*expected = chunk.total_chunks;
-chunks.push(chunk.clone());
+    *expected = chunk.total_chunks;
+} else if *expected != chunk.total_chunks {
+    log::warn!(
+        "Snapshot total changed mid-transfer for {:?} ({} -> {}), resetting buffer",
+        store_type,
+        *expected,
+        chunk.total_chunks
+    );
+    chunks.clear();
+    *expected = chunk.total_chunks;
+}
+if chunk.chunk_index < *expected
+    && chunks.iter().all(|c| c.chunk_index != chunk.chunk_index)
+{
+    chunks.push(chunk.clone());
+}
 
 // Check if we've received all chunks
 if let Some((received_chunks, total)) = snapshot_state.get(&store_type) {
-    if received_chunks.len() as u64 == *total {
+    let mut sorted_chunks = received_chunks.to_vec();
+    sorted_chunks.sort_by_key(|c| c.chunk_index);
+    let complete = sorted_chunks.len() as u64 == *total
+        && sorted_chunks.iter().enumerate().all(|(idx, c)| {
+            c.total_chunks == *total && c.chunk_index == idx as u64
+        });
+    if complete {
         // All chunks received, apply snapshot
         log::info!("All {} chunks received for store {:?}, applying snapshot",
             total, store_type);
-
-        if let Some(ref stores) = stores {
-            // Sort chunks by index
-            let mut sorted_chunks = received_chunks.to_vec();
-            sorted_chunks.sort_by_key(|c| c.chunk_index);
+        if let Some(ref stores) = stores {
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@crates/mesh/src/ping_server.rs` around lines 906 - 931, The current logic
mutates the tracked total every frame and treats len==total as sufficient;
change it so expected total is set only when starting/resetting a transfer (when
inserting the entry or when chunk.chunk_index == 0 and you clear chunks), do not
overwrite *expected on every incoming chunk; validate each incoming chunk:
reject and ignore duplicates (same chunk_index already present) and out-of-range
indices (chunk_index >= expected when expected is known), and only apply the
snapshot in the completion branch (the code that currently sorts and applies)
after verifying the set of received chunk_index values equals exactly
0..expected-1 with no gaps or duplicates. Use the existing snapshot_state entry
tuple (Vec<Chunk>, expected_total) and the variables store_type,
chunk.chunk_index and chunk.total_chunks to locate and implement these checks.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@crates/mesh/src/ping_server.rs`:
- Around line 906-931: The current logic mutates the tracked total every frame
and treats len==total as sufficient; change it so expected total is set only
when starting/resetting a transfer (when inserting the entry or when
chunk.chunk_index == 0 and you clear chunks), do not overwrite *expected on
every incoming chunk; validate each incoming chunk: reject and ignore duplicates
(same chunk_index already present) and out-of-range indices (chunk_index >=
expected when expected is known), and only apply the snapshot in the completion
branch (the code that currently sorts and applies) after verifying the set of
received chunk_index values equals exactly 0..expected-1 with no gaps or
duplicates. Use the existing snapshot_state entry tuple (Vec<Chunk>,
expected_total) and the variables store_type, chunk.chunk_index and
chunk.total_chunks to locate and implement these checks.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 3b10875a-a2b1-4874-ab31-24f9fd87aa10

📥 Commits

Reviewing files that changed from the base of the PR and between a4a7b7c and 2b4e9f3.

📒 Files selected for processing (1)
  • crates/mesh/src/ping_server.rs

@slin1237
slin1237 force-pushed the slin/mesh-snapshot-cleanup branch from 2b4e9f3 to 69a3089 Compare March 20, 2026 23:28

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@crates/mesh/src/ping_server.rs`:
- Around line 919-940: The contiguous-index check on snapshot chunks (in the
snapshot_state handling around store_type / received_chunks / total) should also
validate that every chunk's declared total_chunks matches the expected total to
catch malformed or corrupted chunks; update the validation (where
sorted_chunks.iter().enumerate().all(...) is computed) to require both
c.chunk_index == i as u64 and c.total_chunks == *total, and if that combined
check fails remove snapshot_state for store_type (same behavior as the current
non-contiguous branch) so inconsistent total_chunks are treated as an invalid
snapshot.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 68399b96-d104-40ce-83bd-03d3fe32b3c6

📥 Commits

Reviewing files that changed from the base of the PR and between 2b4e9f3 and 69a3089.

📒 Files selected for processing (1)
  • crates/mesh/src/ping_server.rs

Comment on lines +919 to +940
// Check if we've received all chunks with valid indices
if let Some((received_chunks, total)) =
snapshot_state.get(&store_type)
{
if received_chunks.len() as u64 == *total {
// Verify all indices 0..total are present (no duplicates/gaps)
let mut sorted_chunks = received_chunks.to_vec();
sorted_chunks.sort_by_key(|c| c.chunk_index);
let indices_valid = sorted_chunks.iter().enumerate().all(
|(i, c)| c.chunk_index == i as u64,
);
if !indices_valid {
log::warn!(
"Snapshot for {:?} has {} chunks but indices are not contiguous 0..{}, discarding",
store_type, sorted_chunks.len(), total
);
snapshot_state.remove(&store_type);
continue;
}

log::info!("All {} chunks received for store {:?}, applying snapshot",
chunk.total_chunks, store_type);
total, store_type);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Good addition of contiguous index validation.

The validation that indices are contiguous 0..total correctly catches scenarios like duplicate chunks being pushed. The sort-then-enumerate approach is clean.

Consider also verifying that all chunks agree on total_chunks during the completion check to further harden against malformed/corrupted data:

let indices_valid = sorted_chunks.iter().enumerate().all(
    |(i, c)| c.chunk_index == i as u64 && c.total_chunks == *total,
);
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@crates/mesh/src/ping_server.rs` around lines 919 - 940, The contiguous-index
check on snapshot chunks (in the snapshot_state handling around store_type /
received_chunks / total) should also validate that every chunk's declared
total_chunks matches the expected total to catch malformed or corrupted chunks;
update the validation (where sorted_chunks.iter().enumerate().all(...) is
computed) to require both c.chunk_index == i as u64 and c.total_chunks ==
*total, and if that combined check fails remove snapshot_state for store_type
(same behavior as the current non-contiguous branch) so inconsistent
total_chunks are treated as an invalid snapshot.

The snapshot_state HashMap keyed by (store_type, total_chunks) could mix
chunks from different snapshot attempts if a peer disconnected mid-transfer
and reconnected. If both attempts had the same total_chunks value, old
partial chunks would mix with new ones, producing corrupted state.

What changed:
- crates/mesh/src/ping_server.rs: key snapshot_state by store_type only
  (not total_chunks). When total_chunks changes for a store (new snapshot
  attempt), discard the old partial chunks. This prevents stale chunk
  mixing across reconnections.

Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@slin1237
slin1237 force-pushed the slin/mesh-snapshot-cleanup branch from 69a3089 to aaea31d Compare March 21, 2026 00:19
@slin1237
slin1237 merged commit 7cfc963 into main Mar 21, 2026
28 checks passed
@slin1237
slin1237 deleted the slin/mesh-snapshot-cleanup branch March 21, 2026 00:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant