Skip to content

fix(server): unify dependency freshness on consumed-content snapshots - #507

Merged
16bit-ykiko merged 3 commits into
mainfrom
fix/deps-freshness
Jul 13, 2026
Merged

16bit-ykiko merged 3 commits into
mainfrom
fix/deps-freshness

Conversation

@16bit-ykiko

@16bit-ykiko 16bit-ykiko commented Jul 12, 2026 •

Copy link
Copy Markdown
Member

Problem

Compilation artifacts (PCH / PCM / AST / synthesized header preambles) shared one freshness snapshot whose two layers were both unsound, and the index shards carried their own, subtly different copy of the same check. Audited failure modes (bughunt F44 / F01 / F04 / F33):

  • Same-second save is permanently invisible (F44, Critical). Layer 1 compared each dep's mtime against a single build_at watermark, both truncated to whole seconds, with <=. A header saved in the same second the dependent's compile finished satisfies mtime <= build_at forever — the hash layer never runs, and the dependent keeps compiling against a stale PCH until the header changes again. Reproduced with a natural type-then-save rhythm; no adversarial timing needed.
  • The snapshot describes the wrong bytes (F01, High). capture_deps_snapshot ran on the master after the build and hashed the current disk. A header saved while the build was running yields a snapshot that blesses content the build never read; the poisoned entry persists via cache.json across restarts.
  • Backdated / preserved mtimes are invisible (F04). Any content change whose mtime does not advance past the watermark (rsync -t, git-restore-mtime, clock steps) never reaches the hash layer.
  • PCM deps were never populated (F33, High). The PCM compile() overload never filled out.deps, so every PCM snapshot was empty — and the PCM cache key embeds no content, so "graceful exit → edit a module interface offline → restart" deterministically serves a stale PCM. The module source itself also has to be a dep, which it never was.
  • Index-side gaps: need_update validated only the first compilation context of a shard (a header shard has one per including TU); merge paired the worker's rows with content re-read from disk at merge time (position misalignment when the file changed in between); the same watermark blindness applied to shard staleness.

Design

One freshness vocabulary everywhere, aligned with the one implementation that was already right (FileTracker):

  • Per-dep records instead of a watermark. DepsSnapshot is now a list of DepState { path_id, size, mtime_ns, hash, missing }. Layer 1 passes only when size and nanosecond mtime are equal to the record — backdated or preserved mtimes can no longer masquerade as fresh. Layer 2 hashes the disk against the recorded hash; a match is a touch, not an edit, and repairs the fast path in place.
  • Hashes describe consumed bytes. Workers hash each dependency from the compiler's own in-memory buffers (SourceManager) and ship {path, hash} pairs plus build_at (sampled before the compile) in CompileResult / BuildResult. A later disk read can no longer impersonate what the build saw.
  • Stat baselines are gated by the build start. The master records a {size, mtime_ns} fast path only for deps whose mtime precedes build_at by a 2s filesystem-granularity guard (fs::stat_baseline_before_ns). Anything newer keeps only the consumed hash and re-earns its fast path through one hash comparison. The synthesized header-preamble chain (a third producer of these records) applies the same guard.
  • PCM deps populated (out.deps = unit.deps() + the canonicalized module source with its content hash).
  • Index shards use the same scheme. TUIndex ships per-path consumed hashes; MergedIndex::merge records DepStamp fast paths under the same guard (new, backward-compatible flatbuffer field); need_update validates every compilation context, in-memory and serialized; the indexer skips a merge whose rows no longer hash to the disk content and keeps the last-known snapshot serving instead of stripping it.
  • Formats migrate in place. Old cache.json entries load with zeroed baselines (unknown build_at skipped, missing fields defaulted) and revalidate by hash once; old shards without dep_stamps do the same. No cache or index version bump, nothing is rebuilt on upgrade.

force_revalidate (previously build_at = 0) now zeroes the per-dep fast paths while keeping the hashes, which is the same semantics expressed per record.

Testing

  • New DepsSnapshot unit suite: the full freshness truth table — same-second edit, backdated edit, touch-repairs-fast-path, poisoned-capture detection, no-baseline convergence, missing-file transitions, force_revalidate.
  • New IndexerMerge unit suite: a merge whose rows describe outdated disk content is skipped and the last-known snapshot (rows included) keeps serving; the settled reindex lands.
  • MergedIndex: multi-context need_update exercised for both contexts and through a serialized view; backdated edits; stamp round-trip.
  • Integration: same-second save (deliberately no mtime sleep), backdated header change (no didSave), PCM offline-interface-edit across restart, cache.json PCM entries contain the module source with non-zero hashes, old-format cache.json upgrade path.
  • Existing suites adapted where the schema changed (assertions strengthened, none weakened).
  • Local runs: Debug (ASAN) 909 unit / 271 integration / 3 smoke, RelWithDebInfo 927 unit — all green.
  • Pre-PR review: 3 independent reviewers (correctness / style / tests); all findings addressed in this branch (notably: the chain-snapshot mtime guard, merge-skip sweep protection, deterministic multi-context regression tests).

Per-dep {size, mtime_ns, hash} equality records replace the second-
truncated build_at watermark; hashes come from the worker's own compile
buffers, PCM deps are finally populated (incl. the module source), the
index merge verifies disk content against the indexed hash, and
need_update validates every compilation context.
@coderabbitai

coderabbitai Bot commented Jul 12, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Changes

Dependency freshness pipeline

Layer / File(s) Summary
Hashed dependency model
src/compile/*, src/server/protocol/worker.h, src/server/worker/*, tests/unit/compile/*, tests/unit/server/stateless_worker_tests.cpp
Compilation dependencies now carry consumed-content hashes, build timestamps, and module-source entries.
Cache snapshot and revalidation
src/server/state/*, src/server/compiler/*, src/support/filesystem.h
PCH, PCM, and header snapshots use guarded stat fast paths with content-hash fallback and cache serialization support.
Indexed content validation
src/index/*, src/server/compiler/indexer.cpp, tests/unit/index/*, tests/unit/server/indexer_tests.cpp
Indexes persist path hashes and dependency stamps, validate all contexts, and skip merges when disk content differs.
Freshness and compatibility coverage
tests/integration/compilation/*, tests/unit/server/deps_snapshot_tests.cpp, tests/unit/server/context_resolver_tests.cpp, tests/unit/test/temp_dir.h
Tests cover cache migration, mtime edge cases, snapshot repair, serialized validation, and stale index protection.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Compiler
  participant Workspace
  participant Filesystem
  participant Indexer
  Compiler->>Workspace: capture dependency hashes and stat baselines
  Workspace->>Filesystem: validate dependency metadata and content
  Filesystem-->>Workspace: freshness result
  Workspace-->>Compiler: cache reuse or rebuild decision
  Indexer->>Filesystem: hash indexed TU paths
  Filesystem-->>Indexer: current content hash
  Indexer-->>Compiler: merge accepted or skipped
Loading

Possibly related PRs

  • clice-io/clice#391: Related persistent PCH/PCM cache reuse and dependency snapshot behavior.
  • clice-io/clice#479: Related header-context dependency snapshots and content-hash invalidation.
  • clice-io/clice#485: Related merged-index dependency hashes and persisted staleness validation.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 34.04% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: unifying server dependency freshness around consumed-content snapshots.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/deps-freshness

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 876fc7ec14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/server/state/workspace.cpp
Comment thread src/server/state/workspace.cpp
Comment thread src/server/compiler/indexer.cpp Outdated
Comment thread src/index/merged_index.cpp
Comment thread src/server/state/workspace.cpp

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (2)
tests/integration/compilation/test_persistent_cache.py (1)

165-234: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win

Missing anomaly checks in multi-session tests.

Both test_old_cache_json_upgrades and test_pcm_offline_edit_invalidates bypass the client fixture's automatic check_no_anomaly teardown by using make_client/shutdown_client directly. test_old_cache_json_upgrades never calls assert_no_anomaly for either session, and test_pcm_offline_edit_invalidates only calls it after session 1 (line 220), not after session 2. Given these tests exercise exactly the kind of stat/hash edge cases most likely to surface an internal crash, an anomaly check on every session would be a meaningful safety net, consistent with the fixture's own rationale.

🛡️ Suggested addition
     c2 = await make_client(executable, tmp_path)
     uri2, _ = await c2.open_and_wait(tmp_path / "main.cpp")
     assert_clean_compile(c2, uri2)
     # Loaded, hash-validated, reused — not rebuilt.
     assert list_pch_files(tmp_path)[0].stat().st_mtime == pch_mtime_s1
+    assert_no_anomaly(c2, tmp_path)
     await shutdown_client(c2)
     c2 = await make_client(executable, tmp_path)
     mid_uri2, _ = await c2.open_and_wait(tmp_path / "mid.cppm")
     assert_has_errors(c2, mid_uri2, "Expected errors after offline interface edit")
+    assert_no_anomaly(c2, tmp_path)
     await shutdown_client(c2)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/compilation/test_persistent_cache.py` around lines 165 -
234, Make anomaly validation explicit for every client session in
test_old_cache_json_upgrades and test_pcm_offline_edit_invalidates. Call
assert_no_anomaly with the relevant client and workspace before each
shutdown_client, including both c1 and c2 in the cache-upgrade test and c2 in
the PCM invalidation test, while preserving the existing assertions and shutdown
flow.
src/index/merged_index.cpp (1)

105-115: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate hash_file — consider hoisting to a shared header.

This is byte-identical to workspace::hash_file (src/server/state/workspace.cpp), which the doc comment itself calls out. Since fs::stat_baseline_before_ns/fs::mtime_ns are already shared via src/support/filesystem.h, moving this helper there too would give both subsystems a single source of truth for the freshness hash scheme.

♻️ Suggested consolidation
-namespace {
-
-/// Hash a file's content with the same scheme the server layer uses for its
-/// dependency snapshots (`workspace::hash_file`). Returns 0 on read failure.
-std::uint64_t hash_file(llvm::StringRef path) {
-    auto buffer = llvm::MemoryBuffer::getFile(path);
-    if(!buffer) {
-        return 0;
-    }
-    return llvm::xxh3_64bits((*buffer)->getBuffer());
-}
-
-}  // namespace
+// hash_file now lives in support/filesystem.h (fs::hash_file), shared with
+// workspace.cpp's dependency snapshot capture.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/index/merged_index.cpp` around lines 105 - 115, Consolidate the duplicate
hash_file implementation by moving the shared file-content hashing helper into
src/support/filesystem.h and its implementation location, then update the
merged-index and workspace callers to reuse it. Remove the local
anonymous-namespace hash_file in the merged-index code and preserve the existing
zero-on-read-failure behavior and xxh3 hashing scheme.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/server/compiler/context_resolver.cpp`:
- Around line 566-585: Update resolve_header_context to re-stat each header
after reading its buffer, including the target-header path. Retain the fast-path
size and mtime only when pre-read and post-read metadata match and the post-read
mtime is at or before baseline_before_ns; otherwise store zeroed metadata while
preserving the buffer hash. Apply this consistently to the related
dependency-recording paths near the additional reported locations.

---

Nitpick comments:
In `@src/index/merged_index.cpp`:
- Around line 105-115: Consolidate the duplicate hash_file implementation by
moving the shared file-content hashing helper into src/support/filesystem.h and
its implementation location, then update the merged-index and workspace callers
to reuse it. Remove the local anonymous-namespace hash_file in the merged-index
code and preserve the existing zero-on-read-failure behavior and xxh3 hashing
scheme.

In `@tests/integration/compilation/test_persistent_cache.py`:
- Around line 165-234: Make anomaly validation explicit for every client session
in test_old_cache_json_upgrades and test_pcm_offline_edit_invalidates. Call
assert_no_anomaly with the relevant client and workspace before each
shutdown_client, including both c1 and c2 in the cache-upgrade test and c2 in
the PCM invalidation test, while preserving the existing assertions and shutdown
flow.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 11104059-06b7-404d-8a9c-94988c262802

📥 Commits

Reviewing files that changed from the base of the PR and between 5b2184a and 876fc7e.

📒 Files selected for processing (35)
  • src/compile/compilation.cpp
  • src/compile/compilation.h
  • src/compile/compilation_unit.cpp
  • src/compile/compilation_unit.h
  • src/compile/dep_file.h
  • src/index/include_graph.cpp
  • src/index/include_graph.h
  • src/index/merged_index.cpp
  • src/index/merged_index.h
  • src/index/schema.fbs
  • src/index/tu_index.cpp
  • src/server/compiler/compiler.cpp
  • src/server/compiler/compiler.h
  • src/server/compiler/context_resolver.cpp
  • src/server/compiler/context_resolver.h
  • src/server/compiler/indexer.cpp
  • src/server/protocol/worker.h
  • src/server/state/file_tracker.cpp
  • src/server/state/invalidator.h
  • src/server/state/workspace.cpp
  • src/server/state/workspace.h
  • src/server/transport/master_server.cpp
  • src/server/worker/stateful_worker.cpp
  • src/server/worker/stateless_worker.cpp
  • src/support/filesystem.h
  • tests/integration/compilation/test_persistent_cache.py
  • tests/integration/compilation/test_staleness.py
  • tests/unit/compile/compilation_tests.cpp
  • tests/unit/compile/directive_tests.cpp
  • tests/unit/index/merged_index_tests.cpp
  • tests/unit/server/context_resolver_tests.cpp
  • tests/unit/server/deps_snapshot_tests.cpp
  • tests/unit/server/indexer_tests.cpp
  • tests/unit/server/stateless_worker_tests.cpp
  • tests/unit/test/temp_dir.h

Comment thread src/server/compiler/context_resolver.cpp Outdated
deps() resolves every path through real_path, so on macOS the recorded
dep spelling is /private/var/... while TempDir and pytest hand out the
/var/... symlink form. Exact string equality in BuildPCMRequest and
test_pcm_cache_entry_has_deps only held on Linux; resolve both sides
before comparing.
- Skip the whole TUIndex merge when the main file moved on: header
  shards merged before the main-file check could mix two generations,
  and the sweep stripped contributions the stale result no longer
  mentioned.
- Drop dep-less PCM cache entries at load: they predate populated deps
  and would blindly serve a stale PCM after upgrade.
- Stat chain files and the target after reading them, so a write racing
  the read lands inside the mtime guard and earns no fast path.
- Document the still-missing-dep and identical-stat residuals.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/server/compiler/indexer.cpp (1)

126-140: 🚀 Performance & Scalability | 🔵 Trivial | 💤 Low value

Consider avoiding the std::string copy for header content.

header_buf is alive for the duration of the else block, so header_content could reference its buffer directly instead of copying into header_content_storage. This avoids a heap allocation and copy per header shard.

♻️ Optional refactor
         auto header_path = workspace.path_pool.resolve(global_path_id);
         llvm::StringRef header_content;
-        std::string header_content_storage;
         auto header_buf = llvm::MemoryBuffer::getFile(header_path);
         if(header_buf) {
-            header_content_storage = (*header_buf)->getBuffer().str();
-            header_content = header_content_storage;
+            header_content = (*header_buf)->getBuffer();
         }

The header arbitration logic (lines 133-140) — unconditional content_matches check, touched.insert on mismatch to preserve prior contributions — looks correct.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/server/compiler/indexer.cpp` around lines 126 - 140, Update the header
content setup near content_matches to reference the buffer returned by
header_buf directly, removing header_content_storage and its heap-copy
operation. Preserve the existing buffer lifetime through the content_matches
call and retain the unconditional mismatch handling, touched insertion, and
early return.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@src/server/compiler/indexer.cpp`:
- Around line 126-140: Update the header content setup near content_matches to
reference the buffer returned by header_buf directly, removing
header_content_storage and its heap-copy operation. Preserve the existing buffer
lifetime through the content_matches call and retain the unconditional mismatch
handling, touched insertion, and early return.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 9ec4ab7f-3cab-480f-b03b-63a0f863b3f4

📥 Commits

Reviewing files that changed from the base of the PR and between 37cad8a and 3cda3f6.

📒 Files selected for processing (5)
  • src/server/compiler/context_resolver.cpp
  • src/server/compiler/indexer.cpp
  • src/server/state/workspace.cpp
  • src/server/state/workspace.h
  • tests/integration/compilation/test_persistent_cache.py
🚧 Files skipped from review as they are similar to previous changes (3)
  • src/server/state/workspace.h
  • src/server/compiler/context_resolver.cpp
  • src/server/state/workspace.cpp

@16bit-ykiko
16bit-ykiko merged commit 77a53b8 into main Jul 13, 2026
22 checks passed
@16bit-ykiko
16bit-ykiko deleted the fix/deps-freshness branch July 13, 2026 14:28
16bit-ykiko added a commit that referenced this pull request Jul 13, 2026
Both Windows CI jobs are red on main since #507 (its own merge run
29258064794 failed): `MergedIndex.NeedUpdateChecksAllContexts` and
`MergedIndex.SerializedStampsValidate` fail deterministically with
`need_update() expected true, got false`.

## Root cause

Windows file times advance in ~16ms clock ticks. The failing assertions
perform **same-size** rewrites (`int b2();` → `int b3();`, `int f();` →
`int g();`) moments after the stamp was recorded; when the rewrite lands
in the same tick, the file reproduces the stamp's size **and** mtime
exactly, so the stat fast path — equality-based by design — rightly
reports fresh, and the hash layer the assertions meant to exercise is
never consulted. On Linux/macOS nanosecond timestamps always move, which
is why only Windows fails.

**Production is immune**: `merge` only records a stamp for files
untouched since two seconds before the build
(`stat_baseline_before_ns`), so a stamped file's later edit can never
share its tick. The tests defeat that guard deliberately
(`generous_build_at()` = now + 10s) to earn the fast path, which is what
exposes them to the tick.

## Fix

After each same-size rewrite, bump the file's mtime explicitly
(`set_file_mtime`, +5s) — modelling the reality that a genuine edit
arrives long after the stamp, and pinning those verdicts on the hash
layer as intended. Test-only change; no product code touched.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant