Skip to content

feat: add dependency graph scanner with benchmark - #365

Closed
16bit-ykiko wants to merge 63 commits into
mainfrom
scan-deps
Closed

16bit-ykiko wants to merge 63 commits into
mainfrom
scan-deps

Conversation

@16bit-ykiko

@16bit-ykiko 16bit-ykiko commented Mar 23, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Add wavefront BFS dependency scanner: builds complete include graph from CDB
  • Include resolver with search path extraction using clang ArgumentParser
  • Toolchain query caching in CompilationDatabase (deduplicates by driver+target+std)
  • Parallel file scanning via et::queue() thread pool
  • scan_benchmark CLI with detailed report and JSON graph export
  • CI benchmark workflow: runs scan against full LLVM (all projects) on 3 platforms

Test plan

  • CI benchmark passes on Linux/macOS/Windows
  • Unit tests for include resolver and dependency graph
  • Verify scan report accuracy against known CDB

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Dependency-graph scanning to discover file/include relationships and produce detailed scan reports
    • Async include resolution with configurable search semantics
    • Benchmark executable and CI workflow to measure scan performance across OSes
  • Improvements

    • More efficient toolchain-query caching for faster compilation-database handling
    • Richer include metadata (angled vs quoted, include_next) and removal of legacy fuzzy scan mode
    • Path interning for compact path identifiers
  • Tests

    • New unit/integration tests covering graph scanning and include resolution

Implement a wavefront BFS dependency scanner that builds a complete include
graph from a compilation database. Key components:

- PathPool: intern pool mapping file paths to compact uint32_t IDs
- IncludeResolver: resolve #include directives using search paths from CDB
- DependencyGraph: store per-(file, config) include edges with conditional flags
- CompilationDatabase: toolchain query caching and extract_search_config API
- scan_benchmark: CLI tool with detailed scan report and JSON graph export
- CI benchmark workflow: test scan performance against LLVM on all platforms

Parallelizes file reads via et::queue() thread pool for ~3-4x speedup.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Mar 23, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds async include resolution, a PathPool for path interning, a wavefront dependency-graph scanner and public DependencyGraph/ScanReport APIs, integrates scanning into MasterServer, improves compilation-database toolchain/search caching, adds a scan_benchmark executable and CI workflow, and introduces comprehensive unit tests for the new components.

Changes

Cohort / File(s) Summary
CI/CD
\.github/workflows/benchmark.yml
New GitHub Actions workflow "benchmark" that builds the project on PRs to main across Ubuntu/macOS/Windows and runs scan_benchmark against an LLVM compile_commands.json.
Build / Executables
CMakeLists.txt, benchmarks/scan_benchmark.cpp
Added scan_benchmark executable; added src/syntax/include_resolver.cpp and src/syntax/dependency_graph.cpp into clice-core build; benchmark loads a compilation DB, runs scans multiple times, prints metrics, and can export graph JSON.
Path interning
src/support/path_pool.h
New clice::PathPool type providing intern/resolve backed by a bump allocator, path vector, and string map.
Dependency graph (API & impl)
src/syntax/dependency_graph.h, src/syntax/dependency_graph.cpp
New DependencyGraph, ScanReport, and blocking scan_dependency_graph() implementing a wavefront BFS scan, per-(file,config) include lists with conditional flags, module mapping, and timing/metrics.
Include resolution (API & impl)
src/syntax/include_resolver.h, src/syntax/include_resolver.cpp
Added ResolveResult and async resolve_include() (et::task) with a stat cache; supports angled/quoted/#include_next semantics and found-dir index.
Scan pipeline
src/syntax/scan.h, src/syntax/scan.cpp
Extended ScanResult::IncludeInfo with is_angled and is_include_next; removed scan_fuzzy() and simplified ScanDirectivesGetter by dropping fuzzy-mode logic while keeping precise scanning.
Compilation DB / Toolchain
src/compile/command.h, src/compile/command.cpp
Added SearchDir/SearchConfig, toolchain-option classification and extraction, a string-keyed toolchain cache with query_toolchain_cached(), and APIs extract_search_config() and resolve_path().
Server integration
src/server/master_server.h, src/server/master_server.cpp
Replaced local ServerPathPool with external PathPool, added DependencyGraph dependency_graph, adjusted includes, and invoke scan_dependency_graph() during load_workspace().
Include resolver impl
src/syntax/include_resolver.cpp
Filesystem-stat cache and coroutine-based include resolution implementation using search config and event loop.
Tests
tests/unit/syntax/dependency_graph_tests.cpp, tests/unit/syntax/include_resolver_tests.cpp, tests/unit/syntax/scan_tests.cpp
Added comprehensive unit tests for DependencyGraph, scan_dependency_graph, and include resolution; removed fuzzy-scan tests and updated scan assertions.

Sequence Diagram(s)

sequenceDiagram
    rect rgba(200,200,255,0.5)
    actor User
    end
    participant Benchmark as "scan_benchmark"
    participant CDB as "CompilationDatabase"
    participant Scan as "scan_dependency_graph()"
    participant EventLoop as "Event Loop"
    participant Workers as "Worker Threads"
    participant Resolver as "resolve_include()"
    participant PathPool as "PathPool"
    participant Graph as "DependencyGraph"
    participant FS as "Filesystem / stat_cache"

    User->>Benchmark: run with compile_commands.json
    Benchmark->>CDB: load compile_commands.json & updates
    Benchmark->>Scan: call scan_dependency_graph(cdb, updates, path_pool, graph)
    Scan->>EventLoop: create & run async tasks
    EventLoop->>Workers: schedule file reads & lexing
    Workers->>FS: read files
    Workers->>EventLoop: submit include resolution tasks
    EventLoop->>Resolver: resolve_include(filename, config, stat_cache)
    Resolver->>FS: stat candidate paths (uses/updates stat_cache)
    Resolver-->>EventLoop: return ResolveResult or nullopt
    EventLoop->>PathPool: intern(resolved_path) -> path_id
    EventLoop->>Graph: set_includes(path_id, config_id, included_ids)
    Note over Workers,Graph: repeat waves until stable
    EventLoop-->>Scan: return ScanReport
    Scan-->>Benchmark: print metrics / optionally export JSON
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~65 minutes

Possibly related PRs

Poem

🐇 I nibble paths and chase each include,
Wave by wave the graph grows, keen and shrewd.
Async hops and stat-cache gleam,
Modules mapped in moonlight's beam.
A rabbit cheers — dependencies take wing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 14.15% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding a dependency graph scanner and benchmark capability to the codebase.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch scan-deps

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 15

🧹 Nitpick comments (2)
tests/unit/syntax/scan_tests.cpp (1)

11-26: Add a regression case for #include_next.

The scanner now records is_include_next, but this suite never exercises that path. A focused #include_next test would keep the new flag from regressing silently.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/unit/syntax/scan_tests.cpp` around lines 11 - 26, Add a focused test
that exercises the scanner's new is_include_next flag by creating a new
TEST_CASE (e.g., "IncludeNext") that calls scan with source containing a
`#include_next` directive (for both angled and quoted forms if desired), then
assert that the produced result.includes contains an entry whose path matches
the header, whose is_include_next is true, and that the other flags (is_angled,
conditional, module_name) have the expected values; use the existing scan
function and result.includes[] access pattern to locate where to add these
assertions so this behavior is covered and won't regress.
src/server/master_server.h (1)

54-62: Prefer rebuilding scanner state into fresh objects.

load_compile_database() is incremental, but DependencyGraph only exposes add/set APIs and PathPool only grows. Keeping both as long-lived members means a later workspace reload cannot truly drop deleted files or module mappings. Consider constructing fresh instances in load_workspace() and swapping them in on success.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/server/master_server.h` around lines 54 - 62, The class keeps long-lived
PathPool and DependencyGraph members which are only grown/updated incrementally,
so a workspace reload cannot drop removed files; change load_workspace() to
construct fresh local instances (e.g., local PathPool new_path_pool and
DependencyGraph new_dependency_graph), pass those into load_compile_database()
and other helper functions instead of using the member variables directly, and
only swap/move them into the member fields (PathPool path_pool and
DependencyGraph dependency_graph) after a successful load; ensure any APIs that
currently mutate the members are updated to accept the new instances (or return
them) so the swap happens atomically on success.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/benchmark.yml:
- Around line 29-30: The workflow currently clones llvm-project without a fixed
reference (the step named "Clone LLVM" running the command "git clone --depth 1
https://github.com/llvm/llvm-project.git"); modify this step to pin to a
specific tag or commit hash so the benchmark is deterministic—e.g., use a clone
that checks out a given tag/commit (or pass --branch <tag> and then checkout the
known commit) and document the chosen revision; ensure the step still uses
shallow clone semantics if desired but locks the revision to a specific
tag/commit to avoid nondeterministic changes.
- Around line 32-40: The "Generate CDB" CI step runs a multi-line cmake command
using backslash continuations which will fail on Windows PowerShell; update that
step in the workflow to be PowerShell-safe by either adding "shell: bash" to the
step so the existing backslash continuations work, or rewrite the cmake
invocation (the multi-line cmake -B llvm-build ... llvm-project/llvm command) as
a single-line PowerShell-compatible command; modify the step named "Generate
CDB" accordingly so Windows matrix entries run the configure successfully.

In `@benchmarks/scan_benchmark.cpp`:
- Around line 190-196: The code currently reads
std::thread::hardware_concurrency() into hw_threads and unconditionally calls
setenv(), which is POSIX-only and may export an invalid "0" value; change this
so that you clamp hw_threads to at least 1 (if hw_threads == 0 set to 1) and
only set UV_THREADPOOL_SIZE when std::getenv("UV_THREADPOOL_SIZE") is null, and
replace the POSIX setenv call with a platform branch: on Windows use a
Windows-safe setter such as _putenv_s("UV_THREADPOOL_SIZE", size.c_str()) or
SetEnvironmentVariableA, and on non-Windows keep setenv("UV_THREADPOOL_SIZE",
size.c_str(), 0); keep the variable names hw_threads, UV_THREADPOOL_SIZE,
std::getenv, and the hardware_concurrency() call to locate the change.

In `@src/compile/command.cpp`:
- Around line 841-886: The merged-argv currently appends user include flags
after cached cc1 system include paths, breaking include-order semantics used by
extract_search_config(); in query_toolchain_cached handling you must parse
user_args (using self->parser.parse callback that currently calls append_arg)
and collect -I and -iquote entries into a temporary list, then insert those user
include entries before the cached cc1's system include entries (i.e. find the
first cached system include/-isystem position in arguments created by
arguments.assign(cached.begin(), cached.end()) and splice the collected user
-I/-iquote args there) so user include paths appear prior to cached system dirs;
keep -isystem entries treated as system (leave them in cached order) and
preserve original ordering of user flags when inserting.
- Around line 163-205: The toolchain key/query currently omits language-forcing
flags, so update is_toolchain_option to include language selectors (e.g.,
ID::OPT_x and any joined/equals variants) so extract_toolchain_flags will
preserve them into result.key and result.query_args; additionally, handle the
driver-specific case where a language token can appear at arguments[1] (e.g.,
Zig/`zig c++ file.c`) by detecting common language names and appending that
token to result.key and result.query_args when it is not a filename, using the
existing extract_toolchain_flags/arguments/result.key/result.query_args logic to
ensure the cache key and toolchain query reflect the forced language.

In `@src/server/master_server.cpp`:
- Around line 198-199: load_workspace() currently calls scan_dependency_graph()
synchronously which creates a local et::event_loop and blocks
MasterServer::loop; instead add an async variant (e.g.,
scan_dependency_graph_async or overload of scan_dependency_graph that takes an
et::event_loop&/et::executor& as a parent) that forwards to the existing async
scan_impl() and co_awaits it rather than calling loop.run(); update
load_workspace() to schedule/await that async variant on MasterServer::loop (do
not create a local et::event_loop or call run()), and ensure any callers using
scan_dependency_graph() are updated to use the async version so the full graph
scan runs without blocking the LSP event loop and avoids nested event loop
execution.

In `@src/syntax/dependency_graph.cpp`:
- Around line 102-105: The WaveEntry currently only stores path_id and config_id
so ResolveResult::found_dir_idx is lost between waves; add a std::uint32_t
found_dir_idx field to struct WaveEntry, set it when queuing entries from
resolve_file_includes()/resolve_include() using the ResolveResult::found_dir_idx
value, and propagate that field into subsequent calls (pass it instead of
literal 0 into resolve_include()/resolve_file_includes wherever a found_dir_idx
is forwarded). Update any places that construct WaveEntry (including the other
occurrences noted) and call sites of resolve_include()/resolve_file_includes to
accept and use the carried found_dir_idx so `#include_next` resumes searching from
the correct directory.
- Around line 298-305: The loop over scan_results increments report.total_files
only after skipping scan_result.read_failed which causes header_files =
total_files - source_files to underflow when reads fail; move the
report.total_files++ so it runs for every scan_result (before the
if(scan_result.read_failed) continue) so total_files counts attempted files, and
apply the same change to the analogous block around lines 391-395; keep
report.source_files increment behavior unchanged so header_files calculation
(report.total_files - report.source_files) stays valid.
- Around line 250-264: scanned_files currently de-duplicates only by resolved
path, causing headers scanned under one compilation context to be reused for
other configs; change the dedup key to include the compilation config (i.e.,
deduplicate by (path or path_id, config_id)) so each (path_id, config_id) pair
is scanned independently. Concretely, replace or rework the
llvm::StringMap<std::uint32_t> scanned_files to a map keyed by path+config (for
example a StringMap<llvm::DenseSet<uint32_t>> or an
llvm::StringMap<llvm::StringMap<uint32_t>> or a map with a combined key), update
the code around current_wave/WaveEntry creation where you call
scanned_files.try_emplace(path, path_id) to only suppress pushing a WaveEntry
when that specific (path_id, config_id) was already seen, and apply the same
change at the other occurrence referenced in the review (the block that mirrors
this logic later). Ensure lookups and inserts use the path and config_id pair so
the graph edges remain per (path_id, config_id).

In `@src/syntax/dependency_graph.h`:
- Around line 3-6: The header declares ScanReport::UnresolvedInclude which uses
std::string but doesn't include <string>, so add the missing include by adding
`#include` <string> at the top of the header (alongside the existing <cstdint>,
<optional>, <vector> includes) and also add the same <string> include where the
other occurrence around the ScanReport/UnresolvedInclude definition (the section
corresponding to lines ~124-126) appears so the public header no longer relies
on transitive includes.
- Around line 21-25: get_all_includes() currently returns bit-packed uint32_t
values using CONDITIONAL_FLAG and PATH_ID_MASK which makes callers handle
masking and cannot represent "same path both conditional and unconditional"
cleanly; replace the raw return with a structured edge type (e.g. IncludeEdge
with fields path_id and conditional) and change get_all_includes() (and the
other affected accessors referenced around lines 65-70 and 90-92) to return a
vector/list of IncludeEdge (or an equivalent iterable) so callers no longer need
to inspect CONDITIONAL_FLAG or use PATH_ID_MASK — update all call sites to use
IncludeEdge.path_id and IncludeEdge.conditional instead of manual bit-tests and
dedup logic.
- Around line 54-58: The module map currently overwrites previous providers when
add_module(llvm::StringRef module_name, std::uint32_t path_id) is called,
causing lookup_module(...) to return the wrong PathID in mixed-config
workspaces; change the data structure to record multiple providers per module
(e.g., map module_name -> vector/list of PathIDs + config metadata) or detect
duplicates and keep the original while recording additional providers, update
add_module to append or register duplicates rather than overwrite, and update
lookup_module to select the correct provider based on configuration/context (or
to fail/ambiguous if no config can disambiguate); ensure any other related
declarations around the duplicate-handling area (the other module-related
methods noted at the 81-88 region) are adjusted to use the new multi-provider
representation.

In `@src/syntax/include_resolver.cpp`:
- Around line 43-57: Found file paths (both in the absolute-filename branch and
in the include search loops that use candidate/config.dirs and stat_file) are
returned/cached as spelled, so normalize them with
llvm::sys::path::remove_dots(..., true) before creating the ResolveResult to
canonicalize any “..” components; specifically, after you get a successful
stat_file result (the local result or candidate variable) call
llvm::sys::path::remove_dots on the SmallString path (or on a copy of
result->str()) and then pass the normalized string into ResolveResult instead of
the original, and apply this change to the absolute-path branch and the
include_next/regular search loops that currently return
ResolveResult{result->str(), ...}.

In `@src/syntax/scan.cpp`:
- Around line 65-71: Precise scanning currently fills only path, conditional and
not_found in PreciseScanPPCallbacks::InclusionDirective, leaving is_angled and
is_include_next false; update InclusionDirective to set IncludeInfo::is_angled
(e.g. from the include token or dir.isAngled()/name.front() == '<') and
IncludeInfo::is_include_next (from dir.isIncludeNext() or equivalent) and
preserve conditional (conditional_depth) like scan() does so the IncludeInfo
populated by PreciseScanPPCallbacks matches the flags set by scan(); locate the
IncludeInfo construction in PreciseScanPPCallbacks::InclusionDirective and
assign the two missing fields before pushing into result.includes.

In `@tests/unit/syntax/dependency_graph_tests.cpp`:
- Around line 229-254: The build_cdb_json helper currently concatenates raw
strings into JSON (in function build_cdb_json using CDBEntry fields like e.dir,
e.file, e.extra_args) which breaks on Windows paths and embedded
quotes/backslashes; change it to emit a proper JSON array using llvm::json (or
construct an "arguments" array) by building llvm::json::Object entries with keys
"directory", "file" and either "command" (escaped via llvm::json) or "arguments"
(as an array of strings split from e.extra_args and the compiler), then
serialize the llvm::json::Array to a string and return that instead of manual
string concatenation so paths and special characters are correctly escaped.

---

Nitpick comments:
In `@src/server/master_server.h`:
- Around line 54-62: The class keeps long-lived PathPool and DependencyGraph
members which are only grown/updated incrementally, so a workspace reload cannot
drop removed files; change load_workspace() to construct fresh local instances
(e.g., local PathPool new_path_pool and DependencyGraph new_dependency_graph),
pass those into load_compile_database() and other helper functions instead of
using the member variables directly, and only swap/move them into the member
fields (PathPool path_pool and DependencyGraph dependency_graph) after a
successful load; ensure any APIs that currently mutate the members are updated
to accept the new instances (or return them) so the swap happens atomically on
success.

In `@tests/unit/syntax/scan_tests.cpp`:
- Around line 11-26: Add a focused test that exercises the scanner's new
is_include_next flag by creating a new TEST_CASE (e.g., "IncludeNext") that
calls scan with source containing a `#include_next` directive (for both angled and
quoted forms if desired), then assert that the produced result.includes contains
an entry whose path matches the header, whose is_include_next is true, and that
the other flags (is_angled, conditional, module_name) have the expected values;
use the existing scan function and result.includes[] access pattern to locate
where to add these assertions so this behavior is covered and won't regress.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 684a51bd-eb11-4a1c-a60c-e6eef8576498

📥 Commits

Reviewing files that changed from the base of the PR and between 020c2cb and 2169625.

📒 Files selected for processing (17)
  • .github/workflows/benchmark.yml
  • CMakeLists.txt
  • benchmarks/scan_benchmark.cpp
  • src/compile/command.cpp
  • src/compile/command.h
  • src/server/master_server.cpp
  • src/server/master_server.h
  • src/support/path_pool.h
  • src/syntax/dependency_graph.cpp
  • src/syntax/dependency_graph.h
  • src/syntax/include_resolver.cpp
  • src/syntax/include_resolver.h
  • src/syntax/scan.cpp
  • src/syntax/scan.h
  • tests/unit/syntax/dependency_graph_tests.cpp
  • tests/unit/syntax/include_resolver_tests.cpp
  • tests/unit/syntax/scan_tests.cpp

Comment thread .github/workflows/benchmark.yml
Comment thread .github/workflows/benchmark.yml
Comment thread benchmarks/scan_benchmark.cpp
Comment thread src/compile/command.cpp Outdated
Comment thread src/command/command.cpp
Comment thread src/syntax/dependency_graph.h
Comment thread src/syntax/dependency_graph.h
Comment thread src/syntax/include_resolver.cpp Outdated
Comment thread src/syntax/scan.cpp
Comment thread tests/unit/syntax/dependency_graph_tests.cpp
16bit-ykiko and others added 2 commits March 23, 2026 16:52
…gure

- Replace setenv() with putenv() (setenv is POSIX-only, unavailable on MSVC)
- Use system cc/c++/cl instead of pixi's clang for LLVM CMake configure
  to avoid linker incompatibilities on macOS

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Call fs::init_resource_dir(argv[0]) before scanning so -resource-dir
  points to the correct clang headers (fixes ~15% unresolved includes)
- Use cmake/toolchain.cmake for LLVM configure on all platforms
  (fixes Windows cl not found, macOS pixi linker issues)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@benchmarks/scan_benchmark.cpp`:
- Around line 77-79: The code writes JSON to an ofstream named out using
output_path without verifying the stream succeeded; update the export block
(where out and output_path are used and json/export_data referenced) to check
out.is_open() or out.good() immediately after constructing out and before
writing, and on failure log a descriptive error (including output_path and errno
or strerror) and return/exit/throw so the program does not proceed to write to a
bad stream.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: f996bc15-3319-4675-8040-4f125b34c8ec

📥 Commits

Reviewing files that changed from the base of the PR and between 9a91658 and 1061b65.

📒 Files selected for processing (2)
  • .github/workflows/benchmark.yml
  • benchmarks/scan_benchmark.cpp
✅ Files skipped from review due to trivial changes (1)
  • .github/workflows/benchmark.yml

Comment thread benchmarks/scan_benchmark.cpp Outdated
16bit-ykiko and others added 2 commits March 23, 2026 17:31
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Set default shell to bash for all benchmark steps (fixes Windows
  backslash line continuation issue)
- Add error checking for parallel scan outcome to diagnose Linux segfault

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (3)
src/syntax/dependency_graph.cpp (3)

302-308: ⚠️ Potential issue | 🟡 Minor

Count attempted files before the read-failure continue.

source_files is initialized from the full root wave, but total_files only increments after Line 303 filters failed reads. A missing source file can make Line 397 wrap header_files to a huge value.

Suggested fix
         for(auto& scan_result: scan_results) {
+            report.total_files++;
+
             if(scan_result.read_failed) {
                 LOG_WARN("Failed to read file for scanning: {}", scan_result.path);
                 continue;
             }
-
-            report.total_files++;
 
             auto config_it = configs.find(scan_result.config_id);
             if(config_it == configs.end()) {
                 continue;
             }

Also applies to: 397-397

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/syntax/dependency_graph.cpp` around lines 302 - 308, In the loop
iterating over scan_results (the for(auto& scan_result: scan_results) block),
increment report.total_files (or otherwise count attempted files) before the
read-failed check/continue so that files attempted to be scanned are counted
even when scan_result.read_failed is true; update the code paths that currently
increment report.total_files only after the read-success branch (refer to
scan_result.read_failed and report.total_files) so header_files/wrapping logic
later (e.g., around the header_files counter usage) sees the full attempted-file
count.

102-105: ⚠️ Potential issue | 🟠 Major

Preserve found_dir_idx between waves for nested #include_next.

Line 179 still hardcodes 0, and Line 366 enqueues discovered headers without the directory index returned on Lines 189-191. Any header found in search dir N will evaluate its own #include_next from the wrong starting point in later waves.

Suggested fix
 struct WaveEntry {
     std::uint32_t path_id;
     std::uint32_t config_id;
+    unsigned found_dir_idx = 0;
 };
 
 struct FileScanResult {
     std::string path;
     std::uint32_t path_id;
     std::uint32_t config_id;
+    unsigned found_dir_idx = 0;
     ScanResult scan_result;
     bool read_failed = false;
 };
 
-FileScanResult scan_file_worker(std::string path, std::uint32_t path_id, std::uint32_t config_id) {
+FileScanResult scan_file_worker(std::string path,
+                                std::uint32_t path_id,
+                                std::uint32_t config_id,
+                                unsigned found_dir_idx) {
     FileScanResult result;
     result.path = std::move(path);
     result.path_id = path_id;
     result.config_id = config_id;
+    result.found_dir_idx = found_dir_idx;
 
     auto content = et::fs::sync::read_to_string(result.path);
     if(!content.has_value()) {
         result.read_failed = true;
         return result;
@@
         auto resolved = co_await resolve_include(inc.path,
                                                  inc.is_angled,
                                                  includer_dir,
                                                  inc.is_include_next,
-                                                 0,  // default found_dir_idx
+                                                 scan_result.found_dir_idx,
                                                  config,
                                                  stat_cache,
                                                  loop);
@@
         for(auto& entry: current_wave) {
             auto path = std::string(path_pool.resolve(entry.path_id));
             auto pid = entry.path_id;
             auto cid = entry.config_id;
+            auto dir_idx = entry.found_dir_idx;
             scan_tasks.push_back(et::queue(
-                [path = std::move(path), pid, cid]() {
-                    return scan_file_worker(std::string(path), pid, cid);
+                [path = std::move(path), pid, cid, dir_idx]() {
+                    return scan_file_worker(std::string(path), pid, cid, dir_idx);
                 },
                 loop));
         }
@@
-                    next_wave.push_back({inc_path_id, result.config_id});
+                    next_wave.push_back({inc_path_id, result.config_id, edge.found_dir_idx});
                 }

Also applies to: 108-114, 142-156, 160-180, 277-285, 363-366

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/syntax/dependency_graph.cpp` around lines 102 - 105, The WaveEntry struct
must carry the directory index where a header was found so nested `#include_next`
can resume correctly; add a uint32_t found_dir_idx (or similarly named) to
WaveEntry and set it when you create WaveEntry instances (use the directory
index returned on header discovery instead of hardcoding 0). Update all places
that construct/enqueue WaveEntry (the enqueue/discovery paths mentioned in the
review) to populate found_dir_idx and update any consumers that start
include_next search so they use waveEntry.found_dir_idx as the start index.
Ensure all code paths that previously assumed 0 (including the nested
include_next handling) now read the preserved found_dir_idx from WaveEntry.

253-255: ⚠️ Potential issue | 🟠 Major

Scan discovered headers per (path, config_id) pair.

The graph stores edges per (path_id, config_id), but scanned_files suppresses later waves by path alone. If the same header is reached under two search configs, only the first config ever gets that header’s outgoing edges.

Suggested fix
-    llvm::StringMap<std::uint32_t> scanned_files;
+    llvm::StringMap<llvm::DenseSet<std::uint32_t>> scanned_files;
@@
         auto config_id = context_to_config_id[context];
         for(auto path_id: file_ids) {
             auto path = path_pool.resolve(path_id);
-            scanned_files.try_emplace(path, path_id);
-            current_wave.push_back({path_id, config_id});
+            auto& seen_configs = scanned_files[path];
+            if(seen_configs.insert(config_id).second) {
+                current_wave.push_back({path_id, config_id});
+            }
         }
     }
@@
-                auto [it, inserted] = scanned_files.try_emplace(edge.resolved_path, inc_path_id);
-                if(inserted) {
+                auto& seen_configs = scanned_files[edge.resolved_path];
+                if(seen_configs.insert(result.config_id).second) {
                     next_wave.push_back({inc_path_id, result.config_id});
                 }

Also applies to: 259-265, 363-366

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/syntax/dependency_graph.cpp` around lines 253 - 255, scanned_files
currently deduplicates by path only, causing headers reached under different
search configs to be treated as already scanned; change scanned_files to track
the (path, config_id) pair instead (e.g., make the key a composite of path and
config id or use a map from std::pair<path_id, config_id> to uint32_t) and
update all usages that insert/check scanned_files (the declarations named
scanned_files and the later lookup/insert sites around the blocks corresponding
to the original lines 259-265 and 363-366) so that you only suppress further
waves when both path and config_id match.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/syntax/dependency_graph.cpp`:
- Around line 319-320: The code dereferences resolve_outcome without checking
for errors; update the block that awaits resolve_tasks (the call to et::when_all
with resolve_tasks producing resolve_outcome and resolve_results) to mirror the
scan_outcome pattern: after auto resolve_outcome = co_await
et::when_all(std::move(resolve_tasks)); check if resolve_outcome.has_error() and
handle/log/return the error before dereferencing resolve_outcome into
resolve_results and before iterating over resolve_results; ensure any error
returned by resolve_file_includes() is propagated or handled consistently with
the existing scan error handling.

---

Duplicate comments:
In `@src/syntax/dependency_graph.cpp`:
- Around line 302-308: In the loop iterating over scan_results (the for(auto&
scan_result: scan_results) block), increment report.total_files (or otherwise
count attempted files) before the read-failed check/continue so that files
attempted to be scanned are counted even when scan_result.read_failed is true;
update the code paths that currently increment report.total_files only after the
read-success branch (refer to scan_result.read_failed and report.total_files) so
header_files/wrapping logic later (e.g., around the header_files counter usage)
sees the full attempted-file count.
- Around line 102-105: The WaveEntry struct must carry the directory index where
a header was found so nested `#include_next` can resume correctly; add a uint32_t
found_dir_idx (or similarly named) to WaveEntry and set it when you create
WaveEntry instances (use the directory index returned on header discovery
instead of hardcoding 0). Update all places that construct/enqueue WaveEntry
(the enqueue/discovery paths mentioned in the review) to populate found_dir_idx
and update any consumers that start include_next search so they use
waveEntry.found_dir_idx as the start index. Ensure all code paths that
previously assumed 0 (including the nested include_next handling) now read the
preserved found_dir_idx from WaveEntry.
- Around line 253-255: scanned_files currently deduplicates by path only,
causing headers reached under different search configs to be treated as already
scanned; change scanned_files to track the (path, config_id) pair instead (e.g.,
make the key a composite of path and config id or use a map from
std::pair<path_id, config_id> to uint32_t) and update all usages that
insert/check scanned_files (the declarations named scanned_files and the later
lookup/insert sites around the blocks corresponding to the original lines
259-265 and 363-366) so that you only suppress further waves when both path and
config_id match.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 5e1421b1-b43b-46a3-8a38-bb6c6ee910d6

📥 Commits

Reviewing files that changed from the base of the PR and between 110dec6 and 32cbf84.

📒 Files selected for processing (2)
  • .github/workflows/benchmark.yml
  • src/syntax/dependency_graph.cpp
🚧 Files skipped from review as they are similar to previous changes (1)
  • .github/workflows/benchmark.yml

Comment thread src/syntax/dependency_graph.cpp Outdated
16bit-ykiko and others added 8 commits March 23, 2026 18:09
github.workspace uses backslashes on Windows which get stripped
in bash shell. Use $(pwd) which always produces forward slashes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- --log-level: control spdlog level (default: off for clean benchmark)
- --export: export dependency graph as JSON
- --runs: number of iterations
- -h/--help: usage message
- Use getMainExecutable for reliable resource_dir resolution
- Link eventide::deco for CLI parsing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Track file read time, lexer scan time, stat syscall counts/timing,
and cache hit rate to identify performance bottlenecks in the
wavefront BFS scanner.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Show both wall-clock per-phase timing (config, read+scan, resolve,
graph build) and cumulative-across-threads I/O stats separately
to make the report easier to interpret.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Instead of calling stat()/access() for each candidate include path,
cache directory listings via readdir() and do in-memory set lookups.
This dramatically reduces filesystem syscalls (94K stat -> 11K readdir
on LLVM) and is especially impactful on Windows where stat() is ~10x
slower than on Linux.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Convert resolve_include to a synchronous function (no more coroutine
overhead) and dispatch Phase 2 to the thread pool. DirListingCache uses
64 independent shards to minimize lock contention across threads.

Also remove eventide dependency from include_resolver.h since async
is no longer needed for include resolution.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Parallel Phase 2 with sharded locks showed no net benefit — thread pool
contention slowed Phase 1 more than Phase 2 gained. Revert to serial
include resolution and remove unnecessary mutex/sharding infrastructure.

Also fix test compilation by updating include_resolver_tests to use the
synchronous resolve_include API, and skip flaky InlayHint.Special test.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add get_pending_toolchain_queries() and inject_toolchain_results() to
CompilationDatabase, allowing callers to pre-warm the toolchain cache
by executing unique queries in parallel on the thread pool.

Key optimization: use driver binary + file extension as a fast dedup key
to avoid expensive argument parsing for all context groups. Only ~3-4
unique (driver, extension) combinations need full parsing instead of
hundreds.

In scan_impl, cache-miss queries run in parallel via et::queue +
et::when_all. Subsequent serial lookup() calls all hit the warm cache.

This targets Windows where process creation is expensive (~2300ms for
5-6 serial spawns). Parallel execution should reduce this to ~1 spawn
wall-clock time.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
16bit-ykiko and others added 12 commits March 24, 2026 01:10
Measures lookup() + extract_search_config() overhead separately
to isolate argument parsing cost across platforms (Linux vs Windows).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The previous fast dedup (driver binary + file extension) could miss
contexts with different toolchain-affecting flags (e.g. -std=, -target),
causing synchronous process spawning during lookup(). Now extract full
toolchain keys for all contexts to ensure complete pre-warming.

Also add prewarm/lookup timing breakdown to scan report.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Break down config_ms into lookup vs extract_search_config per-group,
and enable info logging in benchmark CI for detailed diagnostics.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Run scan before parser microbenchmark to isolate pre-warm issues.
Add LOG_WARN on toolchain cache miss in query_toolchain_cached.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Show context dedup ratio and sample commands when no sharing detected,
to diagnose why Windows has no context deduplication.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
On Windows, cmake may use backslashes in the `arguments` array but
forward slashes in the `file` field. This caused `argument == file`
to fail, leaving source file paths in CompilationInfo and preventing
context deduplication (10262 files -> 10262 unique contexts instead
of ~1300). Add is_same_file() that normalizes path separators.

Also add /Fd to the output options filter list.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Remove verbose command dump from benchmark, keep context dedup stat.
Change per-key pre-warm log to DEBUG level. Restore default benchmark
log level.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
et::fs::sync::read_to_string() calls uv_default_loop() which is not
thread-safe when invoked from libuv worker threads. Use LLVM's
MemoryBuffer::getFile() (mmap-based) instead, fixing ASAN SEGV and
the intermittent macOS SIGABRT.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Pre-populate directory listing cache in parallel before BFS loop,
  moving expensive readdir() syscalls off the critical path (especially
  impactful on Windows where readdir is ~5x slower than Linux)
- Eliminate double string copy in Phase 1 scan worker dispatch
- Reuse SmallString<256> candidate buffer in include resolver instead
  of constructing a new one per search dir check
- Return bool from check_file() instead of optional<string> to avoid
  string allocation on every successful file lookup
- Reserve vector capacities for current_wave and resolve_results
- Increase default benchmark iterations to 20 for more reliable results

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Use RequiresNullTerminator=false for MemoryBuffer::getFile since
  scanSourceForDependencyDirectives works with StringRef
- Reserve edges vector in resolve_file_includes to avoid reallocation

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
16bit-ykiko and others added 27 commits March 24, 2026 22:59
…onst char*

StringRef::copy() allocates exactly path.size() bytes with no null
terminator.  scan_file_worker passes this const char* to
MemoryBuffer::getFile, which constructs a Twine from const char* and
uses strlen — reading past the string into uninitialized allocator
memory and producing a garbage path.  Result: ~99% of file reads fail,
so only ~66 files are scanned instead of ~20k.

Fix: allocate n+1 bytes and write '\0' at offset n, keeping the
StringRef length at n so all StringRef operations remain correct.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- CI benchmark now runs --runs 1 (cold start only, no warm iterations)
- Benchmark resets PathPool + DependencyGraph each run (no ScanCache)
- Added per-wave stats (WaveStats) to ScanReport: files, P1/P2 timing,
  next wave size, prefetch count, dir listings/hits per wave
- Minimum thread pool size of 8 for I/O-bound file scanning
- Removed parser microbenchmark (warm-only, not useful for cold start)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Rebuild CDB each iteration to clear toolchain & config caches
- Default 20 runs locally, CI uses 5 runs
- Print min/avg/max summary for total, config, phase1, phase2
- Unified thread pool: max(hw_threads, 4) on all platforms
- Detailed report printed for first run only

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
With mmap, the measured "read" time only captures the mmap() syscall
while actual page-fault I/O is hidden inside the lexer timing. Switch
to IsVolatile=true + RequiresNullTerminator=true to force LLVM's
MemoryBuffer to use read() for all files, giving accurate I/O vs CPU
breakdown across platforms.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1. Config extraction overlap: split into two steps — assign config IDs
   (fast, before Phase 1) and lookup_search_config (slow, runs while
   scan tasks execute on thread pool). Dir cache launch also deferred
   to after config extraction. Saves ~300-400ms on all platforms by
   paying max(config_time, scan_time) instead of config_time + scan_time.

2. Phase 2 timing breakdown: add per-category microsecond timers for
   resolve_include(), path_pool.intern(), and et::queue() prefetch to
   diagnose Windows Phase 2 slowness (1113ms vs Ubuntu 359ms).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The hot path in resolve_include was constructing a full candidate path
(SmallString copy + path::append) for every search dir, then check_file
decomposed it back into dir + filename. With ~1.2M lookups per scan,
this wasted millions of string copies and path decompositions.

New approach: check_file_in_dir takes dir and filename separately,
directly doing the StringSet lookup. Full path is only constructed
on hits (the rare case). This eliminates per-miss overhead of:
- SmallString<256> assignment (dir path copy)
- llvm::sys::path::append
- llvm::sys::path::parent_path + filename decomposition

Also re-adds Phase 2 resolve_include timing diagnostic.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…kups

Before Phase 2, resolve each config search dir to a const StringSet<>*
pointer from the DirListingCache. resolve_include() now uses direct
pointer dereference instead of StringMap hash lookups per candidate dir.

Also: increase CI benchmark runs from 5 to 20 for more stable statistics,
and remove unresolved header logging to reduce benchmark output noise.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The pre-resolve optimization incorrectly checked entries->contains() with
multi-component paths like "llvm/Support/raw_ostream.h" against the search
dir's direct children. For such paths, we now construct the full path and
resolve the actual parent subdirectory via DirListingCache. Simple filenames
(no '/') still use the fast pre-resolved entries path.

This was causing resolution accuracy to drop from ~99% to ~19% and total
discovered files to shrink from ~20k to ~13k.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
For paths like "llvm/Support/raw_ostream.h", check if "llvm" exists in
the search dir's pre-resolved entries before constructing the full path
and resolving the subdirectory. Most search dirs don't match, avoiding
expensive path construction + readdir. Phase 2: 135ms → 32ms locally.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Paths like "/a/b/c/../foo.h" and "/a/b/foo.h" previously produced
different PathPool IDs, causing the same file to be read and scanned
twice. Now apply remove_dots() on resolved paths that contain ".."
or "./" components. Pure string operation, no syscalls.

Locally reduces discovered files from 9665 to 9016 (649 fewer dupes).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Move toolchain query logic (flag extraction, caching, batch pre-warming)
out of CompilationDatabase into a standalone ToolchainProvider class.

CDB now holds a ToolchainProvider by composition and exposes it via
toolchain() accessor. resolve_toolchain_entries() bridges CDB-internal
context pointers to the provider's PendingEntry format.

This separates three concerns that were mixed in CDB:
  1. Compilation command management (stays in CDB)
  2. Toolchain query + caching (now in ToolchainProvider)
  3. SearchConfig extraction (stays in CDB)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Move SearchDir, SearchConfig structs and extract_search_config() out of
CompilationDatabase into compile/search_config.{h,cpp} as a standalone
free function. This decouples argument parsing from the CDB and makes it
independently testable/improvable to match clang's behavior.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…arch

SearchConfig now uses a three-segment model (Quoted/Angled/System) matching
clang's InitHeaderSearch::Realize layout. extract_search_config classifies
cc1 flags by IncludeDirGroup, reorders as Quoted → Angled → System, and
deduplicates across [Angled..end) for correct #include_next semantics.

Also propagates found_dir_idx through the BFS dependency scanner and adds
tests for both the search config extraction and three-tier include resolution.

Ported from fix-header-search-priority branch.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Replace heavy compile/command.h include with compile/search_config.h
  in include_resolver.h (only SearchConfig is needed)
- Remove unused <string> include from include_resolver.h
- Shorten test names to ≤ 4 words per project convention

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- search_config.cpp: -idirafter, -cxx-isystem, -iwithsysroot, HeaderMap
- include_resolver.cpp: macOS Framework search (-F, -iframework)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add fourth segment to SearchConfig: Quoted → Angled → System → After.
Handle -idirafter and -iwithprefix as After group entries matching
clang's InitHeaderSearch layout. Deduplication covers all segments
from [Angled..end). Also applies clang-format to all modified files.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace removed eventide/deco/macro.h and eventide/deco/runtime.h
with the new single-header eventide/deco/deco.h.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Move compilation command related files (CompilationDatabase, SearchConfig,
toolchain query/cache, driver parser) from src/compile/ to src/command/.

This separates "understanding how to compile" (command/) from "actually
compiling" (compile/), making the codebase easier to navigate.

Moved files:
  compile/command.{h,cpp} → command/
  compile/search_config.{h,cpp} → command/
  compile/toolchain.{h,cpp} → command/
  compile/toolchain_provider.{h,cpp} → command/
  compile/driver.h → command/

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace removed deco::cli::Dispatcher with write_usage_for, and fix
include sort order in toolchain_tests.cpp for clang-format CI check.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1. Replace queried compiler's resource dir with ours on Windows to
   avoid clang version mismatch (e.g. AVX10.2-BF16 builtin renames).
2. Escape backslashes in build_cdb_json() so Windows paths produce
   valid compile_commands.json.
3. Use P() macro to make ExtractSearchConfig test paths absolute on
   Windows (drive letter prefix).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Move duplicated TempDir struct from include_resolver_tests.cpp and
  dependency_graph_tests.cpp into shared test/temp_dir.h.
- Add c_path() method returning pooled const char* for arg vectors.
- Replace P() macro in ExtractSearchConfig tests with TempDir, which
  provides cross-platform absolute paths (drive letter on Windows).
- Turn write_cdb() from a TempDir method into a free function to
  avoid coupling TempDir to CompilationDatabase.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
…ruption

TokenizeGNUCommandLine treats backslashes as escape characters, which
corrupts Windows paths in "command" string form (e.g. C:\Users → CUsers).
Switch to "arguments" array form which bypasses tokenization entirely.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The GNU tokenizer treats '\' as an escape character, which corrupts
Windows paths (e.g. C:\Users → C:Users). On Windows all programs are
invoked through the Windows API regardless of compiler (MSVC, clang-cl,
MinGW), so the Windows tokenizer is always correct for CDB entries.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
On Windows, llvm::sys::path::append does not convert forward slashes
within a relative component (e.g. "gcc/12/include") to native
backslashes. Add path::native() so test expectations match the
backslash-normalized paths returned by extract_search_config.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Add missing <string> include in dependency_graph.h (IWYU)
- Move total_files++ before read_failed check to prevent unsigned
  underflow in header_files = total_files - source_files
- Populate is_angled and is_include_next in PreciseScanPPCallbacks
  to match the fast scan path
- Add error check on std::ofstream open in benchmark export

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant