diff --git a/docs/en/design/command-resolve.md b/docs/en/design/command-resolve.md new file mode 100644 index 000000000..41ac99db5 --- /dev/null +++ b/docs/en/design/command-resolve.md @@ -0,0 +1,171 @@ +# Command Resolution + +## Background + +A C++ language server needs to know how to compile every file in a project. This information comes from the compilation database (CDB), typically a `compile_commands.json` file generated by the build system. Each entry in the CDB contains a source file path and a compilation command, for example: + +```bash +g++ -std=c++20 -O2 -fPIC -I../include -DNDEBUG -c src/foo.cpp +``` + +This command records how the build system invokes the compiler during an actual build. However, the language server cannot use it directly, for three reasons. + +**First, this is a driver-level command, not a frontend command.** The `g++` above is a compiler driver -- it selects the correct frontend, linker, and standard library. The language server actually uses Clang's frontend (cc1), and converting from `g++` to cc1 involves a large amount of implicit work: determining the target triple, injecting system header search paths, setting default language standards, and more. None of this information appears explicitly in the CDB -- it is implicitly provided by the compiler. + +**Second, the command contains semantically irrelevant options.** `-O2` and `-fPIC` only affect code generation, not semantic analysis -- the language server does not do code generation and has no use for these options. Similarly, `-c` (compile mode) and `-o` (output file) are build-artifact directives that are meaningless to a language server. Passing them unfiltered to the frontend adds unnecessary complexity and can even cause errors. + +**Third, commands in large projects are highly redundant.** In a project with tens of thousands of source files, the vast majority use the same compiler and semantic options, differing only in include paths and macro definitions. Without deduplication, each file requires an independent toolchain query, wasting memory and startup time. + +Of these three problems, the first has the greatest impact on user experience: missing implicit information. When the language server cannot correctly obtain system header paths, users see errors on standard library headers -- `#include ` reports "file not found," or GCC built-in type traits are flagged as undeclared identifiers. In clangd's issue tracker, these problems are consistently among the most frequently reported ([clangd#1262](https://github.com/clangd/clangd/issues/1262), [clangd#1691](https://github.com/clangd/clangd/issues/1691)). + +clangd's solution is the `--query-driver` flag: users manually specify which compilers need probing, and clangd then queries those compilers for system paths. This approach has two problems. First, it is manual -- users must know which compiler their project uses and correctly configure a glob pattern. Second, the flag is not independently useful -- it relies on existing CDB commands to trigger probing ([clangd#1219](https://github.com/clangd/clangd/issues/1219)). For cross-compilation and embedded development scenarios, users frequently need to debug their configuration repeatedly before clangd correctly identifies the toolchain. + +clice automates the entire command processing pipeline: after reading commands from the CDB, it automatically identifies the compiler family (GCC, Clang, MSVC, etc.), automatically probes the toolchain, and minimizes startup overhead through multi-level deduplication. Users do not need to manually configure any toolchain-related parameters. + +## Design + +The core task of command processing is to convert the raw driver commands in the CDB into cc1 arguments consumable by the Clang frontend. This conversion involves four conceptual layers: argument classification, command separation, toolchain probing, and search path extraction. + +### Argument Classification + +When loading the CDB, each compilation option is classified into one of four categories: + +- **Discarded**: Options related to build artifacts that the language server does not need. These include output file (`-o`), compile mode (`-c`), dependency scanning (`-M` family), PCH building (`-emit-pch`), and C++20 modules (`-fmodule-file`, etc., managed by the language server itself). + +- **Codegen-only**: Options that only affect the code generation backend and do not affect semantic analysis. These include position-independent code (`-fPIC`), stack protection (`-fstack-protector`), frame pointer (`-fomit-frame-pointer`), debug info (`-g` family), LTO, etc. They do not change the AST or diagnostic output. + + > Note: `-O` and `-fsanitize=` are not in this category despite appearing code-generation-related. `-O` defines the `__OPTIMIZE__` macro, and `-fsanitize=address` affects `__has_feature(address_sanitizer)`. They alter preprocessor state and are therefore semantic options. + +- **User-content**: Options that may differ per file but do not affect toolchain probing results. These include include paths (`-I`, `-isystem`, `-iquote`, `-idirafter`), macro definitions (`-D`, `-U`), and forced includes (`-include`). Their meaning is "what this particular file additionally needs." + +- **Semantic**: All remaining options -- they affect compilation semantics and play a role in toolchain probing. Examples include `-std=c++20`, `-Wall`, `-target`, `-march=`, etc. + +Classification is based on Clang's own option table (`OptTable`), using option IDs rather than string matching. + +### Command Separation + +After classification, each compilation command is split into two parts: + +- **`CanonicalCommand`**: Driver path + all semantic options. Represents "the identity and semantic configuration of the compiler." +- **Patch**: All user-content options. Represents "the include paths and macro definitions this file additionally needs." + +Together with the working directory, these form `CompilationInfo` -- an abstract representation of a file's complete compilation configuration. + +The core purpose of this separation is toolchain probing cache efficiency. Toolchain probing requires actually invoking the compiler driver (e.g., running `g++ -dumpmachine` or `clang++ -###`), typically taking 100ms or more. The probing result depends only on the driver and semantic options -- user-content options (`-I`, `-D`) do not affect the system paths or target triple output by the driver. Therefore, regardless of what different `-I` paths files may have, as long as the semantic options are the same, they can share the same probing result. + +In real projects, tens of thousands of files may have only a few dozen distinct `CanonicalCommand` instances, meaning the toolchain needs to be probed only a few dozen times rather than tens of thousands. + +Both `CanonicalCommand` and `CompilationInfo` are deduplicated via `ObjectSet` -- instances with identical content exist only once in memory and are shared by pointer. String arguments are interned through `StringSet`, ensuring pointer stability and enabling direct comparison. + +### Compilation Database + +`CompilationDatabase` loads `compile_commands.json`, parsing, classifying, and deduplicating each entry into `CompilationEntry` (file path ID -> `CompilationInfo`). All entries are sorted by file path ID for binary search lookups. + +On lookup, `CompilationDatabase` assembles a `CompilationInfo` into a `CompileCommand` -- the final output of the command processing pipeline, containing the complete compilation flags and source file path, ready to be submitted to toolchain probing or the Clang frontend. + +For files without a CDB entry (e.g., a file the user opens that is not part of the project), `CompilationDatabase` synthesizes a default command -- selecting `clang` or `clang++ -std=c++20` based on the file extension. + +`CompilationDatabase` also provides the ability to group by configuration: `ConfigGroup` aggregates files that share the same `CompilationInfo`. This is the right granularity for extracting search path configurations during dependency scanning -- different `-I` paths produce different groups. For toolchain probing, the granularity is coarser (user-content options don't affect probing results), so `Toolchain` further deduplicates on top of `ConfigGroup`. + +### Configuration Rules + +Beyond the compilation commands in the CDB itself, users can append or remove compilation options via `[[rules]]` in `clice.toml`. Each rule contains file matching patterns (globs) and lists of options to append or remove. + +When looking up a file's compilation command, matching rules are applied on top of the CDB command -- specified options are first removed from the base command, then new options are appended. This allows users to fine-tune compilation flags at the project level without modifying the build system's output. + +### Toolchain + +`Toolchain` converts driver-level commands into cc1 arguments. Its design centers on two core capabilities: + +**Compiler family identification.** `Toolchain` identifies the compiler family from the executable name -- the `CompilerFamily` enum includes GCC, Clang, MSVC, ClangCL, NVCC, Intel, and Zig. The identification logic handles various naming variants: version suffixes (`clang++-17`), architecture prefixes (`arm-none-eabi-g++`), and Windows `.exe` suffixes. The family determines which probing strategy is used. + +**Caching strategy.** Probing results are cached by (driver path, file extension, non-user-content flags). File extension is part of the cache key because `.c` and `.cpp` may trigger different driver rules. Failed probes are also cached (negative caching) to avoid retrying the same nonexistent compiler repeatedly. + +### Search Paths + +`SearchConfig` extracts header search paths from cc1 arguments, organizing them into a four-tier structure: + +1. **Quoted** (`-iquote`): Search paths for `#include "foo.h"` +2. **Angled** (`-I`): Search paths for `#include ` +3. **System** (`-isystem`, `-internal-isystem`, etc.): System header paths +4. **After** (`-idirafter`): Paths searched after system directories + +This four-tier model corresponds to Clang's internal search layout. Paths within tiers are deduplicated (starting from the Angled tier), with the deduplication algorithm replicating Clang's behavior: if the same path appears in both Angled and System tiers, the one in Angled is kept. This ensures `#include_next` correctness. + +`SearchConfig` is the foundational input for include path completion, include path resolution, [dependency graph construction](dependency-scanning.md), and other features. + +## Implementation + +### Loading and Parsing + +CDB loading uses simdjson for streaming JSON parsing, processing entries one by one: + +1. Read each entry's `directory`, `file`, and `arguments` (or `command`) fields +2. Filter out non-C/C++ files (e.g., `.rc`, `.asm`, `.def`) +3. Resolve relative file paths to absolute paths +4. Classify each option in the `arguments` field, routing them into canonical or patch +5. Absolutize relative paths in include path options (resolved against `directory`) +6. Deduplicate `CanonicalCommand` and `CompilationInfo` via `ObjectSet` +7. Sort all entries by file path ID + +The parsing process also handles a special case: CMake-generated CDBs sometimes contain `-Xclang -include-pch -Xclang ` sequences (CMake's PCH workaround), which are detected and discarded during loading. + +### Toolchain Probing + +Different compiler families use different probing strategies: + +**GCC**: A two-step process. First, call the GCC driver to obtain two key pieces of information -- the target triple (`-dumpmachine`) and the install path (`-print-search-dirs`). Second, inject this information into Clang's driver (`--target=` and `--gcc-install-dir=`), letting Clang's driver emulate GCC's behavior and produce cc1 arguments. This allows the Clang frontend to correctly locate GCC's standard library and system headers. + +**Clang / Zig**: Invoke the driver with the `-###` option, which prints the complete cc1 command line without actually executing compilation. Parse the first cc1 line from the output. For Zig, the driver path consists of two parts (`zig cc` or `zig c++`), requiring special handling during probing. + +> The external driver's version may be newer than clice's embedded LLVM version, and its cc1 output may contain options that clice does not recognize. During parsing, unknown options are silently dropped to ensure compatibility. + +**MSVC / ClangCL**: Switch Clang's driver to MSVC-compatible mode via the `--driver-mode=cl` directive, then use Clang's driver to obtain cc1 arguments. + +After probing completes, the temporary probe file's path and module-output-related flags (which reference the deleted temp file) are stripped from the results. If clice's resource dir differs from what the probe returned, all related paths are replaced to ensure the frontend uses builtin headers from a matching version. + +**Startup warm-up.** During the dependency scanning phase at server startup, all unique toolchain cache keys are collected and probed in parallel. Pipe draining and process waiting for probe subprocesses also run concurrently to avoid deadlocks from full pipes. Once warm-up completes, all subsequent toolchain queries hit the cache. + +### Search Path Extraction + +After toolchain probing yields cc1 arguments, search path extraction iterates through these arguments, routing include path options by type into the four tiers. All paths are resolved to absolute paths and normalized (removing `.` and `..`). + +The `-iprefix` / `-iwithprefix` / `-iwithprefixbefore` options must be processed cooperatively in order of appearance: `-iprefix` sets a prefix, subsequent `-iwithprefix` prepends the prefix to its path and places it in the After tier, and `-iwithprefixbefore` places it in the Angled tier. + +After the four tiers are concatenated, deduplication begins from the Angled tier. The Quoted tier does not participate in deduplication -- the same path appearing in both Quoted and Angled is legitimate, and both are preserved. This replicates Clang's internal behavior. + +## FAQ + +- **Why classify by option ID rather than string matching?** + + clice's argument parsing is based on Clang's own option table (generated from `Options.inc`), using option IDs for classification. String matching is prone to missing edge cases: Clang's option syntax has many forms -- `-std=c++20` is joined form, `-I /path` is separate form, `-Wall` is flag form, and some options have `/` prefixes (MSVC style). Classifying by ID means all these syntax variants are correctly handled by Clang's parser, and the classification logic only needs to care about "what is this option," not "how is it spelled." + +- **Why don't user-content options participate in the toolchain cache key?** + + This is the core benefit of the two-level separation design. `-I` and `-D` do not change the system paths, target triple, or language defaults output by the compiler driver. Excluding them from the cache key reduces the number of keys from "one per file" to "one per configuration." A project with tens of thousands of files typically has only a few dozen distinct cache keys, requiring only a few dozen subprocess calls at startup. + +- **Why must search path deduplication precisely match Clang's behavior?** + + If the language server's header search order differs from the actual compiler's, `#include` may resolve to different files -- when identically named headers exist in different directories, search order determines which one is used. The semantics of `#include_next` depend even more directly on the deduplication result of search directories. Strict matching ensures that the language server sees exactly the same code as the compiler. + +- **Why does GCC probing use a two-step process instead of directly parsing GCC's `-v` output?** + + The Clang frontend needs GCC's target triple and install path to correctly locate GCC's standard library headers. Directly parsing GCC's `-v` output is feasible but would introduce additional text parsing logic that is fragile across different GCC versions with varying output formats. By injecting the information into Clang's driver via `--target=` and `--gcc-install-dir=`, Clang itself handles the search path assembly, which is more reliable. + +- **Why is compiler family identification based on executable filename rather than path resolution?** + + Compiler driver behavior is affected by the name used to invoke it. For example, `/usr/bin/clang++` is typically a symlink to `/usr/lib/llvm-20/bin/clang`, but invoking it as `clang++` automatically enables C++ mode and links C++ libraries. Using `realpath` to resolve to the actual path before identification would lose the semantic information carried by the invocation name. Similarly, if `arm-none-eabi-g++` were resolved to some generic GCC binary path, the cross-compilation context would be lost. + +- **How are files without a CDB entry handled?** + + A default command is synthesized. Based on the file extension, either `clang` or `clang++ -std=c++20` is selected, and the resource dir is injected. This ensures basic semantic analysis remains available even when a file is not in the CDB. For header files, the system also attempts to find a source file that includes it via the dependency graph and uses that source file's compilation command as context (see [Compilation Context](compilation-context.md)). + +## Known Limitations + +- **Incomplete support for some compiler families.** NVCC and Intel compilers (`icc`, `icx`, `dpcpp`) are recognized but currently fall through to the generic Clang driver path without dedicated probing logic. This means these compilers' special system paths may not be correctly discovered. + +- **SearchConfig does not support all search path options.** `-cxx-isystem` (system directories effective only in C++ mode), `-iwithsysroot` (prepends sysroot to path), and HeaderMap support are not yet implemented. These options are uncommon in practice but may appear in specific Apple or cross-compilation toolchains. + +- **Global impact of configuration rules.** `[[rules]]` in `clice.toml` can append or remove options from compilation commands. If the user modifies a rule that affects all files (e.g., appending a global `-I`), all files' compilation configurations change, potentially triggering a full re-index. There is currently no mechanism to detect which rule changes actually affect which files. + +- **MSVC-style option parsing.** On non-Windows systems, MSVC-style option prefixes (`/U`, `/D`, `/I`) must be handled specially to prevent Unix absolute paths (such as `/Users/...`) from being misparsed as MSVC options. This is currently resolved by dynamically adjusting option visibility based on the driver name, but edge cases may still exist. diff --git a/docs/en/design/command.md b/docs/en/design/command.md deleted file mode 100644 index 91c10e924..000000000 --- a/docs/en/design/command.md +++ /dev/null @@ -1,97 +0,0 @@ -# Command Processing - -## Background - -A language server needs to know how to compile every file. This information comes from the compilation database (`compile_commands.json`), generated by build systems (CMake, Bazel, Meson, etc.). However, the raw compilation commands in the CDB cannot be fed directly to the Clang frontend for several reasons: - -**Compilation commands are driver-level, not frontend-level.** The CDB records the commands that the build system uses to invoke the compiler (e.g., `g++ -std=c++17 -O2 -c foo.cpp`), but the language server needs Clang frontend (cc1) arguments. Converting from driver commands to cc1 arguments requires querying the compiler toolchain to obtain the target triple, system header search paths, and other information. - -**Compilation commands contain semantically irrelevant options.** Build systems pass options that only affect code generation (e.g., `-fPIC`, `-fomit-frame-pointer`), which do not affect semantic analysis but add complexity. They also include build-artifact-related options (e.g., `-o foo.o`, `-emit-pch`), which are entirely meaningless to a language server. - -**Compilation commands lack implicit information.** System header paths, default language standards, compiler built-in macro definitions, and similar details do not appear explicitly in the CDB -- they are implicitly provided by the compiler toolchain. The language server must probe the toolchain to fill in this information. - -**Large projects have highly redundant compilation commands.** In a project with tens of thousands of files, the vast majority share the same compiler and semantic options (`-std=c++17`, `-Wall`, etc.), differing only in include paths and macro definitions. Without deduplication, this wastes significant memory and startup time. In clangd issues, users have reported Bazel-generated `compile_commands.json` files exceeding 10GB, with individual entries reaching 250KB. - -## Design - -### Overall Pipeline - -Command processing is a multi-stage pipeline: - -```text -compile_commands.json (raw commands) - | load and parse -Argument classification (codegen-only / discarded / user-content / semantic) - | two-level separation -CanonicalCommand (shared semantic options) + Patch (per-file user content) - | toolchain probing -cc1 arguments (complete frontend-consumable compilation command) - | search path extraction -SearchConfig (four-tier header search paths) -``` - -### Argument Classification - -When loading the CDB, each compilation option is classified into one of four categories: - -**Discarded**: Options related to build artifacts that the language server does not need. Examples include `-o` (output file), `-c` (compilation mode), `-M` (dependency scanning), `-emit-pch` (PCH building), etc. These are simply dropped. - -**Codegen-only**: Options that only affect the code generation backend and do not affect semantic analysis. Examples include `-fPIC`, `-fomit-frame-pointer`, `-funwind-tables`, debug info option groups (`-g*`), etc. These do not change the AST or diagnostic output and are dropped. Note that `-O` and `-fsanitize=` may appear to be code-generation-related, but they define macros (such as `__OPTIMIZE__` and `__has_feature(address_sanitizer)`), so they are retained. - -**User-content**: Options that may differ per file but do not affect toolchain probing results. Primarily include paths (`-I`, `-isystem`, `-iquote`, `-idirafter`) and macro definitions (`-D`, `-U`). These are extracted as patches and reattached after toolchain probing. - -**Semantic**: All remaining options. These affect compilation semantics and play a role in toolchain probing. Examples include `-std=c++17`, `-Wall`, `-target`, `-march=`, etc. - -### Two-level Separation: Canonical and Patch - -After classification, the compilation command is split into two parts: - -- **CanonicalCommand**: The driver name + all semantic options. Represents "the semantic aspects of how to invoke the compiler." -- **Patch**: All user-content options. Represents "what this particular file additionally needs." - -**Why this separation?** - -The core reason is **toolchain probing cache efficiency**. Toolchain probing requires actually invoking the compiler driver (e.g., `g++ -v`), which is an expensive operation (typically 100ms+). The probing result depends only on the driver and semantic options -- user-content options (`-I`, `-D`) do not affect the cc1 arguments output by the driver. Therefore, as long as the semantic options are the same, regardless of what different `-I` paths files may have, they can share the same probing result. - -In real projects, tens of thousands of files may have only 5-50 distinct CanonicalCommands. This means the toolchain only needs to be probed 5-50 times instead of tens of thousands. - -This also enables memory deduplication of compilation commands. Identical CanonicalCommand and Patch combinations are merged into the same CompilationInfo instance and shared via pointers. - -### Path Absolutization - -During the loading phase, all include path options in user content are absolutized -- if a path is relative, it is resolved to an absolute path based on the CDB entry's `directory` field. This ensures consistency in subsequent usage, regardless of the current working directory. - -### Toolchain Probing - -Toolchain probing converts driver-level commands to cc1 arguments: - -**Probing process**: For a given CanonicalCommand, the corresponding compiler driver (GCC, Clang, MSVC, etc.) is invoked to obtain the complete cc1 arguments -- including the target triple, system header search paths, default macro definitions, and all other implicit information. - -**Caching strategy**: Probing results are cached by (driver, file extension, non-user-content flags). The file extension is part of the cache key because `.c` and `.cpp` files may trigger different driver rules. User-content options are not part of the cache key because they do not affect driver output. - -**Negative caching**: Failed toolchain probes (e.g., a nonexistent GCC version) also cache the failure result, avoiding repeated attempts for every file. - -**Startup warm-up**: At server startup, probing is initiated in parallel for all unique cache keys, pre-populating the cache. By the time LSP requests start arriving, all toolchain probing results are already available. - -**Compiler adaptation**: Different compiler families (GCC, Clang, MSVC/ClangCL, etc.) have different driver invocation methods and output parsing logic. Some compiler families (NVCC, Intel, Zig) are recognized but fall through to a generic Clang-driver path rather than having fully dedicated handling. - -### Search Path Extraction - -After toolchain probing, the cc1 arguments contain the complete header search paths. SearchConfig extracts and organizes these paths into a four-tier structure: - -1. **Quoted** (`-iquote` directories): Search paths for `#include "foo.h"` -2. **Angled** (`-I` directories): Search paths for `#include ` -3. **System** (`-isystem` directories): System header paths where diagnostics are suppressed -4. **After** (`-idirafter` directories): Searched after system directories - -This four-tier model is largely consistent with Clang's internal search logic (some less common options like `-cxx-isystem`, `-iwithsysroot`, and Framework search paths are not yet supported). Paths within each tier are deduplicated (starting from the Angled tier, matching Clang's `RemoveDuplicates` algorithm), ensuring search behavior matches the actual compiler. - -SearchConfig is the foundational input for include path resolution, include path completion, dependency graph construction, and other features. - -## Design Decisions and Trade-offs - -**Why classify by option ID rather than string matching?** Classification uses Clang's own option table (`OptTable`), ensuring all options are correctly identified, including those with complex syntax (e.g., `-Wno-error=deprecated`, `-isystem=/usr/include`). String matching is prone to missing edge cases. - -**Why don't user-content options participate in the toolchain cache key?** This is the core of the two-level separation design. `-I` and `-D` do not change the system paths, target triple, or other information output by the driver. Excluding them reduces the number of cache keys from "one per file" to "one per configuration," yielding an order-of-magnitude improvement in large projects. - -**Why must search path deduplication match Clang's algorithm?** If the language server's search path order differs from the actual compiler's, include resolution might find different files, producing confusing inconsistencies. Strict matching ensures behavioral consistency. diff --git a/docs/en/design/compilation-context.md b/docs/en/design/compilation-context.md index d4c399725..f1fd23b8a 100644 --- a/docs/en/design/compilation-context.md +++ b/docs/en/design/compilation-context.md @@ -102,9 +102,9 @@ Beyond automatic resolution, clice also provides three LSP extension requests th The header context is conceptually clear (host source file + include position), but given a header context, how to actually make Clang compile the header under the correct preprocessor state is an engineering problem that requires choosing an approach. -> The prefix synthesis discussed here only applies to headers the user has opened. Headers on disk that are not open do not need special handling — they are included and processed normally when each source file is compiled. The indexing system collects symbol information from headers as part of each source file's indexing pass, and MergedIndex merges the index data produced for the same header under different source files (see [Index Design](index-design.md) for the merging mechanism). +> The prefix synthesis discussed here only applies to headers the user has opened. Headers on disk that are not open do not need special handling — they are included and processed normally when each source file is compiled. The indexing system collects symbol information from headers as part of each source file's indexing pass, and MergedIndex merges the index data produced for the same header under different source files (see [Index Design](symbol-index.md) for the merging mechanism). -clice uses **prefix synthesis + `-include` injection**: based on the host source file and include position from the header context, it extracts all content before the target header along the include chain, synthesizes it into prefix code, writes it to a prefix file on disk, and then injects that file into the compilation command via Clang's `-include` flag. The core advantage of this approach is its natural compatibility with PCH optimization — the prefix file's content is the target header's preamble, which can be compiled into a PCH and cached. When the user subsequently edits the header body, the large volume of header inclusions in the prefix does not need to be reprocessed each time. Detailed rationale for this design choice is in the FAQ section below. PCH construction, caching, and invalidation mechanisms are described in [Incremental Compilation Design](incremental.md). +clice uses **prefix synthesis + `-include` injection**: based on the host source file and include position from the header context, it extracts all content before the target header along the include chain, synthesizes it into prefix code, writes it to a prefix file on disk, and then injects that file into the compilation command via Clang's `-include` flag. The core advantage of this approach is its natural compatibility with PCH optimization — the prefix file's content is the target header's preamble, which can be compiled into a PCH and cached. When the user subsequently edits the header body, the large volume of header inclusions in the prefix does not need to be reprocessed each time. Detailed rationale for this design choice is in the FAQ section below. PCH construction, caching, and invalidation mechanisms are described in [Incremental Compilation Design](incremental-parse.md). The synthesis process has four stages. The following example illustrates the process. Suppose the project has these files: @@ -168,7 +168,7 @@ Compilation context affects not only live compilation but also index constructio #endif ``` -If the index only records the result from one context, go-to-definition or find-references would miss information from the other context. MergedIndex merges and stores index data produced from the same file under different compilation contexts, returning the union of all contexts on queries — for `config.h`, find-references can find references to both `AsyncHandler` and `SyncHandler`. See [Index Design](index-design.md) for the specific merging and deduplication mechanisms. +If the index only records the result from one context, go-to-definition or find-references would miss information from the other context. MergedIndex merges and stores index data produced from the same file under different compilation contexts, returning the union of all contexts on queries — for `config.h`, find-references can find references to both `AsyncHandler` and `SyncHandler`. See [Index Design](symbol-index.md) for the specific merging and deduplication mechanisms. ## FAQ diff --git a/docs/en/design/incremental-parse.md b/docs/en/design/incremental-parse.md new file mode 100644 index 000000000..d8fb12db1 --- /dev/null +++ b/docs/en/design/incremental-parse.md @@ -0,0 +1,199 @@ +# Incremental Compilation + +## Background + +C++ `#include` is textual substitution -- the preprocessor inserts the contents of included header files verbatim into the source file. A source file with only a few dozen lines of user code can expand to tens of thousands of lines or more after `#include` expansion. For example, a single `#include ` directive pulls in thousands of lines of standard library code; add in project headers and the preprocessed output can easily exceed a hundred thousand lines. + +A language server must recompile files after every user edit to provide up-to-date diagnostics, completions, and semantic information. Fully compiling those hundred thousand lines every time would take several seconds -- clearly unacceptable. + +However, a key characteristic emerges when observing actual editing behavior: the `#include` region at the top of a file rarely changes, while the user code at the bottom changes constantly. In a typical editing session, the user might modify code hundreds of times without touching the `#include` region once. + +```cpp +#include +#include +#include +#include "project/config.h" +#include "project/logging.h" +// ── preamble ↑ changes slowly, compilation result can be cached ── +// ── user code ↓ changes rapidly, must be recompiled each time ── + +void process(const std::vector& data) { + // ... +} +``` + +Based on this observation, C++ language servers commonly adopt a **preamble separation** strategy: split the file into the preamble (the preprocessor directive region at the top) and the remaining user code. The preamble is compiled into a precompiled header (PCH) and cached; subsequent compilations load the PCH directly and only reprocess the user code. This reduces post-edit recompilation to just a few dozen to a few hundred lines, typically bringing latency under one second. + +clangd also adopts this strategy, but has several design-level shortcomings in invalidation detection and lifecycle management: + +- **High memory usage**. clangd keeps preamble compilation artifacts in process memory. In large projects, preamble ASTs from multiple open files consume substantial memory, and usage grows over long-running sessions (clangd [#251](https://github.com/clangd/clangd/issues/251), [#115](https://github.com/clangd/clangd/issues/115)). + +- **PCH file leaks on crash**. clangd stores on-disk PCH in temporary files. On crash, RAII cleanup cannot execute and temp files persist in `/tmp`. On multi-user servers, accumulated leaked files can exhaust `/tmp` space (clangd [#209](https://github.com/clangd/clangd/issues/209), [#255](https://github.com/clangd/clangd/issues/255)). + +- **Imprecise invalidation detection**. When a header file is touched without changing its content (common during build tool dependency scanning, VCS branch switching), relying solely on modification time (mtime) triggers unnecessary PCH rebuilds. In large projects, these false positives noticeably degrade the editing experience. + +- **Cold start on restart**. PCH is not persisted across sessions. After a server restart, PCH must be rebuilt for all open files. + +clice redesigns the incremental compilation mechanism to address these issues: disk-persisted content-addressed PCH storage, two-layer invalidation detection, a pull-based compilation model, and preamble completeness checking. + +## Design + +Incremental compilation is organized around four core concepts: preamble separation defines "what to cache," two-layer invalidation detection defines "when to rebuild," pull-based compilation defines "when to trigger," and content-addressed storage defines "how to store." + +### Preamble Separation + +The preamble is the region at the top of a source file consisting of preprocessor directives (`#include`, `#define`, `#pragma`, etc.) and module declarations (`module;`). Its end position is identified by a byte offset (bound) -- the byte position just before the first line of non-preprocessor content. + +The preamble is compiled into a PCH file and cached on disk. Subsequent compilations load the PCH and only need to process user code after the bound. If the preamble is empty (bound is zero), no PCH is needed and the entire PCH flow is skipped. + +`PCHState` is a PCH cache entry, containing: + +- The PCH file path on disk +- A hash of the preamble content +- The preamble byte boundary (bound) +- A dependency snapshot (`DepsSnapshot`, see below) +- DocumentLink information extracted from the PCH (`#include` directive positions and targets, for the editor to display clickable links) + +### Two-Layer Invalidation Detection + +A PCH caches the preprocessing results of all headers included in the preamble. When the content of any dependency header changes, the PCH is stale and needs rebuilding. The challenge is precisely determining "whether the content actually changed." + +The most direct approach is checking file modification times (mtime): if all dependency files have mtimes no later than the PCH's build timestamp, no file has been modified. This check requires only `stat` system calls and is very fast. However, mtime checks produce false positives: build tool dependency scanning, VCS branch switching, editor auto-save, and similar operations update mtime without changing file content. + +An alternative is directly comparing content hashes: recompute the hash of every dependency file on each check and compare against the hashes recorded at build time. This approach is perfectly precise but requires reading and hashing the contents of all dependency files. A typical C++ file may depend on hundreds of headers, making the I/O cost of full hashing on every check non-negligible. + +clice combines both into a two-layer detection strategy: + +- **Layer 1 (mtime fast screening)**: Iterate over all dependency files and compare each file's mtime against the PCH's build timestamp. If all mtimes are no later than the build timestamp, the PCH is valid and can be reused directly. +- **Layer 2 (content hash precise verification)**: For files flagged as "possibly modified" by Layer 1 (mtime later than the build timestamp), recompute their xxh3 content hash and compare against the hash recorded at build time. Only trigger a rebuild when hashes differ. + +Layer 1 filters out the vast majority of unchanged files (the common-case path). Layer 2 eliminates mtime false positives (build tool touches, VCS checkouts, etc.). The combined effect: PCH is rebuilt only when the content of a dependency file has actually changed. + +`DepsSnapshot` is the underlying data structure for two-layer detection, captured when a PCH build completes. It records the path identifiers, content hashes, and build timestamp of all dependency files. + +### Pull-Based Compilation + +clice uses a pull-based compilation model: compilation is not triggered immediately on file change but on demand when a feature request (hover, completion, semantic highlighting, etc.) requires an up-to-date AST. + +When the user edits a file (`didChange`), the master process only updates the in-memory file content and marks the AST as dirty (`ast_dirty`), without starting any compilation. When a feature request arrives, the compilation service checks whether the AST is dirty or has become stale due to external file changes. If recompilation is needed, it first ensures the PCH and module dependencies are ready, then dispatches the compilation task to a worker process. + +The benefit of this model is avoiding wasteful compilations during rapid successive keystrokes. The user may trigger a dozen `didChange` events per second, but compilation only executes when an actual result is needed -- such as a hover or completion request. + +> Note that "external file changes" and "user edits" are two independent dirty-marking paths. User edits mark `ast_dirty` via `didChange`; external file changes (e.g., a dependency header modified on disk) are discovered dynamically via two-layer invalidation detection before compilation. + +### Content-Addressed PCH Storage + +PCH files on disk are named by the hash of the preamble content (e.g., `a3f7e8c1d2b4f6e9.pch`), implementing content-addressed storage. This provides two benefits: + +- **Disk sharing**: Different files with identical preamble content naturally share the same PCH file on disk, with no additional deduplication logic needed. +- **Cross-session persistence**: PCH cache metadata (path, hash, boundary, dependency snapshot) is serialized to a `cache.json` file on disk. On server restart, this metadata is loaded and each PCH's validity is verified through two-layer invalidation detection, avoiding the need to rebuild all PCHs on a cold start. + +When preamble content changes, the new PCH uses a different hash for its filename and the old file becomes orphaned. A cleanup mechanism periodically reclaims orphaned PCH files that have not been used beyond a certain age. + +## Implementation + +### Preamble Boundary Computation + +The preamble boundary is determined through lexical scanning: the project's `Lexer` scans line by line from the start of the file, identifying preprocessor directives beginning with `#` and `module;` global module fragment declarations. Scanning stops at the first line of non-preprocessor content, and the byte offset at that point is returned as the boundary. + +This lexer-based detection does not require starting a full preprocessor and is very fast. + +### Preamble Completeness Check + +Before triggering a PCH rebuild, the preamble must be checked for syntactic completeness. Two typical incomplete states: + +```cpp +#include "lib // unclosed quote, user is typing a filename +import std.core // missing semicolon, user is typing a module declaration +``` + +Building a PCH from an incomplete preamble produces a PCH containing erroneous preprocessor state. Subsequent compilations loading this erroneous PCH would see a flood of spurious errors. Therefore, when the preamble is detected as incomplete, the rebuild is deferred and the old PCH (if one exists) continues to be used. + +### PCH Build Pipeline + +When a compilation request requires a PCH, the following pipeline executes: + +``` +Compute preamble boundary and hash + │ + ▼ + ┌─ Cache hit? ──── yes → Reuse cached PCH + │ │ + │ no + │ │ + │ ▼ + │ Preamble complete? ── no → Defer rebuild, keep old PCH + │ │ + │ yes + │ │ + │ ▼ + │ Another coroutine building? ── yes → Wait for completion, use result + │ │ + │ no + │ │ + │ ▼ + │ Dispatch to stateless worker to build PCH + │ │ + │ ▼ + └─ Update cache, capture dependency snapshot +``` + +Cache hit requires two conditions: the preamble hash matches the cached value (preamble content unchanged), and two-layer invalidation detection passes (dependency file contents unchanged). Both conditions must be satisfied simultaneously. + +PCH builds are executed by stateless worker processes (see [multi-process architecture](multi-process.md)). The worker uses Clang's Preamble compilation mode, processing only the preamble portion before the bound. Upon completion, it returns the PCH file path and list of dependency files. + +### Concurrent Build Serialization + +Multiple feature requests may simultaneously trigger a PCH build for the same file. `PCHState` contains a shared event (`building`): the first coroutine to initiate a build sets this event, and subsequent coroutines that find the event present wait for its completion and then use the build result. This ensures the PCH for a given file is built only once. + +### Dependency Snapshot Timing Guarantee + +The `DepsSnapshot` build timestamp (`build_at`) is obtained **before** computing file hashes. This ordering ensures there is no time window in which a modification could be missed: + +If a file is modified during hash computation, its mtime will be later than `build_at`. On the next two-layer detection, Layer 1 will flag this file as "possibly modified," and Layer 2 will recompute its hash and discover the change. + +If the order were reversed -- hashing first, then obtaining the timestamp -- a window could arise: a file modified between hash computation and timestamp acquisition would have an mtime no later than `build_at`, causing the modification to be missed. + +### Overall Compilation Flow + +When a feature request arrives, the compilation pipeline executes in this order: + +1. Check whether the AST is cached and not stale -- if so, reuse it directly +2. If C++20 modules are used, ensure module dependencies are ready (see [module compilation](module-graph.md)) +3. Ensure the PCH is ready +4. Dispatch the compilation task to a stateful worker process with the PCH path and module file paths +5. The worker loads the PCH and compiles only the user code after the preamble + +AST dependency files (headers `#include`d after the preamble) are also tracked via `DepsSnapshot`, using the same two-layer invalidation detection. Even if the user has not edited the current file, if a dependency header is modified on disk, the next feature request will trigger recompilation. + +### Interaction with Compilation Context + +For non-self-contained headers, the compilation context system synthesizes a prefix file and injects it via `-include` into the compilation arguments (see [compilation context](compilation-context.md)). This injected prefix file is processed by Clang during the preamble compilation phase and is therefore naturally covered by PCH caching. The PCH build pipeline handles source files and headers uniformly. + +### Cache Persistence + +PCH and PCM cache metadata are persisted to disk via a `cache.json` file. This file is updated after each successful build and loaded on server startup. Writes use a write-to-temp-then-atomic-rename pattern to prevent file corruption from mid-write crashes. + +After loading the cache on startup, all PCH entries are validated through two-layer invalidation detection. Stale entries are automatically rebuilt on the next compilation, requiring no special cache consistency recovery logic. + +## FAQ + +- **Why two-layer detection instead of content hashing alone?** Content hashing is precise but requires reading the contents of all dependency files. A typical C++ file may depend on hundreds of headers, making the I/O cost of full hashing on every check non-negligible. The mtime fast-screening layer reduces the number of files that need hashing to those "touched since the last build," which in the common case is typically zero. + +- **Why full PCH rebuild each time? Can it be updated incrementally?** Clang supports chained PCH: each `#include` in the preamble is built as an independent PCH link, with each link depending on the previous link's compilation artifact. When the user adds a new `#include` at the end of the preamble, only one new link needs to be appended to the existing chain rather than rebuilding the entire preamble. Benchmarks show (PR [#405](https://github.com/ykiko/clice/pull/405)) that for a preamble of 70 C++ standard library headers, incrementally appending one `#include` takes about 36ms compared to about 1230ms for a full rebuild (35x speedup). Chained PCH AST load latency is virtually unaffected (+2% to +6%). clice plans to adopt chained PCH to optimize incremental rebuild performance; this is currently in the experimental stage. + +- **Why pull-based compilation instead of push-based?** The key reason is that clice persists all files' PCH to disk, resulting in far more cached files than clangd's in-memory model (clangd keeps only a few active files' preambles via LRU). When a header file is modified, a large number of files' PCHs may be affected. With push-based compilation, all affected PCHs would need to be rebuilt immediately on header modification -- an unacceptable volume. Pull-based compilation defers rebuilding until a feature request arrives, rebuilding only the PCH for the file the user currently needs. Since PCH loading itself is fast (see next item), the latency introduced by on-demand rebuilding is small. + +- **Is disk PCH slower than in-memory PCH?** The practical impact is minimal. Clang loads PCH using mmap to map the file into memory, avoiding a full read-copy. More importantly, Clang deserializes AST nodes from the PCH lazily -- only nodes actually referenced during compilation are deserialized, and most PCH content is never accessed. Thus PCH loading performance depends primarily on how the binary file is mapped, with no fundamental difference between a disk file and an in-memory buffer. The benefits of disk PCH -- cross-restart persistence, no resident process memory usage, content-addressed sharing -- make this trade-off worthwhile. + +## Known Limitations + +- **Per-file independent caching**. The PCH cache is keyed by file path identifier. Even if two files have identical preamble content, they have independent cache entries -- each performing its own invalidation detection and build. While the PCH files on disk are shared through content-addressed naming, cache metadata (dependency snapshots, build state, etc.) is not shared. The improvement direction is to key the cache by preamble content hash plus compilation flags, enabling cross-file cache metadata sharing. + +- **Full rebuild**. Any content change in a dependency file triggers a full PCH rebuild, with no way to rebuild only the affected portion. The improvement direction is to adopt chained PCH (see FAQ), limiting the rebuild scope to the chain links after the point of change. + +- **Compilation flags not part of cache key**. PCH disk filenames and cache lookups do not consider compilation flags. When two files have the same preamble text but different compilation flags (e.g., `-D`), they may incorrectly share a PCH. The improvement direction is to include preprocessing-relevant compilation flags in the cache key. + +- **Incomplete preamble completeness check**. The current completeness check only covers unclosed quotes and missing semicolons in `#include`/`import` directives. Other types of incomplete edits (e.g., typing a `#define` value) are not detected. The impact of building a PCH from such an incomplete preamble on subsequent compilation has not been thoroughly tested. Further investigation is needed into Clang's behavior when processing incomplete preprocessor directives, to determine whether the completeness check scope should be extended. + +- **No proactive diagnostics push after header save**. Under the current pure pull-based model, after a user saves a header file, open source files that depend on it do not immediately update diagnostics -- the user must trigger an action on the source file (e.g., hover, edit) for the two-layer invalidation detection to discover the change and trigger recompilation. The improvement direction is to check affected open sessions on `didSave` and proactively trigger compilation for them (a hybrid push/pull model). diff --git a/docs/en/design/incremental.md b/docs/en/design/incremental.md deleted file mode 100644 index cd1716a27..000000000 --- a/docs/en/design/incremental.md +++ /dev/null @@ -1,139 +0,0 @@ -# Incremental Compilation - -## Background - -A language server must recompile files every time the user edits code in order to provide up-to-date diagnostics, completions, and semantic information. Full compilation of a C++ file can take several seconds -- a typical source file may expand to tens or even hundreds of thousands of lines after `#include` expansion. Recompiling from scratch on every keystroke would produce an unacceptable user experience. - -clice achieves incremental compilation through **preamble separation**: each source file is split into two parts -- the leading `#include` region (the preamble) and the remaining user code. The preamble is compiled into a precompiled header (PCH) and cached on disk; subsequent compilations load the PCH directly and only reprocess the user code. For example, a file containing `#include ` expands to roughly 20,000 lines; once the PCH is built, recompilation covers only the few dozen lines the user actually wrote. - -The core challenges of this mechanism lie in PCH **invalidation detection** and **lifecycle management**: - -**Precision of invalidation detection**: A PCH caches the preprocessing results of every `#include`d header. When any header is modified, the PCH is out of date. However, naively checking file modification times (mtime) causes many false positives -- build tools frequently touch files without changing their contents (e.g., `cmake --build` dependency scanning, VCS branch switches), triggering unnecessary PCH rebuilds. - -**Controlling rebuild overhead**: Building a PCH requires full preprocessing and serialization, typically taking hundreds of milliseconds to several seconds. If every save triggers a PCH rebuild, users will notice perceptible lag. The system must accurately determine whether the PCH is truly stale and only rebuild when necessary. - -**Concurrency coordination**: Multiple files may share the same preamble. When a PCH needs to be rebuilt, several compilation requests may attempt to trigger the rebuild simultaneously. The system must ensure only one build executes while other requests wait for the result. - -**Incomplete edits**: While the user is typing `#include "`, the preamble is in an incomplete state. Triggering a PCH rebuild at this point would produce an erroneous PCH, causing a flood of spurious errors in subsequent compilations. - -In clangd, incremental compilation issues have long plagued users: diagnostics not updating after header modifications (requiring the user to reopen the file or restart the server), overly frequent PCH rebuilds causing editing lag in large projects, and occasional PCH corruption during simultaneous multi-file editing that prevents the entire project from compiling correctly. - -## Design - -### Core Idea - -The fundamental strategy of incremental compilation is to split a source file into the **slowly changing part** (preamble) and the **rapidly changing part** (user code), handling each separately. The preamble is compiled into a PCH cached on disk and only rebuilt when the content of a dependent header actually changes; the user code is recompiled by loading the PCH after each edit, which is very cheap. - -The core challenge of PCH management is **precise invalidation detection** -- finding the right balance between "rebuilding too aggressively" and "using a stale cache." clice achieves precise detection through two layers of invalidation checking (mtime fast screening + content hash verification), combined with preamble completeness checking and serialized concurrent builds, ensuring the PCH is always in a correct and efficient state. - -### Preamble Boundary Detection - -At the heart of PCH is the preamble -- the region at the top of a source file consisting of preprocessor directives. The language server must precisely determine where the preamble ends (as a byte offset) so that only this portion is compiled into the PCH. - -Preamble boundary detection is based on lexical scanning: scanning line by line from the start of the file, identifying preprocessor directives beginning with `#` (`#include`, `#define`, `#pragma`, etc.) and global module fragment declarations (`module;`). Scanning stops at the first line of non-preprocessor content. The returned byte offset marks the preamble boundary. - -If the boundary is zero (the file has no preprocessor directives), no PCH is needed -- the PCH build step is skipped entirely and any existing PCH cache entry for that file is cleared. - -### Two-Layer Invalidation Detection - -PCH invalidation detection must answer the question: "Since the last build, has the content of any header file included in the PCH changed?" - -clice uses a two-layer detection strategy that balances precision and performance: - -**Layer 1: Modification Time (mtime) Fast Screening** - -The system iterates over all dependency files of the PCH and retrieves each file's mtime. If every file's mtime is no later than the PCH's build timestamp (build_at), no file has been modified and the PCH is still valid -- it is reused directly. This layer requires only stat system calls, making it very cheap. - -**Layer 2: Content Hash Precise Verification** - -For files whose mtime is later than build_at (flagged as "possibly modified" by the first layer), the system recomputes their content's xxh3 hash and compares it against the hash recorded at build time. If the hashes match, the file content has not actually changed (it was merely touched), and the PCH is still valid. A rebuild is triggered only when the hashes differ. - -The core value of this two-layer strategy is **eliminating false positives**: build tool touch operations, VCS checkout operations, and similar actions update mtime without changing content. Relying solely on mtime would cause many unnecessary PCH rebuilds. The second layer's content hashing filters out these false positives with modest additional overhead (hashing only the "suspect" files). - -### DepsSnapshot: Dependency Snapshot - -The foundation of two-layer detection is DepsSnapshot -- a snapshot of dependency state captured when a PCH build completes. It records: - -- **path_ids**: Path identifiers for all dependency files -- **hashes**: The content hash (xxh3_64bits) of each dependency file at build time -- **build_at**: The timestamp when the snapshot was captured - -**Timing correctness**: build_at is obtained **before** computing file hashes. This ensures that if a file is modified during hash computation (its mtime becomes later than build_at), the next detection cycle will flag it as "possibly modified" in the first layer and pass it to the second layer for verification. There is no time window (TOCTOU) in which a modification could be missed. - -### PCH Build Pipeline - -When a compilation request requires a PCH, the following pipeline is triggered via ensure_pch(): - -**1. Compute preamble boundary and hash** - -The preamble portion of the file content is extracted and its xxh3 hash is computed. This hash is used for PCH file naming on disk and also serves as a cache validity check. - -**2. Look up the cache** - -The system checks whether a cache entry already exists for the file in the Workspace's pch_cache. If an entry exists and: - -- The preamble hash matches (preamble content has not changed) -- Two-layer invalidation detection passes (dependency file contents have not changed) - -The cached PCH is reused directly with no rebuild needed. - -**3. Check preamble completeness** - -If the cache misses or is invalidated, the system checks whether the preamble is complete before triggering a rebuild -- specifically, whether there are unclosed quotes (e.g., `#include "lib`) or unterminated module declarations (e.g., `import foo` missing a semicolon). If incomplete, the user is likely editing the preamble region; the rebuild is deferred and the old PCH (if one exists) continues to be used. - -**4. Serialize concurrent builds** - -The system checks whether another coroutine is already building a PCH for this file. If so, it waits for the completion event and then uses the build result. This is implemented through a shared event (kota::event) in PCHState -- the first coroutine to initiate the build sets the event, and subsequent coroutines wait on the same event. - -**5. Dispatch to a stateless worker process** - -The PCH build is dispatched as a high-priority task to a stateless worker process. The worker uses Clang's Preamble compilation mode (CompilationKind::Preamble), passing the full file content and preamble boundary to Clang via file remapping; Clang processes only the preamble portion before the boundary. Upon completion, the worker returns the PCH file path and the list of dependency files. - -**6. Update the cache** - -The build result is stored in pch_cache: the PCH file path, preamble hash, preamble boundary, and dependency snapshot. DocumentLink information from the PCH (positions and targets of `#include` directives) is also captured, enabling the editor to display clickable header links. - -### PCH File Naming on Disk - -PCH files on disk are named by the hash of the preamble content, implementing content-addressable storage. This allows different files with identical preamble content to **share the same PCH file on disk**. When preamble content changes, the new PCH file uses a different hash for its name, and the old file becomes orphaned. A cleanup mechanism periodically reclaims orphaned PCH files that have not been used beyond a certain age. - -### PCH Cache Storage Model - -The PCH cache (pch_cache) is currently keyed by the source file's path_id. This means that even if two files have identical preambles, they have independent cache entries -- each performing its own invalidation detection and rebuild. While the PCH files on disk are shared through content-addressable naming, the cache metadata (dependency snapshots, build timestamps, etc.) is not shared. - -PCH cache metadata (path, hash, boundary, dependency snapshot) is persisted to a cache.json file on disk. This allows recovery on server restart, avoiding the need to rebuild all PCHs on a cold start. - -### Compilation Triggering and Overall Flow - -clice uses a **pull-based** compilation model -- compilation is not triggered immediately on file change but is triggered on demand when a feature request (hover, completion, semantic highlighting, etc.) requires an up-to-date AST. This avoids wasteful compilations during rapid successive keystrokes. - -When a feature request requires an up-to-date AST, the compilation pipeline proceeds through these steps: - -1. Check whether the AST is cached and not dirty -- if so, reuse it directly -2. Check whether the AST's dependencies are stale (using the same two-layer invalidation detection) -3. If recompilation is needed, first ensure C++20 module dependencies are ready (via CompileGraph; see the module compilation documentation) -4. Then ensure the PCH is ready (via ensure_pch) -5. Finally, dispatch the compilation task to a stateful worker process with the PCH path - -The stateful worker loads the PCH and only needs to compile the user code after the preamble -- this is the core benefit of incremental compilation, reducing compilation time from several seconds to typically under one second. - -### PCH and Header Compilation Context - -For non-self-contained headers, the compilation context system synthesizes a prefix file (see the compilation context documentation) and injects it via `-include` to simulate the header's inclusion environment within its host source file. This injected prefix file appears in the header's compilation preamble and is therefore covered by PCH caching. - -ensure_pch() handles headers and source files uniformly: whether the preamble contains the file's own `#include` region or injected prefix content, the same hashing, caching, invalidation detection, and build pipeline applies. - -## Design Decisions and Trade-offs - -**Why two-layer detection instead of content hashing alone?** While content hashing is precise, it requires reading and hashing the entire contents of every dependency file. A typical C++ file may depend on hundreds of headers, making the I/O cost of hashing all of them non-trivial. The mtime fast-screening layer reduces the number of files that need hashing to those "touched since the last build," which is typically far smaller than the total dependency count. - -**Why not incremental PCH patching?** Ideally, when only one header undergoes a small change, it should be possible to patch the PCH incrementally rather than fully rebuilding it. However, Clang's PCH format does not support incremental patching -- it is a monolithic serialized snapshot of AST and preprocessor state. Incremental patching would require deep modifications to Clang's serialization layer, and the complexity is disproportionate to the benefit. - -**Why doesn't the PCH hash include compilation flags?** Currently, the PCH hash is based solely on the preamble text content and does not include compilation flags (`-D`, `-std`, etc.). This is a deliberate trade-off -- it allows files with the same preamble text but different compilation flags to share a PCH file on disk. In practice, the preprocessing result of the same preamble under different compilation flags is usually identical (differences mainly come from macro definitions, which are typically reflected in `#define` directives within the preamble text). However, there are edge cases where this is incorrect -- for example, when a command-line `-D` defines a macro that affects header behavior, different compilation flags should use different PCHs. - -**Why is preamble completeness checking important?** Building a PCH from an incomplete preamble produces a PCH containing erroneous preprocessor state. Subsequent compilations using this erroneous PCH will exhibit a flood of spurious errors. Deferring the rebuild until the preamble is complete avoids this cascade of errors. - -## Known Limitations - -- **Per-file independent caching**: pch_cache is keyed by path_id and cannot share cache metadata (dependency snapshots, build state, etc.) across different files with identical preambles. A future improvement is to key the cache by preamble content hash plus compilation flags, enabling true cross-file cache sharing while also addressing the issue of compilation flags not participating in the cache key, along with introducing LRU eviction. -- **Full rebuild**: Any content change in a dependency file triggers a full PCH rebuild. There is no way to rebuild only the affected portion. This is an inherent limitation of the Clang PCH format. diff --git a/docs/en/design/index-design.md b/docs/en/design/index-design.md deleted file mode 100644 index 84df5a333..000000000 --- a/docs/en/design/index-design.md +++ /dev/null @@ -1,157 +0,0 @@ -# Index Design - -## Background - -Cross-file features of a language server — go-to-definition, find-references, call hierarchy, symbol search, etc. — all rely on a symbol index. The indexing system must address several core challenges: - -**Performance at scale**: C++ projects can contain tens of thousands of files, each producing a large number of symbols after compilation. The indexing system must build, persist, and query the index within reasonable time and memory constraints. - -**Incremental updates**: After a user modifies a file, the relevant index entries should update incrementally rather than re-indexing the entire project. - -**Compilation-context awareness**: The same header file can produce different symbols under different compilation contexts. Existing solutions like clangd store only the result of the last compilation, causing users to see stale index data after switching contexts. - -**Real-time feedback for open files**: Files the user is actively editing should be immediately reflected in query results, without waiting for background indexing to finish. - -The following issues come up repeatedly in clangd's issue tracker: - -- Go-to-definition jumps to the wrong location (when multiple translation units define identically named symbols) -- Background indexing is slow — large projects need tens of minutes for the initial index pass -- Symbol search results are incomplete or contain stale entries -- Template code in header files has different reference relationships under different instantiation contexts - -## Design - -### Three-Level Index Hierarchy - -clice uses a three-level index structure, each level serving a different purpose: - -```text -TUIndex (compilation artifact) - ↓ merge -ProjectIndex (global symbol directory) + MergedIndex (per-file sharded relation data) - ↑ overlay -OpenFileIndex (real-time override for open files) -``` - -### TUIndex: Single-Compilation Output - -A TUIndex is the raw index data produced by a single compilation. When a translation unit is compiled, SemanticVisitor traverses the AST, records all symbol occurrences and relations, and produces a TUIndex. - -A TUIndex contains: - -- **Per-file index data**: Each file involved in the compilation (the main file plus all included headers) has its own FileIndex, recording occurrences and relations within that file -- **Symbol table (SymbolTable)**: Maps symbol hashes to symbol names and kinds -- **Include graph**: Records include relationships between files - -A TUIndex is transient data — it is no longer needed after being merged into ProjectIndex and MergedIndex. In the background indexing scenario, a TUIndex is generated and serialized in a worker process, then transmitted to the master process for merging. - -### ProjectIndex: Global Symbol Directory - -ProjectIndex is the global symbol directory, aggregating symbol information from all indexed files. It functions like a search engine's inverted index — given a symbol hash, you can quickly find its name, kind, and which files it appears in. - -ProjectIndex stores: - -- **Symbol table**: Symbol hash → symbol name, symbol kind, reference file bitmap (Bitmap) -- **Path pool**: Internalized mapping of file paths - -**Key design point**: ProjectIndex does not store concrete location information (line and column numbers). Location information is stored in MergedIndex's per-file shards. ProjectIndex's role is that of a "directory" — it tells you which files a symbol exists in, then you look up the exact locations in the corresponding MergedIndex shard. - -This separation keeps ProjectIndex compact. Reference file bitmaps are stored using Roaring Bitmap compression, which is highly memory-efficient. - -### MergedIndex: Per-File Sharded Index - -MergedIndex is the core storage layer of the indexing system. It maintains one shard per file in the project, storing all symbol occurrence positions and relation information for that file. - -The core feature of MergedIndex is **compilation-context merging**. The same header file may be included by multiple source files, and each compilation may produce different symbol relations (e.g., due to conditional compilation, template instantiation, etc.). MergedIndex merges index data from these different compilation contexts into the same shard. - -Merging uses content-addressed deduplication: a content hash is computed for each FileIndex produced by a compilation, and different compilation contexts with identical content share the same data. This avoids redundant storage when a header file is indexed multiple times. - -MergedIndex shards also store file contents, used for converting byte offsets to LSP positions (line/column numbers). - -### Symbol Identity: SymbolHash - -Every symbol in the system is identified by a SymbolHash (`uint64_t`). It is generated from Clang's USR (Unified Symbol Resolution) — a canonical string representation of symbol identity that encodes namespace, class name, function signature, template parameters, and other information. - -Key properties of SymbolHash: - -- **Cross-file consistency**: The same symbol has the same SymbolHash across different files — this is the foundation for cross-file navigation -- **Compact and efficient**: A 64-bit integer is better suited as a DenseMap key than a variable-length USR string -- **Deterministic**: The same declaration always produces the same hash - -### Symbol Relation Model - -The core data in the index consists of **relations** between symbols. Each relation contains: - -- **Relation kind (RelationKind)**: Definition, declaration, reference, weak reference, read, write, interface, implementation, type definition, base class, derived class, constructor, destructor, caller, callee, etc. -- **Location**: The source range where the relation occurs -- **Target symbol**: The other end of the relation (for type relations such as inheritance and calls) - -Complementing relations are **occurrences** — records of a symbol's simple presence at a location, without a relation kind. - -This distinction allows different queries to be served efficiently: - -- **Go to definition**: Look up RelationKind::Definition -- **Find references**: Look up RelationKind::Reference -- **Call hierarchy**: Look up RelationKind::Caller / Callee -- **Type hierarchy**: Look up RelationKind::Base / Derived -- **Symbol under cursor**: Look up Occurrence - -## Query Flow - -### Cross-File Queries - -Taking "go to definition" as an example, the query flow is: - -1. **Locate the cursor symbol**: In the current file, find the Occurrence at the cursor position via byte offset to obtain the SymbolHash -2. **Query the symbol directory**: Look up all files referencing this SymbolHash in ProjectIndex -3. **Per-file relation query**: For each referencing file, determine whether it is currently open: - - If the file **is open** (has a Session), use its OpenFileIndex to look up relation data, **skipping** the corresponding MergedIndex shard - - If the file **is not open**, use the MergedIndex shard to look up relation data - -Key point: For any given file, OpenFileIndex is **preferred** over MergedIndex — open files are queried via OpenFileIndex first, falling back to MergedIndex when the session AST is dirty or unavailable. Closed files are always queried via MergedIndex. This ensures that query results for open files reflect the latest buffer contents whenever possible. - -### Real-Time Override for Open Files - -Each open file maintains an OpenFileIndex in its Session, storing index data from the most recent compilation. It contains the same types of data as a MergedIndex shard (relations and occurrences), but sourced from in-memory AST compilation results rather than background indexing. - -Design principle: The open file's index is not written to the Workspace (does not affect global state); only after the file is saved does background indexing update the MergedIndex. This preserves the stability of the global index — half-typed code that may be incomplete should not pollute query results for other files. - -During queries, if a referenced file is open, the query prefers that file's OpenFileIndex over the MergedIndex shard, falling back to MergedIndex when the AST is dirty or unavailable. This ensures the user sees results corresponding to the latest buffer contents whenever possible. - -## Background Indexing - -### Scheduling Strategy - -Background indexing is managed by the Indexer, which maintains a queue of files awaiting indexing. It uses the following scheduling strategy: - -- **Idle-timeout batching**: Files are not indexed immediately upon entering the queue. Instead, an idle timer is set. When the timer fires, all files currently in the queue are processed as a batch. This avoids frequent small indexing runs. -- **Pause/resume**: When the user initiates interactive requests (hover, completion, etc.), background indexing can be paused to prioritize user requests. Nested pause semantics are supported. - -### Index Merging - -After background indexing completes, the TUIndex merge process is: - -1. Merge symbols from the TUIndex into ProjectIndex's global symbol table -2. Merge each file's FileIndex from the TUIndex into the corresponding MergedIndex shard -3. Update reference file bitmaps - -Merging is incremental — only new data is processed; there is no need to rebuild the entire index. - -## Serialization and Persistence - -The index is serialized using FlatBuffers, supporting efficient on-disk persistence and lazy loading. FlatBuffers' zero-copy property allows the index to be accessed directly via memory mapping, avoiding the overhead of full deserialization. - -Each MergedIndex shard is stored as an independent file on disk and loaded on demand. ProjectIndex is serialized as a whole since it is global data. - -## Design Decisions and Trade-offs - -**Why three levels instead of two?** The separation of ProjectIndex and MergedIndex is critical. If location information were also stored in ProjectIndex, its size would balloon, making it unusable as a lightweight "directory." Without ProjectIndex, every cross-file query would need to scan all MergedIndex shards to locate files containing a symbol — unacceptable for large projects. - -**Why don't open-file indexes write to global state?** This is part of the two-layer state model (Workspace vs. Session). Unsaved edits may consist of incomplete code with syntax errors. Writing them to the global index would pollute query results for other files. Only stable state that has been saved to disk should affect the global index. - -**Why use content-addressed deduplication?** Header files are frequently included by multiple source files, each compilation producing identical index data. Content addressing avoids storing N copies of the same data. When header content changes, old data is naturally replaced by the new version. - -## Known Limitations and Future Directions - -- **Symbol table locality**: Currently all symbols (including function-local variables, etc.) are merged into ProjectIndex's global symbol table. This causes a large number of symbols that are meaningful only within a single file to be stored globally, wasting memory and merge time. The improvement direction is to introduce file-level SymbolTables — local symbols stay in TUIndex / MergedIndex shards, and only truly cross-file-referenced symbols enter ProjectIndex. -- **Fuzzy symbol search**: Current symbol search uses simple substring matching. For scenarios like agent and workspace symbol queries, a more efficient fuzzy search index is needed (tokenization, prefix trees, n-gram indexes, etc.). diff --git a/docs/en/design/module-graph.md b/docs/en/design/module-graph.md new file mode 100644 index 000000000..cb634ca90 --- /dev/null +++ b/docs/en/design/module-graph.md @@ -0,0 +1,170 @@ +# Module Compilation + +## Background + +C++20 introduced modules, changing the independent compilation model that C++ has used since its inception. In traditional C++, each source file compiles independently into an object file, sharing declarations between files via headers. Modules break this independence: + +```cpp +// math.cppm — module interface unit +export module math; +export int add(int a, int b) { return a + b; } + +// main.cpp — imports the module +import math; +int main() { return add(1, 2); } +``` + +Before compiling `main.cpp`, the module interface unit `math.cppm` must be compiled first, producing a precompiled module file (PCM). If a project has multi-level module dependencies — A imports B, B imports C — the compilation order must be C → B → A. This forms a directed acyclic graph (DAG), where nodes are module files and edges are `import` relationships. + +Build systems (CMake, Ninja, etc.) are naturally suited for this kind of DAG scheduling: scan all files, build a complete dependency graph, compile in topological order. But a language server faces different challenges: + +**Real-time requirements.** After a user opens a file that imports modules, they expect editing feedback within hundreds of milliseconds. Waiting for the entire module graph to compile is unacceptable. A language server needs a lazy, on-demand compilation strategy — compile only the modules the current file actually needs. + +**Cascading file changes.** When the user modifies a module interface file and saves, all PCMs of modules that directly or indirectly depend on it become stale. The language server must detect this, cancel in-progress compilations, mark affected modules as dirty, and recompile them when next needed. + +**Concurrency and cancellation.** Multiple files may simultaneously need the same module's PCM. The language server must avoid duplicate compilations and let later requests wait for earlier ones to finish. At the same time, when the user closes the file that triggered a compilation and no longer needs a module, the in-progress compilation should be cancellable to free resources. + +**Temporary cyclic dependencies.** Although C++ modules do not allow circular `import`s, users may temporarily introduce cycles during editing. The language server must detect and report errors rather than deadlocking. + +In clangd, C++20 module support has long been in an experimental stage, lacking dedicated design. [clangd/clangd#1293](https://github.com/clangd/clangd/issues/1293) is the clangd team's summary of the module support problem — concluding that "these are not just bugs that can be fixed, a design and new infrastructure is needed," and at the time "nobody has plans/availability to work on this soon." Years later, related issues continue to appear: [clangd/clangd#2569](https://github.com/clangd/clangd/issues/2569) lists critical missing features in modules (rename, find references, etc.); [clangd/clangd#2292](https://github.com/clangd/clangd/issues/2292) and [clangd/clangd#2497](https://github.com/clangd/clangd/issues/2497) report file-locking conflicts caused by clangd sharing PCM files with the build system — clangd holds the PCM open, preventing the build system from overwriting it. + +The common root cause of these problems is the lack of a module compilation scheduling system designed specifically for language server use cases. clice addresses this through CompileGraph — an interest-counted compilation DAG with lazy construction, on-demand compilation, real-time cancellation, and dependency cascading. + +## Design + +CompileGraph is a compilation dependency graph built at runtime. Each node (CompileUnit) represents a module file, and edges represent `import` dependency relationships. + +### Lazy Construction + +Unlike build systems that scan all files and build a complete DAG before compilation, CompileGraph is lazily constructed — a node's dependencies are resolved and added to the graph only when first touched by a compilation request. When the user opens a file, only the module chain actually needed by that file is scanned and compiled, not the entire project's module graph. + +CompileGraph has two compilation entry points: + +- **Compile module**: Compile the specified module and all its transitive dependencies, producing a PCM file. Used for module interface units themselves. +- **Compile dependencies**: Compile all module dependencies of the specified file, but not the file itself. Used for ordinary source files — they are not part of the module DAG but may use modules via `import`. + +### Interest Counting + +Interest counting is CompileGraph's core scheduling mechanism, tracking "how many active requests currently care about a given compilation unit." It answers two questions: should compilation be started? Can compilation be cancelled? + +When a file needs a module, that module's reference count is incremented (acquire); when a request finishes or is cancelled, the count is decremented (release). A reference count reaching zero means no active request currently cares about this module. + +Reference counts are managed through RAII guards: creating a guard automatically acquires, and its destructor automatically releases. This naturally aligns with coroutine cancellation semantics — cancellation destroys the coroutine frame, the destructor releases the reference count, with no extra cleanup code needed. + +Interest counting operates at two levels: request-level and compilation-task-level. Request-level references mean "this request is waiting for a module's PCM"; compilation-task-level references mean "this module's compilation task is waiting for its direct dependencies to complete." The latter ensures a module is not prematurely cancelled when another module's compilation task depends on it, even if the original request has been cancelled. + +### Compilation Rounds + +Each compilation attempt constitutes a compilation round. Waiters learn of completion through a completion event and decide their next action based on the result. A round has three possible outcomes: + +- **Success**: PCM produced successfully, dirty flag cleared +- **Failed**: Compilation error (dependency failure, cycle detection, etc.), dirty flag retained +- **Stale**: Compilation cancelled (file modified during compilation, or reference count reached zero), waiters automatically drive a new round + +Failure is not sticky — retaining the dirty flag means that after the user fixes an error, the next request naturally triggers a retry without requiring a server restart. + +### Dirty State and Generation Counter + +CompileUnit uses a dirty flag to indicate that compilation is needed. The generation counter is a monotonically increasing value, incremented on each file update, used to detect asynchronous races: the compilation task records the current generation at start and compares upon completion — a mismatch means the file was modified during compilation, and the result is stale. + +### Dependencies + +Each CompileUnit maintains forward dependencies (modules it imports) and reverse dependencies (modules that import it). Forward dependencies are obtained via lazy resolution; reverse dependencies are back-filled at resolution time. Reverse dependencies are the basis for cascading notifications on file changes — starting from the modified module, following reverse edges finds all affected modules. + +The module name to file mapping is maintained by DependencyGraph (see [Dependency Scanning](dependency-scanning.md)) — the startup fast scan discovers all module declarations and builds a module name → file path registry. CompileGraph uses this registry to resolve `import` statements to concrete file paths. + +## Implementation + +### Compilation Flow + +The complete flow of a module compilation request: + +``` +Request enters + │ + ├─ RAII guard acquires on target module + │ + ├─ Target not dirty? ──→ Return immediately (PCM available) + │ + ├─ No compilation in progress? ──→ Start compilation task + │ │ + │ ├─ Lazily resolve dependencies (scan import declarations) + │ ├─ Check for self-cycle + │ ├─ Acquire on direct dependencies + │ ├─ Wait for all deps to compile (parallel) + │ ├─ Dispatch to worker process (produce PCM) + │ └─ Check generation counter ──→ Mismatch = Stale + │ + ├─ Wait for compilation round to complete + │ + └─ Based on result: Success → return / Failed → error / Stale → retry +``` + +Lazy resolution uses the Clang preprocessor for precise scanning, which differs from the fast lexer-based scan used during startup [dependency scanning](dependency-scanning.md). The fast scan does not expand macros or evaluate conditionals, suitable for building a global overview of include relationships; precise scanning expands all preprocessor directives to obtain the file's actual module dependencies under its current compile command. Resolution results are cached and reused by subsequent compilations until reset by a file update. + +> Module implementation units (those with `module X;` but no `export`) implicitly depend on their corresponding module interface unit. Precise scanning detects this and automatically adds the dependency. + +### Deferred Zero-Interest Cancellation + +When the reference count drops to zero, compilation is not cancelled immediately. Instead, a check is deferred by one event loop tick — if the reference count is still zero at that point, cancellation proceeds. + +This handles the scenario where a compilation request is superseded. When the user edits during compilation, a new request replaces the old one: the old request releases its references (count momentarily drops to zero), then the new request establishes references within the same tick (count rises back). Immediate cancellation would unnecessarily terminate shared dependency compilations. The one-tick delay allows such reference handoffs to complete smoothly, avoiding disruption to in-progress shared compilations. + +### Cascading Updates + +When a module file is saved (didSave), CompileGraph performs a cascading update: + +1. Reset the modified file's resolved flag — the next compilation will rescan dependencies, since the file may have added or removed `import` statements +2. Clear old forward dependency edges +3. Mark as dirty, increment the generation counter +4. Traverse all transitive dependents along reverse dependency edges; for each affected module: cancel its compilation round, mark dirty, increment generation +5. Return the list of all files marked dirty, so the caller can clear corresponding PCM caches + +Cascading updates do not modify reference counts — existing waiters retain their references. Upon observing a Stale result, they automatically drive a new compilation round. + +### Cycle Detection + +Before waiting for a dependency's compilation to complete, CompileGraph checks for wait cycles: starting from the target node, it searches along the dependency chain, following only nodes currently being compiled, checking whether the chain leads back to the current waiter. If a cycle is detected, it returns failure immediately, avoiding deadlock. + +### RAII Guards and Structured Concurrency + +Coroutine cancellation in kotatsu means destroying the coroutine frame — code after suspension points never executes; only destructors of already-constructed objects are guaranteed to run. CompileGraph leverages this by placing all cleanup logic in two layers of RAII guards: + +- **RefGuard** (request level): Holds root references from a request to modules. Releases reference counts on destruction when the request completes or is cancelled. +- **UnitGuard** (compilation round level): Manages all state for one compilation round. On destruction: publishes the outcome, clears the compiling flag, releases all acquired dependency references, and fires the completion event to notify waiters. + +All compilation tasks are managed through kota::task_group, providing structured concurrency guarantees: shutdown cancels all tasks first, then waits for their frames to unwind, ensuring no dangling compilation tasks remain. + +### PCM Caching + +PCM files use content-addressed path naming — the filename is determined by the module name and a hash of the compilation arguments, stored in a dedicated cache directory. This is fully isolated from build system artifacts, avoiding file-locking conflicts. + +PCM cache uses two-layer staleness detection: first comparing dependency files' modification times (mtime), then re-hashing content when times have changed. Recompilation only occurs when dependency content has actually changed, avoiding unnecessary rebuilds caused by "touch without modification." Cache metadata is persisted to `cache.json` on disk and can be restored on server restart. + +### Integration with the Compilation Pipeline + +CompileGraph is initialized during server startup. Before every file compilation, Compiler uses CompileGraph to ensure all module dependencies are ready — this is the first step of compilation preparation, executed before PCH construction. + +For `import` statements that the user has added in the editor but not yet saved (which CompileGraph is unaware of), Compiler performs an additional buffer scan to identify new module dependencies and attempts to build the corresponding PCMs. This is a compensation mechanism that covers the experience when the user is editing but has not saved. + +## FAQ + +- **Why interest counting instead of a task queue?** A task queue cannot express "no one needs this compilation anymore." When the user closes the file that triggered a compilation, continuing wastes resources. Interest counting precisely tracks demand, making cancellation decisions grounded — not based on timeouts or heuristics, but on whether any request is still waiting for the result. + +- **Why defer by one tick instead of cancelling immediately?** In an event loop model, multiple steps of an operation complete within the same tick. When a compilation request is superseded, the old reference is released and the new one established shortly after, with a momentary zero-reference in between. Immediate cancellation would unnecessarily terminate shared dependencies — deferring by one tick ensures reference handoffs within the same tick don't trigger cancellation. + +- **Why a generation counter instead of locks?** The master process is a single-threaded event loop with no data races. The "races" to detect come from asynchronous timing — "was the file updated during compilation?" The generation counter answers this with minimal overhead, without introducing locks. + +- **Why lazy dependency resolution?** Dependency resolution requires running the Clang preprocessor (precise scanning) on each module file, which is not cheap. If all module dependencies were resolved at startup, it would add to startup time. Lazy resolution ensures only modules actually needed for compilation are scanned — opening a file doesn't trigger scanning the entire module graph. + +- **Why are failures not sticky?** Users are actively editing code — syntax errors are the norm. If failures were marked as persistent, users couldn't get correct results after fixing errors without restarting the server. Retaining the dirty flag lets the next request naturally trigger a retry. + +- **Why store PCMs in a separate cache directory?** clangd shares PCM files with the build system, which causes file-locking conflicts — the language server holds the PCM open, preventing the build system from overwriting it (see [clangd/clangd#2292](https://github.com/clangd/clangd/issues/2292)). clice uses content-addressed isolated caching, avoiding this problem. The trade-off is additional disk space usage and extra time for first-time compilation. + +## Known Limitations + +- **Compilation task memory accumulation.** Each compilation round creates a task in the task_group. Completed task frames are only reclaimed when the task_group is destroyed, not immediately upon completion. In a long-running server, completed task frames accumulate over time. + +- **Dependency resolution determinism.** A time window exists between resolving dependencies and the actual compilation. If a file is modified during this window, the resolved dependencies may not match the actual dependencies at compilation time. The generation counter detects this and triggers retries, but at the cost of extra compilation overhead. + +- **Limited coverage for unsaved module dependencies.** CompileGraph builds dependency relationships from disk files. When the user adds an `import` in the editor without saving, Compiler compensates via buffer scanning, but if the imported module itself is not yet managed by CompileGraph (e.g., a new module file that hasn't been saved), the corresponding PCM cannot be built. The user needs to save the module interface file first, then save the importing file. diff --git a/docs/en/design/module.md b/docs/en/design/module.md deleted file mode 100644 index cb005a262..000000000 --- a/docs/en/design/module.md +++ /dev/null @@ -1,143 +0,0 @@ -# Module Compilation - -## Background - -C++20 introduced modules, the largest change to the C++ compilation model since the language's inception. Traditional C++ compilation is "compile each file independently" — each source file compiles to an object file, and files share declarations via header files. Modules break this independence: when a file `import`s another module, the imported module's **module interface unit** must be compiled first, producing a **precompiled module file (PCM)**, before the importing file can be compiled. - -This introduces **compile-time dependency relationships** — a directed acyclic graph (DAG), where nodes are module files and edges are `import` relationships. Build systems (CMake, Ninja, etc.) naturally support this kind of DAG scheduling, but a language server faces fundamentally different challenges: - -**Real-time requirements**: A build system can scan all files at once, build a complete dependency graph, and compile in topological order. A language server cannot — after a user opens a file, they expect editing feedback in milliseconds. Waiting for the entire module graph to compile is unacceptable. - -**On-demand compilation**: The user has only opened a few files in the project; there is no need to compile the entire module graph. What the language server needs is "compile only the modules the current file requires" — a lazy, incremental compilation strategy. - -**Cascading updates on file changes**: When the user modifies a module interface file and saves it, all PCMs of modules that directly or indirectly depend on it become stale. The language server must transparently cancel in-progress compilations, mark affected modules as dirty, and recompile them when next needed. - -**Concurrency and races**: Multiple files may simultaneously need the same module's PCM. If two requests trigger compilation of the same module at the same time, duplicate compilation must be avoided while ensuring the later request can wait for the earlier compilation to finish. - -**Cyclic dependencies**: Although C++ modules do not allow circular `import`s, users may temporarily introduce cycles during editing. The language server must detect this and report an error gracefully, rather than deadlocking. - -In clangd, C++20 module support has long been in an experimental stage. Recurring issues in the community include: incomplete dependency resolution for module compilation, leading to missing PCM files; module files modified without dependents being recompiled, leading to stale diagnostics; lack of real-time tracking of module dependencies, requiring a manual server restart to get correct compilation results. - -The root cause of these problems is the lack of a dedicated module compilation scheduling system. clice solves these problems at their source through CompileGraph — a reference-counted compilation DAG. - -## Design - -### Core Idea - -CompileGraph is a compilation dependency DAG built at runtime. Each node (CompileUnit) represents a module file, and edges represent `import` dependency relationships. Unlike a build system's static DAG, CompileGraph is **lazily built** — a module's dependencies are resolved and added to the graph only when the module is actually needed. - -CompileGraph's scheduling is based on **interest counting**: when a file needs a module, that module's reference count is incremented; when it is no longer needed, the count is decremented. A reference count reaching zero means no active request currently cares about this module, and its in-progress compilation can be cancelled to free resources. - -### Key Concepts of CompileUnit - -Each CompileUnit is organized around the following core concepts: - -**Dependencies**: Each unit maintains forward dependencies (modules it `import`s) and reverse dependencies (modules that depend on it). Forward dependencies are obtained via lazy resolution; reverse dependencies are automatically populated during resolution. Reverse dependencies are the basis for cascading notifications when files change. - -**Dirty state and generations**: The dirty flag indicates the unit needs recompilation. A generation counter is incremented on each file update, used to detect whether an asynchronous compilation result is stale — the compilation task captures the generation value at start and compares it upon completion. - -**Compilation rounds**: Each compilation produces a round containing a completion event and a result. Waiters learn of completion through the event and decide their next action based on the result. A cancellation source is used to cancel the current round's compilation task. - -**Reference count**: Tracks how many active requests care about this unit. When it reaches zero, compilation can be cancelled to free resources. - -### Reference Counting and RAII Guards - -Reference counting is the core scheduling mechanism of CompileGraph. It tracks "how many active requests currently care about a given compilation unit" and determines when compilation tasks are started and cancelled. - -CompileGraph provides two RAII guards for managing reference counts: - -**RefGuard (request-level guard)**: The entry point of a compilation request creates a RefGuard, which performs acquire (increment reference count) on all required compilation units. When the RefGuard is destroyed, it automatically performs release (decrement reference count). This ensures that even if a request is cancelled (coroutine frame destroyed), reference counts are correctly released. - -**UnitGuard (compilation-round guard)**: Each compilation task internally uses a UnitGuard to maintain the compilation round's state. It is responsible for: acquiring reference counts on all direct dependencies after dependency resolution; upon completion or cancellation, publishing the result, clearing the compiling flag, releasing all acquired dependency references, and firing the completion event to notify waiters. - -Using destructors rather than coroutine finally blocks for cleanup is a key design choice — cancellation is implemented via coroutine frame destruction, so cleanup logic must be placed in destructors to guarantee execution. - -### Lazy Dependency Resolution - -Dependency resolution is an expensive operation — it requires scanning the module file's contents and parsing `import` declarations. CompileGraph does not resolve dependencies at node creation time; instead, it calls resolve_fn for the first time the node needs to be compiled. - -resolve_fn is provided externally (typically based on precise scanning of the module file) and returns the list of files that the module directly depends on. Resolution results are cached in the CompileUnit's dependencies, and the dependents lists are back-filled at the same time. Subsequent compilations reuse the cached dependency information unless a file update resets the resolved flag. - -### Compilation Flow - -CompileGraph provides two entry points: - -**compile(path_id)** — compile the specified module and all its transitive dependencies: - -1. RefGuard performs acquire on the target module -2. If the target module is not dirty, return immediately -3. If no compilation is in progress, start a compilation task -4. Wait for the compilation round to complete -5. Based on the result, return success, failure, or retry (retry if stale) - -**compile_deps(path_id)** — compile all transitive module dependencies of the specified file, but not the file itself. This is used for ordinary source files (non-module files) — they are not part of the module DAG themselves but may `import` modules. - -The internal flow of a compilation task (unit_body): - -1. Resolve dependencies (ensure_resolved) -2. Check for self-cycles (self-importing) -3. Acquire reference counts on all direct dependencies -4. Recursively wait for all dependencies to finish compiling -5. Dispatch the actual compilation work to a stateless worker (dispatch_fn) -6. Check the generation counter — if the file was updated during compilation, the result is stale and the dirty flag is not cleared - -### Zero-Reference Deferred Cancellation - -When a compilation unit's reference count drops to zero, it means no active request currently cares about it. But cancelling immediately may be too aggressive — during dependency switching, the reference count might drop to zero and then increase again within the same event loop tick. - -CompileGraph uses a strategy of deferring by one event loop tick: when the reference count drops to zero, it does not cancel immediately but instead schedules a deferred check. The deferred check runs on the next event loop tick — if the reference count is still zero at that point, cancellation actually proceeds. This avoids transient zero-references triggering unnecessary cancellation and restart. - -### Generation Counter and Race Detection - -In an asynchronous system, a compilation task dispatched to a worker process must wait for the result to come back. During this time, the user may have modified the file, meaning the file content has already changed. If the old compilation result were marked as "clean," it would lead to use of stale PCMs. - -The generation counter solves this problem: the compilation task captures the current generation value at start, and upon completion, compares it against the current value. If they differ, a file update occurred during compilation, and the result is stale — the dirty flag is not cleared, and subsequent requests will trigger recompilation. - -### File Updates and Cascading Propagation - -When a module file is modified (via didSave), CompileGraph's update() method performs a cascading update: - -1. Reset the resolved flag of the modified file (forcing the next compilation to re-resolve dependencies) -2. Clear old dependency edges -3. Mark as dirty and increment the generation counter -4. Traverse all transitive dependents along reverse dependency edges -5. For each affected dependent: cancel its compilation round, mark it as dirty, and increment its generation counter - -update() returns the list of all modules marked dirty, allowing the caller to clear the corresponding PCM caches. - -Note that update() does not modify reference counts — existing waiters retain their references. When they observe that the compilation result is "stale," they automatically retry, driving a new round of compilation. - -### Cycle Detection - -Although legal C++20 modules do not allow cyclic dependencies, they may be introduced temporarily during editing. CompileGraph detects wait cycles before waiting for a dependency's compilation to complete: it searches along the target's dependency chain, following only nodes that are currently compiling, checking whether the chain leads back to the current waiter. If a cycle is detected, it returns failure immediately, avoiding deadlock. - -### Compilation Results and Retry Semantics - -A compilation round has three possible results: - -- **Success**: Compilation succeeded; the dirty flag is cleared -- **Failed**: Compilation failed (dependency failure, dispatch failure, or cyclic dependency). The dirty flag is not cleared — the next explicit request can retry -- **Stale**: Compilation was cancelled (due to a file update or reference count reaching zero). Waiters automatically retry a new compilation round - -Stale retries are bounded — each update() increments the generation only once, so retries do not accumulate. Without a continuous stream of updates, retries terminate naturally. - -### Structured Concurrency and Shutdown - -All compilation tasks are managed through kota::task_group. task_group provides a structured concurrency guarantee — it waits for all child tasks to complete upon destruction. CompileGraph's shutdown() method achieves graceful shutdown by cancelling all tasks in the task_group and waiting for them to exit. - -## Design Decisions and Trade-offs - -**Why reference counting instead of a simple task queue?** A task queue cannot express the semantics of "no one needs this compilation anymore." In module compilation scenarios, a user might close the file that triggered a compilation while that compilation is still in progress; continuing to compile at that point wastes resources. Reference counting precisely tracks demand, giving cancellation decisions a solid basis. - -**Why defer by one tick instead of cancelling immediately?** In an event loop model, multiple steps of an operation may complete within the same tick. If an old reference is released before a new one is established (transient zero-reference), immediate cancellation would cause unnecessary rework. Deferring by one tick ensures that reference handoffs within the same tick do not trigger cancellation. - -**Why use a generation counter instead of locks?** The master process is a single-threaded event loop — there are no data races. The "races" here come from the timing of asynchronous operations. The generation counter detects "whether a file update occurred during compilation" with minimal overhead, without introducing the complexity of locks. - -**Why is dependency resolution lazy?** Dependency resolution requires reading and scanning file contents, which is not cheap. If all dependencies were resolved at node creation time, opening a single file could trigger scanning of the entire module graph. Lazy resolution ensures that only modules actually needed for compilation are scanned. - -**Why are failures non-sticky?** Module compilation failures are usually transient — the user is editing code, and syntax errors are the norm. Making failures persistent would prevent the user from getting correct results after fixing an error. Not clearing the dirty flag means the next request naturally retries. - -## Known Limitations - -- **Compilation graph memory accumulation**: Each compilation round creates a task in the task_group. Completed task frames are only reclaimed when the task_group is destroyed, not immediately upon completion. In a long-running server, a large number of completed task frames may accumulate. -- **Dependency resolution determinism**: The result of resolve_fn depends on the file's current content. If the file is modified between resolve and the actual compilation, the resolved dependencies may be inconsistent with those at compilation time. The generation counter can detect this, but it results in additional retries. diff --git a/docs/en/design/overview.md b/docs/en/design/overview.md index 797244b48..34a8fceb3 100644 --- a/docs/en/design/overview.md +++ b/docs/en/design/overview.md @@ -4,163 +4,107 @@ This document describes the responsibilities and role of each module in the clic ## Project Vision -clice is a ground-up redesign of a C++ language server, architected to solve fundamental problems that have long plagued existing C++ language servers. Core innovations include: +clice is a brand-new C++ language server, redesigned from the architecture level to solve long-standing problems in previous C++ language servers. Key features: -- **Compilation Context**: The same file can produce different results under different compilation contexts. clice treats this concept as a first-class citizen throughout the entire design -- from compilation to indexing to query responses. No existing language server (not just C++) formally addresses this. +- **Compilation Context**: clice is the first language server to introduce compilation context as a formal concept. Every step of compilation, indexing, and querying explicitly distinguishes the current compilation context, and users can query and switch between them. See [Compilation Context](compilation-context.md). -- **Multi-process Architecture**: A master + worker process model isolates Clang's memory leaks and crashes while enabling priority-aware scheduling and real-time memory monitoring. +- **Multi-process Architecture**: A master + worker process model isolates Clang's memory leaks and crashes while enabling priority-aware scheduling and real-time memory monitoring. See [Multi-process Architecture](multi-process.md). - **Coroutine-based Async Model**: Built on C++20 coroutines and the kotatsu library, replacing traditional callback-style async and making business logic clearer. -- **Real-time Module Compilation System**: A reference-counted C++20 module compilation DAG with support for real-time cancellation and dependency cascading -- the most advanced real-time module compilation scheduling approach known to date. +- **Real-time Module Compilation**: A reference-counted C++20 module compilation DAG with support for real-time cancellation and dependency cascading. See [Module Compilation Graph](module-graph.md). -clice's long-term goal extends beyond being a language server. It will integrate concurrent scheduling for tools like clang-tidy, enable cross-translation-unit optimizations (such as duplicate header detection), and become a unified, high-level platform for the C++ tooling ecosystem. +## Module Overview -## Foundation Layer +### `src/support/` — Foundation Utility Library -### `src/support/` +General-purpose utilities and infrastructure shared by all other modules. -General-purpose utility library. Provides logging, filesystem abstractions, string operations, pattern matching, and other infrastructure. +- `PathPool`: Internalizes file paths as `uint32_t` identifiers, used as stable file identifiers throughout the system +- `FuzzyMatcher`: Token-aware fuzzy matching for code completion and symbol search +- Markup / Doxygen: Parsing and formatting of documentation comments +- Logging, filesystem abstractions, string utilities, etc. -Key components: +### `src/command/` — Compilation Command Processing -- **PathPool**: Internalizes file paths as compact `uint32_t` identifiers, used as stable file identifiers throughout the system. This is globally shared infrastructure that nearly every module dealing with file paths depends on. -- **Logging**: Structured logging based on spdlog. -- **FuzzyMatcher**: Token-aware fuzzy matching for scenarios like code completion. -- **Markup / Doxygen**: Parsing and formatting of documentation comments. -- **StringSet / ObjectSet**: Deduplicating internalization of strings and objects, with pointer stability guaranteed by a bump allocator. +Commands read from the compilation database (CDB) are raw commands generated by the build system and cannot be fed directly to the Clang frontend -- they may contain options meant only for code generation, lack system header search paths, or include parameters irrelevant to a language server. This module classifies, filters, probes toolchains, and deduplicates the raw commands, transforming them into compilation parameters consumable by the language server. -### `src/command/` +- `CompilationDatabase`: Loads `compile_commands.json` and separates compilation commands into flags that affect semantics (canonical) and flags that only affect user content (patch, such as `-I` and `-D`), enabling cross-file command deduplication and sharing +- `Toolchain`: Queries system compilers for complete compilation parameters (such as system header search paths), caching results by (driver, file extension, non-user-content flags) +- `SearchConfig`: A four-tier model for header search paths (Quoted / Angled / System / After), matching Clang's internal search logic -CLI parsing and compilation command processing. Commands read from the compilation database (CDB) are raw commands generated by the build system and cannot be fed directly to the Clang frontend -- they may contain options meant only for code generation, lack system header search paths, or include parameters irrelevant to a language server. This module classifies, filters, probes toolchains, and deduplicates the raw commands, ultimately transforming them into compilation parameters consumable by the language server. +See [Compilation Command Resolution](command-resolve.md). -Key components: +### `src/compile/` — Compilation Abstraction -- **CompilationDatabase**: Loads `compile_commands.json` and performs a two-level separation of compilation commands -- splitting flags that affect semantics (canonical) from flags that only affect user content (patch, such as `-I` and `-D`), enabling cross-file command deduplication and sharing. -- **Toolchain**: Queries system compilers for complete compilation parameters (such as system header search paths), caching results by (driver, file extension, non-user-content flags) to avoid redundant probing of the same toolchain. -- **ArgumentParser**: Classifies Clang compilation options into three categories -- codegen-only (discarded), discarded, and user-content (extracted as patches) -- ensuring only semantically meaningful parameters are retained. -- **SearchConfig**: A four-tier model for header search paths (Quoted / Angled / System / After), matching Clang's internal search logic. +Wraps the Clang compiler, abstracting Clang APIs into safe, unified compilation interfaces. This layer is purely a compilation abstraction with no server logic. -## Compilation Abstraction Layer +- `CompilationUnit` / `CompilationUnitRef`: RAII wrappers around the Clang AST context. `CompilationUnitRef` provides a unified read-only view for accessing source location mappings, preprocessor directives, AST nodes, and more. This is the primary input for `src/feature/` and `src/semantic/`. +- `CompilationParams`: Describes the complete configuration for a single compilation, including compilation type (Preamble / Content / Completion / Indexing, etc.), file remapping, PCH/PCM reuse, etc. -### `src/compile/` +### `src/syntax/` — Lightweight Syntax Processing -Wraps the Clang compiler. Abstracts raw Clang APIs into safe, unified compilation interfaces that hide low-level details. +Syntax-level processing that does not require a full AST. Runs before compilation to quickly obtain structural information and dependency relationships for files. -This layer is purely a compilation abstraction with no server logic. It takes compilation parameters, drives Clang through compilation, and produces results -- including AST, preprocessor state, source location mappings, and more. +- `Lexer`: A token-level utility built on Clang's raw lexer. Does not run the preprocessor. Used for directive scanning, include path resolution, etc. +- `DependencyGraph`: A global include/module dependency graph supporting forward queries, reverse queries, host source file search, include chain lookup, etc. +- Dependency scanning: Wraps Clang's `DependencyDirectivesScanner` to quickly extract include and module dependencies +- `IncludeResolver`: Resolves include paths to actual files based on search path configuration -Key components: +See [Dependency Scanning](dependency-scanning.md). -- **CompilationUnit / CompilationUnitRef**: RAII wrappers around the Clang AST context. `CompilationUnitRef` provides a unified read-only view for accessing source location mappings, preprocessor directives, file contents, AST nodes, and more. This is the primary input for `src/feature/` and `src/semantic/`. -- **CompilationParams**: Describes the complete configuration for a single compilation operation, including compilation type (Preprocess / Content / Preamble / ModuleInterface / Completion / Indexing), file remapping, PCH/PCM reuse, cancellation flags, etc. -- **Compilation Types**: Different compilation types produce different outputs. Preamble produces PCH, ModuleInterface produces PCM, Content produces a full AST, and Completion produces completion candidates. They share the same parameter framework but have different semantics. +### `src/semantic/` — Semantic Analysis -### `src/syntax/` +Semantic analysis capabilities beyond Clang's native APIs. Takes a `CompilationUnitRef` and extracts higher-level semantic information. -Lightweight syntax processing that does not require a full AST. +- `SemanticVisitor`: An AST traverser that records each symbol's occurrence location and relations (definition, reference, inheritance, call, etc.). This is the primary producer of index data. +- `TemplateResolver`: Resolves dependent names through pseudo-instantiation, enabling semantic analysis to see through template contexts. See [Template Resolver](template-resolver.md). +- `SymbolKind` / `RelationKind`: Fine-grained symbol kinds and relation types -This layer handles tasks that can be completed without full compilation: lexical analysis, dependency scanning, include path resolution, and more. It runs before compilation to quickly obtain structural information and dependency relationships for files. +### `src/index/` — Symbol Index -Key components: +Symbol indexing system with cross-translation-unit query support. Uses a three-tier structure: -- **Lexer**: A token-level utility built on Clang's raw lexer, providing a streaming interface. Does not run the preprocessor or expand macros. Used for directive scanning, include path resolution, and similar tasks. -- **DependencyGraph**: A global include/module dependency graph. Records the include relationships for each file (under each SearchConfig), supporting forward queries (what does this file include?), reverse queries (what includes this file?), and module-name-to-file mappings. Also provides BFS upward search for host source files, shortest include chain lookup, and other navigation features. -- **Scan**: Wraps Clang's `DependencyDirectivesScanner` to quickly extract include and module dependencies from files without full preprocessing. -- **IncludeResolver**: Resolves include paths to actual files based on search path configuration. -- **Completion**: Generates candidates for include path completion and module import completion. +- `TUIndex`: Index data produced from a single compilation, generated by `SemanticVisitor` +- `ProjectIndex`: The global symbol table. Aggregates symbol information from all indexed files, supporting lookup by symbol hash +- `MergedIndex`: Per-file sharded index storage. Merges index data produced from the same file under different compilation contexts -## Semantic Analysis Layer +See [Symbol Index](symbol-index.md). -### `src/semantic/` +### `src/feature/` — LSP Feature Implementations -Semantic analysis capabilities beyond Clang's native APIs. +Concrete implementations of LSP features. Each feature takes a `CompilationUnitRef` and returns the corresponding LSP response data. This layer is purely computational -- it has no involvement with network communication, state management, or process scheduling. -This layer takes a `CompilationUnitRef` and extracts higher-level semantic information -- symbol relations, symbol classification, template resolution, and more. Its output is the core data source for the indexing system and LSP features. +Includes: code completion, hover information, signature help, semantic highlighting, inlay hints, document symbols, document links, folding ranges, formatting, diagnostics, etc. -Key components: +> `feature/` only covers single-file, AST-based feature implementations. Cross-file navigation features (go to definition, find references, etc.) are handled by the `Indexer` using index data. Some features involve multi-phase processing -- for example, include path completion in code completion can be resolved at the syntax layer without full compilation. -- **SemanticVisitor**: An AST traverser that records each symbol's occurrence location and relations (definition, reference, read, write, inheritance, call, etc.). This is the primary producer of index data. -- **TemplateResolver**: Resolves dependent names through pseudo-instantiation, enabling semantic analysis to see through template contexts. This is a key innovation of clice -- clangd has limited capability in this area. -- **SymbolKind / RelationKind**: Fine-grained symbol kinds and relation types, far more detailed than those defined by the LSP protocol. -- **find_target**: Resolves AST nodes to their declaration targets, handling various indirect references and implicit conversions. +### `src/server/` — Server Runtime -## Index Layer +The language server's core runtime, responsible for assembling all the layers above into a runnable service. -### `src/index/` +**`protocol/`** — Protocol definitions. Describes the message formats for communication between the master process and worker processes, as well as between the server and clients. Includes Worker protocol (compilation/query/build requests), LSP extension protocol (compilation context switching, etc.), and the agentic protocol for AI agents. -Symbol indexing system with cross-translation-unit query support. +**`workspace/`** — Project-level global state. `Workspace` holds the compilation database, toolchain, path pool, dependency graph, PCH/PCM cache, project index, and all other project-level state. Core invariant: unsaved buffer contents of open files never modify the `Workspace` -- it only reflects the state on disk. -The indexing system uses a three-tier structure, corresponding to different query scenarios and lifecycles: +**`compiler/`** — Compilation scheduling and index management. -- **TUIndex**: Index data produced from a single compilation. Records symbol occurrences, relations, and include graphs for each file involved in the compilation. This is the raw index data source, generated by `SemanticVisitor` during compilation. -- **ProjectIndex**: The global symbol table. Aggregates symbol information from all indexed files, supporting lookup by symbol hash for names, kinds, and referencing files. Used for cross-file symbol search and navigation. -- **MergedIndex**: Per-file sharded index storage. Merges index data produced from the same file under different compilation contexts. MergedIndex is the core embodiment of compilation-context awareness -- a header file may be included by multiple source files, each inclusion producing different symbol relations, and MergedIndex unifies them. +- `Compiler`: The scheduler for compilation lifecycles. Coordinates the ordering of PCH builds, module dependency resolution, and AST compilation, dispatching compilation tasks to worker processes +- `CompileGraph`: The DAG scheduler for C++20 module compilation. Uses reference counting for interest tracking, supporting dependency cascade cancellation +- `Indexer`: Handles background indexing scheduling and serves as the entry point for cross-file queries. Combines `ProjectIndex`, `MergedIndex`, and in-memory indexes for open files to provide query results -Index data is serialized with FlatBuffers, supporting on-disk persistence and lazy loading. +**`service/`** — Service entry point and session management. -## Feature Layer +- `MasterServer`: The top-level coordinator. Holds `Workspace`, `Session` map, `WorkerPool`, `Compiler`, and `Indexer`, routing LSP requests to the appropriate handling logic +- `Session`: The editing state for each open file (buffer contents, compilation version number, in-memory index, PCH reference, etc.). Created on didOpen, destroyed on didClose +- `LSPClient` / `AgentClient`: Request handlers for the LSP protocol and agentic protocol -### `src/feature/` +**`worker/`** — Worker process management. -Concrete implementations of LSP features. Each feature takes a `CompilationUnitRef` (or `CompilationParams`) and returns the corresponding LSP response data. +- `WorkerPool`: Manages worker process lifecycles and scheduling. Stateful workers (`StatefulWorker`) hold AST and serve query requests. Stateless workers (`StatelessWorker`) execute one-shot tasks (PCH/PCM builds, completion, indexing, etc.) +- Process fault tolerance: Workers are automatically restarted on crash. When a stateful worker crashes, its owned documents are automatically reassigned -This layer is purely computational -- it has no involvement with network communication, state management, or process scheduling. It only concerns itself with "given a compilation result, how to produce the correct LSP response." Features include: - -- Code completion, hover information, signature help -- Semantic highlighting, inlay hints -- Document symbols, document links -- Folding ranges, formatting, diagnostics - -Each feature is implemented as a standalone function following a uniform pattern. When adding a new LSP feature, follow the existing implementations. - -> **Note**: `feature/` only covers single-file, AST-based feature implementations. Cross-file language features (such as go to definition, find references, call hierarchy, etc.) are handled by the Indexer in `src/server/compiler/` using index data. Additionally, some features involve multi-phase processing pipelines and do not necessarily reach the AST layer -- for example, code completion can sometimes be resolved at the syntax layer (include path completion, module import completion), and document links may be extractable during the preprocessing phase. The implementations in `feature/` correspond only to the phase that requires full compilation results. - -## Server Layer - -### `src/server/` - -The language server's core runtime. This is the most complex layer, responsible for assembling all the layers above into a runnable service. - -#### `src/server/protocol/` - -Protocol definitions. Describes the message formats for communication between the master process and worker processes, as well as between the server and clients. - -- **Worker Protocol**: Defines request/response messages for stateful and stateless worker processes -- compilation requests, query requests, build requests, document update notifications, etc. Uses bincode serialization. -- **Extension Protocol**: Extension requests beyond the LSP standard, such as compilation context query and switching. -- **Agentic Protocol**: High-level query interface for AI agents, providing compilation command queries, project file listings, file dependency analysis, impact analysis, symbol search, symbol detail retrieval, and more. These interfaces are accessible directly via the command line without needing an intermediary like MCP. - -#### `src/server/workspace/` - -Project-level global state -- the single source of truth from disk. - -- **Workspace**: Holds the compilation database, toolchain, path pool, dependency graph, PCH/PCM cache, project index, and all other project-level state. Core invariant: unsaved buffer contents of open files never modify the Workspace. Workspace state changes come from three paths: initial load, cascading updates triggered by file saves (didSave), and index merges after background indexing completes. -- **Config**: Loading and validation of TOML/JSON configuration files. - -#### `src/server/compiler/` - -Compilation scheduling and index management. This layer bridges `src/compile/` (low-level compilation abstraction) and `src/server/service/` (service layer). - -- **Compiler**: The scheduler for compilation lifecycles. Coordinates the ordering of PCH builds, module dependency resolution, and AST compilation, dispatching compilation tasks to worker processes. It holds no persistent data -- all data resides in Workspace and Session; the Compiler is solely responsible for orchestrating execution flow. -- **CompileGraph**: The DAG scheduler for C++20 module compilation. Uses reference counting for interest tracking -- when no requestor cares about a module unit, its compilation is automatically cancelled. Supports dependency cascade cancellation and recompilation triggered by file updates. -- **Indexer**: Handles background indexing scheduling and serves as the entry point for cross-file queries. Manages the indexing queue, deduplication, and idle-timeout batch processing. During queries, it combines three data sources: ProjectIndex serves as a directory (locating which files contain a symbol), MergedIndex provides per-file sharded relation data, and in-memory indexes for open files provide real-time coverage of unsaved state. - -#### `src/server/service/` - -Service entry point and session management. - -- **MasterServer**: The top-level coordinator. Holds Workspace, Session map, WorkerPool, Compiler, and Indexer, routing LSP requests to the appropriate handling logic. Manages the server lifecycle (Uninitialized -> Initialized -> Ready -> ShuttingDown -> Exited) and starts file-watching tasks. -- **Session**: The editing state for each open file. Holds the current buffer contents, compilation version number, ABA-protected generation counter, in-memory symbol index, PCH reference, dependency snapshot, and more. Sessions are created on didOpen and destroyed on didClose. Session modifications do not affect other files -- all cross-file dependencies point to on-disk files. -- **LSPClient**: The LSP protocol request handler, registering handlers for each LSP method. -- **AgentClient**: The Agentic protocol request handler. - -#### `src/server/worker/` - -Worker process management. - -- **WorkerPool**: Manages worker process lifecycles, routing, and scheduling. - - **StatefulWorker**: Holds AST in memory and serves query requests (hover, semantic tokens, etc.). Each open file is affinity-bound to a stateful worker by path_id; new files are assigned to the worker with the lightest current load. - - **StatelessWorker**: Executes one-shot tasks (PCH/PCM builds, completion, indexing, etc.). Uses priority-aware scheduling -- interactive requests (completion, signature help) take priority over background tasks (indexing). Concurrency is dynamically adjusted based on memory pressure and worker crash rates. -- **Process Fault Tolerance**: Workers are automatically restarted on crash (with a maximum restart limit); when a stateful worker crashes, its owned documents are automatically reassigned. +See [Multi-process Architecture](multi-process.md). ## Inter-module Relationships @@ -168,14 +112,14 @@ The data flow roughly follows this direction: ```text command (compilation command parsing) - | + ↓ compile (drives Clang compilation) - | -semantic (extracts semantic information) --> index (builds indexes) - | + ↓ +semantic (extracts semantic information) ──→ index (builds indexes) + ↓ feature (produces LSP responses) -The server layer assembles all of the above via Workspace/Session/WorkerPool into a runnable service +The server layer assembles all of the above via Workspace / Session / WorkerPool into a runnable service ``` -`support` and `syntax` are cross-cutting layers shared by multiple modules. `server/protocol` defines the message contracts between processes and between server and client. +`support` and `syntax` are cross-cutting layers shared by multiple modules. diff --git a/docs/en/design/symbol-index.md b/docs/en/design/symbol-index.md new file mode 100644 index 000000000..f4e94dced --- /dev/null +++ b/docs/en/design/symbol-index.md @@ -0,0 +1,246 @@ +# Symbol Index + +## Background + +Many features of a language server need to work across files. A user triggers "go to definition" in one file, and the target may be in any other file in the project; "find references" needs to scan every file in the project that might mention the symbol; call hierarchy and type hierarchy involve chains of symbol relationships across multiple files. To support these features, the language server must maintain a project-wide symbol index that records which files each symbol appears in and the semantic relationships between them (definition, reference, call, inheritance, etc.). + +C++ makes index construction and maintenance particularly difficult. + +The first level of difficulty is the compilation-context problem with header files. C++'s `#include` is textual substitution — the contents of a header file are inserted verbatim into the source file that includes it at compile time. This means the same header file can produce entirely different symbols under different compilation contexts: + +```cpp +// crypto.h +#ifdef USE_OPENSSL + using TLSContext = OpenSSLContext; +#else + using TLSContext = BoringSSLContext; +#endif +``` + +When `crypto.h` is included by a source file that defines `USE_OPENSSL`, `TLSContext` is `OpenSSLContext`; when included by another source file, it is `BoringSSLContext`. Conditional compilation is the most obvious example, but include order and template instantiation can also cause header files to produce different symbol relationships under different contexts. If the index records only one context's results, users will see incorrect jump targets or incomplete reference lists after switching contexts. + +The second level is scale. A mid-sized C++ project (a few thousand source files) produces hundreds of thousands of symbols after compilation; large projects (LLVM, Chromium) reach millions. The index system must maintain reasonable memory usage, build time, and query latency at this scale. + +clangd's handling of both levels is insufficient. On the compilation-context front, clangd's background index stores only the last compilation result for each header — whichever source file is compiled last overwrites the previous index data. If a symbol reference exists only under a particular compilation context that is not the last one indexed, that reference is lost. + +On the cross-file lookup front, clangd's background index stores symbol information in the index shard corresponding to the file where the symbol is declared. This declaration-file-centric storage model prevents reference counts from accumulating correctly across files (clangd [#23](https://github.com/clangd/clangd/issues/23)). Users also frequently encounter incomplete "find references" results — certain references only appear after manually opening the relevant files, at which point the dynamic index fills in the missing data (clangd [#516](https://github.com/clangd/clangd/issues/516), [#802](https://github.com/clangd/clangd/issues/802)). Additionally, when compilation commands change, clangd's staleness detection does not trigger re-indexing (clangd [#199](https://github.com/clangd/clangd/issues/199)), leaving the index data out of sync with the actual compilation state for extended periods. + +clice's index system is redesigned to address these problems: it adopts a three-level index structure separating the global symbol directory from per-file sharded relation data, uses content-addressed deduplication to merge indexes from different compilation contexts, leverages FlatBuffers for on-demand lazy loading to control memory usage, and provides a real-time overlay for open files to ensure query result freshness during editing. + +## Design + +### Symbol Identity + +The index system needs a way to identify the same symbol across files and translation units. clice uses `SymbolHash` (a 64-bit integer) as the unique identifier for each symbol. + +`SymbolHash` is generated from Clang's USR (Unified Symbol Resolution). USR is a canonical string representation of symbol identity that encodes the symbol's fully qualified name: namespace, class name, function signature, template parameters, etc. For example, `std::vector::push_back` and `std::vector::push_back` produce different USRs. `SymbolHash` is the hash of the USR string. + +`SymbolHash` has two key properties. First, cross-file consistency: the same symbol always has the same `SymbolHash` regardless of which file it appears in. `std::string` seen in file A and `std::string` seen in file B have the same hash, allowing all definitions and references to be associated through it — this is the foundation of cross-file navigation. Second, compactness: a 64-bit integer is better suited as a hash table key and for serialized storage than a variable-length USR string. + +### Symbol Occurrences and Relations + +The index stores two kinds of core data: symbol occurrences (`Occurrence`) and symbol relations (`Relation`). + +An `Occurrence` records a symbol's presence at a source location, containing only a source range and the target symbol's `SymbolHash`. It answers the question "what symbol is under the cursor." + +A `Relation` records richer semantic information, consisting of three elements: the relation kind (`RelationKind`), a source location, and a target symbol. Relation kinds cover common inter-symbol semantics: + +- Definition and declaration (Definition, Declaration) +- References (Reference, WeakReference) +- Inheritance (Base, Derived) +- Calls (Caller, Callee) +- Type relationships (Interface, Implementation, TypeDefinition) +- Construction and destruction (Constructor, Destructor) + +The two are stored separately because their query patterns differ. `Occurrence` is indexed by position — given a byte offset, binary search quickly locates the symbol under the cursor. `Relation` is indexed by `SymbolHash` — given a symbol, look up all its definitions, references, and call relationships. These two queries have contradictory sorting requirements; separate storage allows both to execute efficiently. + +### Three-Level Index Hierarchy + +clice's index is organized into three levels, each with a different lifecycle and responsibility: + +``` +TUIndex Raw artifact from a single compilation, discarded after merging + ↓ merge +ProjectIndex Global symbol directory (which files a symbol appears in), resident in memory +MergedIndex Per-file sharded relation data (exact positions and relations), loaded on demand + ↑ overlay +FileIndex Real-time overlay for open files (from in-memory AST) +``` + +**TUIndex** is the raw index data produced by compiling a translation unit. `SemanticVisitor` traverses the AST, generating `Occurrence` and `Relation` records for each symbol, organized by file into a `TUIndex`. Since a compilation involves the main file and all included headers, `TUIndex` internally maintains a separate `FileIndex` for each file involved. `TUIndex` also contains a `SymbolTable` (mapping symbol hashes to names and kinds) and an `IncludeGraph` (include relationships from this compilation). `TUIndex` is transient data, discarded after being merged into the persistent indexes. + +**ProjectIndex** is the global symbol directory. It aggregates symbol information from all indexed translation units, maintaining a global symbol table: `SymbolHash` → symbol name, symbol kind, reference file bitmap. The reference file bitmap records which files the symbol appears in, stored using Roaring Bitmap compression. + +`ProjectIndex` does not store exact symbol positions (offsets, line numbers). Its role is that of a "directory" — it tells you which files a symbol exists in, then you look up the exact positions in the corresponding `MergedIndex` shard. This separation keeps `ProjectIndex` compact enough to reside in memory at all times. + +**MergedIndex** is the per-file sharded index storage layer. Each file in the project corresponds to one `MergedIndex` shard, storing all symbol occurrence positions and relation information for that file. This is the largest part of the index system by volume and the layer that actually serves queries. It supports lazy loading from disk — shards that are never queried need not be loaded into memory. + +`MergedIndex`'s core capability is merging and deduplicating index data from different compilation contexts of the same file, detailed in the Implementation section. + +**FileIndex** (the open-file overlay) resides in each open file's `Session`, produced by the most recent in-memory compilation. It stores the same types of data as a `MergedIndex` shard (`Occurrence` and `Relation`), but is never written to global state — it is used only as an overlay during queries, overriding disk-indexed data with results from the current edit buffer's compilation. + +### Symbol Table + +`SymbolTable` maps `SymbolHash` to symbol metadata — name and kind (Class, Function, Variable, etc.). It appears in two places: the global symbol table in `ProjectIndex` and the local symbol table in each `Session`. When looking up a symbol's name, the `Session` is checked first (more current); if not found, `ProjectIndex` is consulted. + +### IncludeGraph + +`IncludeGraph` records the include relationships from a single compilation. It consists of two parts: a path list (all file paths involved in the compilation) and `IncludeLocation` records (each indicating a file was included at a certain line, along with the source of the inclusion). + +`IncludeGraph` serves two purposes. During `TUIndex` merging, it provides the mapping from compilation-unit-internal file IDs to project-global path IDs. Within `MergedIndex`, it is stored as part of the compilation context, with the include chain used for staleness detection. + +## Implementation + +### Index Construction + +`TUIndex` construction is performed by `SemanticVisitor`: given a compilation unit, it traverses the AST, generating `Occurrence` and `Relation` records for each named declaration and macro. After traversal, each file's data is deduplicated and sorted — `Occurrence` entries are sorted by position to support binary search, `Relation` entries are sorted by kind and position for efficient filtering. + +During construction, the main file's (source file's) `FileIndex` is extracted separately. This allows different treatment during merging — the main file is merged as a source-file context, while other files are merged as header contexts. + +### Index Merging + +`TUIndex` merging into the persistent indexes proceeds in two steps. + +Step one: symbol information is merged into `ProjectIndex`. All symbols from the `TUIndex` are inserted into the global symbol table, and each symbol's reference file bitmap is updated — file IDs involved in this compilation are added to the bitmap. Path mapping is also completed at this step: path IDs internal to the `TUIndex` are converted to `ProjectIndex`'s global path IDs. + +Step two: each file's `FileIndex` is merged into the corresponding `MergedIndex` shard. For the main file, compilation context information (build timestamp, include chain) is attached; for header files, header context information (include location identifier) is attached. + +### Compilation-Context Deduplication + +The core problem for `MergedIndex` is: the same header file is included by N source files, producing N `FileIndex` entries. If each were stored in full, storage would grow linearly with the number of translation units. But in practice, the vast majority of headers produce identical index data under different compilation contexts — the same symbols appear at the same positions, producing the same relations. Only headers like the `crypto.h` example above, affected by conditional compilation, produce different index content under different contexts. + +`MergedIndex` solves this through content-addressed deduplication. Each `FileIndex` has its SHA-256 content hash computed before merging. `FileIndex` entries with the same hash have identical content and share the same canonical ID (an auto-incrementing integer identifier). + +Specifically, `Occurrence` and `Relation` entries inside `MergedIndex` are not simple lists — each entry is associated with a Roaring Bitmap recording which canonical IDs it belongs to. When a new `FileIndex` is merged in: + +1. Compute its SHA-256 hash +2. Check the cache: if this hash already exists, the data is identical — reuse the existing canonical ID and increment its reference count +3. If it is a new hash, allocate a new canonical ID, insert all `Occurrence` and `Relation` entries, and associate them with this new ID + +When a compilation context is removed (e.g., a source file is deleted from the project), the corresponding canonical ID's reference count is decremented. Canonical IDs whose reference count reaches zero are marked into the "removed" set. During queries, data belonging to the removed set is filtered out. + +This design means storage depends on the number of distinct index contents rather than the number of compilation contexts. For most headers, regardless of how many source files include them, only one copy of the data is stored. + +### Compilation Context Types + +`MergedIndex` internally distinguishes two types of compilation contexts: + +- **CompilationContext**: Produced when a file is compiled directly as a source file. Records the build timestamp and include chain (for staleness detection), along with the corresponding canonical ID. A file can have multiple `CompilationContext` entries, corresponding to different compilation commands in the compilation database. +- **HeaderContext**: Produced when a file is included as a header by another source file. Records the including source file and include location, along with the corresponding canonical ID. + +These two context types work in concert with the [compilation context](compilation-context.md) system. During queries, context types need not be distinguished — all context data has already been unified through canonical ID Bitmaps. During staleness detection, the `CompilationContext`'s include chain is used to determine whether re-indexing is needed. + +### Lazy Loading + +`MergedIndex` is serialized using FlatBuffers. FlatBuffers' design allows queries to be executed directly on serialized data without deserializing into in-memory structures. `MergedIndex` leverages this to implement a two-tier access model: + +- **Read-only path**: After loading from disk, `MergedIndex` remains as the raw memory-mapped buffer. Query operations execute directly on the FlatBuffers data, with zero deserialization overhead. +- **Read-write path**: When modifications are needed (merging new data or removing old contexts), the FlatBuffers data is first deserialized into in-memory structures, and subsequent operations are performed on those structures. Modified shards are re-serialized when saved. + +At startup, only `ProjectIndex` (relatively compact) needs to be loaded. `MergedIndex` shards are loaded on demand, and most shards are never accessed in a single session. + +### Query Flow + +Using "find references" as an example to illustrate the full cross-file query flow: + +1. In the current file, use the cursor's byte offset to binary-search the `Occurrence` list and obtain the `SymbolHash` of the symbol under the cursor +2. Look up the `SymbolHash`'s reference file bitmap in `ProjectIndex` to get all files containing the symbol +3. Query each file in the list individually: + - If the file is currently open (has an active `Session`), use the `Session`'s `FileIndex`, skipping the corresponding `MergedIndex` shard + - If the file is not open, load the corresponding `MergedIndex` shard and look up relations in it +4. Aggregate all `Relation` entries found across files (filtering by the target `RelationKind`), convert to LSP positions, and return to the client + +In step 3, open files preferentially use the `Session`'s `FileIndex` rather than `MergedIndex`, because the buffer content may differ from disk. The `Session`'s `FileIndex` comes from in-memory compilation results that more accurately reflect the code the user is currently seeing. Only when the `Session`'s AST is in a dirty state (the user has edited but the file has not been recompiled) does the system fall back to `MergedIndex`. + +> Converting offsets to LSP positions requires the file content and a line-start offset table. `MergedIndex` shards also store the corresponding file's content and line-start table, so this conversion can be performed even for files that are not open. + +### Staleness Detection + +Staleness detection determines whether a file needs to be re-indexed. The `MergedIndex` shard stores the build timestamp and include chain. During detection, the last modification time (mtime) of each file in the include chain is checked. If any file's mtime is later than the build timestamp, the dependency has been updated and re-indexing is needed. + +This detection is conservative — an mtime change does not necessarily mean the content changed (e.g., a `touch` operation, branch switching). But a false positive only results in one extra indexing pass, never a missed update. + +### Background Indexing Scheduling + +Background indexing scheduling must balance index timeliness against interference with user interaction. The index module employs the following strategies: + +- **Queue with idle delay**: Files that need indexing are added to a queue, and processing begins only after the editor has been idle for a configurable period. This avoids triggering index tasks during rapid editing. +- **Concurrency control with memory monitoring**: The number of concurrent index tasks has a configurable upper limit. During indexing, system memory usage is dynamically monitored — concurrency is automatically reduced under memory pressure and gradually restored when memory recovers. +- **Priority management**: User-initiated operations (such as compiling an open file) pause background indexing. Indexing resumes after the operation completes, ensuring user request latency is not affected by background indexing. +- **Result merging and persistence**: Each index task compiles a file and builds a `TUIndex` in a stateless subprocess. The result is serialized and sent back to the main process, which merges it into `ProjectIndex` and `MergedIndex`. After indexing completes, modified shards are written back to disk so they can be loaded directly on the next startup. + +## FAQ + +- **Why separate `ProjectIndex` and `MergedIndex` instead of using a single unified index?** + + If position information were also stored in `ProjectIndex`, its size would balloon dramatically, making it impossible to keep in memory. Without `ProjectIndex`, every cross-file query would need to traverse all `MergedIndex` shards to locate files containing the symbol — in a project with tens of thousands of files, loading that many shards is unacceptable. `ProjectIndex` serves as a lightweight directory layer that first narrows the search scope to a handful of specific files, then precise lookup happens in the corresponding shards. + +- **Why don't open-file indexes write to global state?** + + Buffer content being edited by the user may be incomplete code with syntax errors. If this temporary state were written to the global index, it would pollute query results for other files. For example, a symbol definition temporarily disappearing from a header being edited would affect find-references results for every file that references that symbol. The global index only accepts stable state saved to disk, built through background indexing from disk files. + +- **Is there a hash collision risk with content-addressed deduplication?** + + Theoretically SHA-256 collisions are possible, but the probability is negligible (on the order of 2^-128). In practice, treating SHA-256 collisions as "will not happen" is standard. Even if a collision occurred, the only consequence would be two different `FileIndex` entries sharing data — it would not cause a crash or data corruption. + +- **Why FlatBuffers rather than Protocol Buffers or a custom format?** + + FlatBuffers allows queries to be executed directly on serialized data without deserializing first. For data like `MergedIndex` that may have thousands of shards, most shards are never accessed in a single session. FlatBuffers' zero-copy property makes loading a shard nearly free — only a memory-mapped file is needed, and only actually accessed data is read into memory. Protocol Buffers requires a full deserialization step, making it unsuitable for this on-demand loading model. + +- **Why store `Occurrence` and `Relation` separately?** + + `Occurrence` is indexed by position — given an offset, binary search locates the symbol under the cursor, requiring position-sorted data. `Relation` is indexed by `SymbolHash` — given a symbol, look up all its relationships, requiring symbol-grouped data. Combining them into a single data structure would inevitably sacrifice efficiency in at least one of the two query patterns. + +- **Why are cross-file queries performed in the main process rather than subprocesses?** + + clice uses a multi-process architecture where each open file is compiled in its own stateful subprocess. Cross-file queries (such as find-references) need to iterate over all open files' `Session` instances and aggregate their `FileIndex` results. If these Sessions were spread across different subprocesses, each query would require cross-process communication with multiple workers and then result aggregation — unacceptable in both latency and complexity. Instead, subprocesses send their `FileIndex` back to the main process after compilation, and the main process performs queries uniformly — it can access all open files' `FileIndex` entries as well as `ProjectIndex` and `MergedIndex`, completing the full query flow in a single process. + + This design also has a semantic consideration: query results for open files should reflect the editor's buffer state, not the disk state. Even if a file on disk has been modified by an external tool, as long as the editor has not sent a `didChange` notification, query results should be based on the version the editor holds. Centralizing all `FileIndex` entries in the main process makes this semantic invariant easier to maintain. + +- **Why not use a database for index storage?** + + clice needs to persist multiple types of cache files: index shards, PCH, PCM, etc. PCH and PCM files are large (potentially hundreds of MB) but few in number (roughly proportional to the number of open files or modules), and have a simple lifecycle — create, read, delete when stale. There are no complex query or transaction requirements. The capabilities databases excel at (transactions, indexing, complex queries) are irrelevant for these files; filesystem management is sufficient. + + Index shards are the only part that could potentially benefit from a database: numerous (equal to the number of project files), small in size, and could benefit from atomic writes and automatic LRU eviction. But the current filesystem-based approach already handles these needs adequately. Whether the added complexity of introducing a database dependency is justified for this one use case needs to be evaluated when an actual bottleneck is encountered. For very large projects (tens of thousands of files), storing that many index shards in a single directory may create filesystem-level pressure; hierarchical storage or a lightweight database could be considered in the future. + +## Known Limitations + +- **Symbol table locality**. Currently all symbols (including function-local variables) are merged into `ProjectIndex`'s global symbol table. This causes a large number of symbols meaningful only within a single file to be stored globally, increasing hash table insertion overhead during merging and memory usage. + + The improvement direction is to introduce multi-level symbol tables — not only `ProjectIndex` should have a `SymbolTable`, but `MergedIndex` shards should have their own as well. The rule for determining which level a symbol belongs to is: a symbol belongs to the `SymbolTable` of the file where it is defined, provided it is internal (will not be referenced by other files). For example: + + ```cpp + // utils.h + inline int helper(int x) { + auto temp = x * 2; // temp is a local symbol of utils.h + return temp + 1; + } + ``` + + ```cpp + // main.cpp + #include "utils.h" + static int counter = 0; // counter is a local symbol of main.cpp + + int main() { + counter = helper(42); + } + ``` + + `temp` is defined in `utils.h` and will never be referenced by any other file — it should be in the `SymbolTable` of `utils.h`'s `MergedIndex` shard, not in `ProjectIndex`'s global symbol table. `counter` is a static variable in `main.cpp` and similarly should be in `main.cpp`'s `MergedIndex` shard. Note that although `temp` appears in a header file, it belongs to the header's `SymbolTable` rather than the including source file's `SymbolTable`, because it is defined in the header. Only symbols like `helper` and `main` that may be referenced cross-file need to enter `ProjectIndex`. + + The goal is to minimize `ProjectIndex`'s size and merging overhead, while avoiding duplicate storage of the same local symbol across multiple translation units. + +- **Staleness detection precision**. The current staleness detection uses only mtime — re-indexing is triggered whenever a dependency file's mtime is later than the build timestamp. This produces unnecessary re-indexes in scenarios like `touch`, branch switching, or CI restores (file mtime changed but content is actually unchanged). The improvement direction is mtime + content hash dual-layer detection: the first layer uses mtime for a fast check — if unchanged, skip immediately (zero I/O); the second layer computes the content hash for files whose mtime changed — if the hash is unchanged, the content was not actually modified and can also be skipped. This approach is already used for compilation artifact staleness detection (PCH, AST); the index staleness detection should be aligned. + +- **Fuzzy symbol search**. The current workspace symbol search (workspace/symbol) is a simple substring match that does a linear scan over all symbols in `ProjectIndex`. This is insufficient for large projects and does not support fuzzy matching. + + C++ symbol names have structure: `getSymbolHash` is camelCase, `get_symbol_hash` is snake_case, `std::vector::push_back` has namespace qualification. When searching, users typically type abbreviations or fragments (e.g., `symhash`, `gSH`, `vec_pb`), expecting them to match the full symbol name. Substring matching cannot handle these queries. + + The improvement direction is to build a dedicated search index over symbol names. A tokenizer is needed to split symbol names by naming conventions (`getSymbolHash` → `[get, Symbol, Hash]`, `push_back` → `[push, back]`), then build an inverted index over the tokens. For example, trigrams (three-character groups) can be used as index keys, and at query time trigram intersections produce a candidate set that is then scored precisely. clangd's Dex index uses this trigram posting list approach and serves as a useful reference implementation. Another direction is to adopt a mature full-text search library, though the cost of introducing an external dependency needs to be evaluated. + +- **PCH-induced index split**. When using PCH (precompiled header) optimization, a file's compilation is effectively split into two phases: first the preamble (the `#include` directives at the top of the file) is compiled to produce the PCH, then the PCH is used to compile the rest of the file. The PCH itself is a compilation unit and produces its own index data. + + This split affects the index. Take document links (clickable `#include` directives in the editor) as an example: `#include` directives in the preamble belong to the PCH compilation phase, and the main file's compilation cannot see them. Since the index system does not currently store document link information, it cannot reconstruct the PCH portion's results through the index. The current workaround is to pre-serialize the PCH's document links as JSON during PCH construction and store it in the PCH metadata, then manually splice it into the main file's results at query time. This works but is not clean. + + A better approach would be to incorporate document links into the PCH metadata system (which already stores dependency file lists and other information), or to leverage the include relationship information already present in the index to reconstruct document links. The latter has the problem that index construction takes time, and after PCH compilation completes it should be put to use as quickly as possible — waiting for indexing to complete would add latency. Neither approach is fully implemented yet. diff --git a/docs/en/index.md b/docs/en/index.md index 4986b9107..53db59d14 100644 --- a/docs/en/index.md +++ b/docs/en/index.md @@ -20,16 +20,16 @@ hero: alt: clice features: - - icon: T - title: Better Template Handling - details: Use pseudo-instantiation to handle dependent template names, with code completion even for complex templates - - icon: H - title: Header File Context - details: Support header file state switching between different source file contexts, and fully support non-self-contained files - - icon: M - title: Modules - details: Excellent C++20 module support, from code completion to highlighting to navigation, all adapted - - icon: I - title: Better Performance - details: Excellent asynchronous task scheduling, support for compilation task cancellation, caching necessary information, avoiding meaningless CPU waste + - icon: 📝 + title: Compilation Context + details: The first language server to introduce compilation context as a formal concept. Users can query and switch compilation contexts, with support for non-self-contained headers and multi-configuration projects + - icon: 📦 + title: C++20 Modules + details: Reference-counted real-time module compilation DAG with cancellation and dependency cascading. Code completion, semantic highlighting, and go-to-definition fully adapted for module syntax + - icon: 🔍 + title: Template Resolution + details: Resolves dependent names through pseudo-instantiation, providing accurate code completion and navigation even inside template definitions + - icon: ⚡ + title: Multi-Process Architecture + details: Master + Worker process model isolating Clang crashes and memory leaks. Supports priority scheduling, real-time memory monitoring, and automatic process recovery --- diff --git a/docs/en/sidebar.yaml b/docs/en/sidebar.yaml index 2b7ad08fd..fa57283b6 100644 --- a/docs/en/sidebar.yaml +++ b/docs/en/sidebar.yaml @@ -28,17 +28,17 @@ design: items: - overview - compilation-context - - command - - index-design + - command-resolve + - symbol-index - multi-process - - module - - incremental + - module-graph + - incremental-parse - template-resolver - dependency-scanning dev: label: Development - collapsed: true + collapsed: false items: - build - contribution diff --git a/docs/zh/design/command-resolve.md b/docs/zh/design/command-resolve.md new file mode 100644 index 000000000..92f3c4222 --- /dev/null +++ b/docs/zh/design/command-resolve.md @@ -0,0 +1,171 @@ +# 命令解析 + +## 背景 + +C++ 语言服务器需要知道如何编译项目中的每一个文件。这些信息来自编译数据库(compilation database,简称 CDB),通常是构建系统生成的 `compile_commands.json` 文件。CDB 中的每条记录包含一个源文件路径和一条编译命令,例如: + +```bash +g++ -std=c++20 -O2 -fPIC -I../include -DNDEBUG -c src/foo.cpp +``` + +这条命令记录了构建系统在实际编译时调用编译器的方式。然而,语言服务器不能直接使用它,原因有三个。 + +**第一,这是驱动级命令,不是前端命令。** 上面的 `g++` 是编译器驱动程序(driver),它负责选择正确的前端、链接器和标准库。语言服务器实际使用的是 Clang 的前端(cc1),而从 `g++` 到 cc1 的转换涉及大量隐式操作:确定目标三元组(target triple)、注入系统头文件搜索路径、设置默认语言标准等。这些信息在 CDB 中不会显式出现——它们是编译器隐式提供的。 + +**第二,命令中混杂了语义无关的选项。** `-O2` 和 `-fPIC` 只影响代码生成,不影响语义分析——语言服务器不做代码生成,也不需要这些选项。类似地,`-c`(编译模式)和 `-o`(输出文件)是构建产物相关的指令,对语言服务器毫无意义。如果不加过滤地传给前端,它们会增加不必要的复杂度,甚至引起错误。 + +**第三,大型项目中的命令高度冗余。** 一个拥有上万源文件的项目,绝大多数文件使用相同的编译器和语义选项,只有 include 路径和宏定义因文件而异。不做去重意味着每个文件都要独立查询工具链,浪费内存和启动时间。 + +这三个问题中,最影响用户体验的是第一个:隐式信息的缺失。当语言服务器无法正确获取系统头文件路径时,用户会看到标准库头文件报错——`#include ` 报 "file not found",或者 GCC 内置的 type traits 被标记为未声明标识符。在 clangd 的 issue 中,这类问题长期占据最高频率([clangd#1262](https://github.com/clangd/clangd/issues/1262)、[clangd#1691](https://github.com/clangd/clangd/issues/1691))。 + +clangd 的解决方案是 `--query-driver` 参数:用户手动指定哪些编译器需要探测,clangd 再去查询这些编译器获取系统路径。这个方案有两个问题:首先,它是手动的——用户必须知道自己的项目使用了哪个编译器,还要正确配置 glob 模式。其次,它作为一个 flag 并不独立可用——它依赖 CDB 中已有的命令来触发探测([clangd#1219](https://github.com/clangd/clangd/issues/1219))。对于交叉编译、嵌入式开发等场景,用户经常需要反复调试才能让 clangd 正确识别工具链。 + +clice 将整个命令处理流程自动化:从 CDB 中读取命令后,自动识别编译器家族(GCC、Clang、MSVC 等),自动探测工具链信息,并通过多级去重将启动开销降到最低。用户不需要手动配置任何工具链相关的参数。 + +## 设计 + +命令处理的核心任务是将 CDB 中的原始驱动命令转换为 Clang 前端可消费的 cc1 参数。这个转换涉及四个概念层次:参数分类、命令分离、工具链探测、搜索路径提取。 + +### 参数分类 + +加载 CDB 时,每个编译选项被分为四类: + +- **丢弃(Discarded)**:与构建产物相关的选项,语言服务器不需要。包括输出文件(`-o`)、编译模式(`-c`)、依赖扫描(`-M` 系列)、PCH 构建(`-emit-pch`)、C++20 模块(`-fmodule-file` 等,由语言服务器自行管理)。 + +- **仅代码生成(Codegen-only)**:只影响代码生成后端、不影响语义分析的选项。包括位置无关代码(`-fPIC`)、栈保护(`-fstack-protector`)、帧指针(`-fomit-frame-pointer`)、调试信息(`-g` 系列)、LTO 等。这些不会改变 AST 或诊断结果。 + + > 注意:看起来像代码生成选项的 `-O` 和 `-fsanitize=` 不在此列。`-O` 会定义 `__OPTIMIZE__` 宏,`-fsanitize=address` 会影响 `__has_feature(address_sanitizer)` 的结果。它们改变预处理器状态,属于语义选项。 + +- **用户内容(User-content)**:每个文件可能不同的选项,但不影响工具链探测结果。包括 include 路径(`-I`、`-isystem`、`-iquote`、`-idirafter`)、宏定义(`-D`、`-U`)和强制包含(`-include`)。这些选项的含义是"这个特定文件额外需要什么"。 + +- **语义选项(Semantic)**:剩余的所有选项,它们影响编译语义并在工具链探测中起作用。例如 `-std=c++20`、`-Wall`、`-target`、`-march=` 等。 + +分类基于 Clang 自身的选项表(`OptTable`),按选项 ID 进行判断,而不是按字符串匹配。 + +### 命令分离 + +分类完成后,每条编译命令被拆分为两部分: + +- **`CanonicalCommand`**:驱动程序路径 + 所有语义选项。代表"编译器的身份和语义配置"。 +- **Patch**:所有用户内容选项。代表"这个文件额外需要的 include 路径和宏定义"。 + +两者组合加上工作目录构成 `CompilationInfo`,是一个文件完整编译配置的抽象表示。 + +这个分离的核心目的是工具链探测的缓存效率。工具链探测需要实际调用编译器驱动(如运行 `g++ -dumpmachine` 或 `clang++ -###`),耗时通常在 100ms 以上。探测结果只取决于驱动程序和语义选项——用户内容选项(`-I`、`-D`)不会影响驱动输出的系统路径或目标三元组。因此,无论文件有什么不同的 `-I` 路径,只要语义选项相同,就可以共享同一个探测结果。 + +在实际项目中,上万个文件可能只有几十个不同的 `CanonicalCommand`,这意味着工具链只需要探测几十次而非上万次。 + +`CanonicalCommand`、`CompilationInfo` 都通过 `ObjectSet` 去重——内容相同的实例在内存中只存在一份,由指针共享。字符串参数通过 `StringSet` 内部化,保证指针稳定且可直接比较。 + +### 编译数据库 + +`CompilationDatabase` 负责加载 `compile_commands.json`,将每条记录解析、分类、去重后存储为 `CompilationEntry`(文件路径 ID → `CompilationInfo`)。所有条目按文件路径 ID 排序,支持二分查找。 + +查找时,`CompilationDatabase` 将 `CompilationInfo` 组装为 `CompileCommand`——这是命令处理流水线的最终输出,包含完整的编译选项和源文件路径,可直接提交给工具链探测或 Clang 前端。 + +对于没有 CDB 条目的文件(例如用户打开了一个不在项目中的文件),`CompilationDatabase` 会合成一个默认命令——根据文件扩展名选择 `clang` 或 `clang++ -std=c++20`。 + +`CompilationDatabase` 还提供了按配置分组的能力:`ConfigGroup` 将共享相同 `CompilationInfo` 的文件聚合在一起。这是依赖扫描中提取搜索路径配置的正确粒度——不同的 `-I` 路径产生不同的分组。对于工具链探测,粒度更粗(用户内容选项不影响探测结果),因此 `Toolchain` 会在 `ConfigGroup` 的基础上进一步去重。 + +### 配置规则 + +除了 CDB 本身的编译命令,用户可以通过 `clice.toml` 中的 `[[rules]]` 配置追加或移除编译选项。每条规则包含文件匹配模式(glob)和要追加(append)/移除(remove)的选项列表。 + +在查找文件的编译命令时,匹配的规则会被应用到 CDB 命令之上——先从基础命令中移除指定的选项,再追加新选项。这使得用户可以项目级地微调编译参数,而不需要修改构建系统的输出。 + +### 工具链 + +`Toolchain` 负责将驱动级命令转换为 cc1 参数。它的设计围绕两个核心能力: + +**编译器家族识别。** `Toolchain` 通过可执行文件名识别编译器家族——`CompilerFamily` 枚举包括 GCC、Clang、MSVC、ClangCL、NVCC、Intel、Zig。识别规则处理了各种命名变体:版本后缀(`clang++-17`)、架构前缀(`arm-none-eabi-g++`)、Windows 的 `.exe` 后缀等。家族信息决定了后续使用哪种探测策略。 + +**缓存策略。** 探测结果以 (驱动路径, 文件扩展名, 非用户内容选项) 为键缓存。文件扩展名参与缓存键,因为 `.c` 和 `.cpp` 可能触发不同的驱动规则。失败的探测也会被缓存(负缓存),避免对同一个不存在的编译器重复尝试。 + +### 搜索路径 + +`SearchConfig` 从 cc1 参数中提取头文件搜索路径,组织为四段式结构: + +1. **Quoted**(`-iquote`):`#include "foo.h"` 的搜索路径 +2. **Angled**(`-I`):`#include ` 的搜索路径 +3. **System**(`-isystem`、`-internal-isystem` 等):系统头文件路径 +4. **After**(`-idirafter`):在系统目录之后搜索的路径 + +这个四段模型对应 Clang 内部的搜索布局。段内路径会去重(从 Angled 段开始),去重算法复制了 Clang 的行为:如果同一路径出现在 Angled 和 System 两个段中,保留 Angled 中的那个。这确保了 `#include_next` 的正确性。 + +`SearchConfig` 是 include 路径补全、include 路径解析、[依赖图构建](dependency-scanning.md)等功能的基础输入。 + +## 实现 + +### 加载与解析 + +CDB 加载使用 simdjson 流式解析 JSON,逐条处理: + +1. 读取每条记录的 `directory`、`file`、`arguments`(或 `command`)字段 +2. 过滤非 C/C++ 文件(如 `.rc`、`.asm`、`.def`) +3. 将相对文件路径解析为绝对路径 +4. 对 `arguments` 字段中的每个选项进行分类,分别放入 canonical 和 patch +5. 将 include 路径选项中的相对路径绝对化(基于 `directory` 解析) +6. 通过 `ObjectSet` 去重 `CanonicalCommand` 和 `CompilationInfo` +7. 所有条目按文件路径 ID 排序 + +解析过程中还处理了一个特殊情况:CMake 生成的 CDB 中有时会包含 `-Xclang -include-pch -Xclang ` 序列(CMake 的 PCH 变通方案),加载时会识别并丢弃这个模式。 + +### 工具链探测 + +不同编译器家族使用不同的探测策略: + +**GCC**:分两步。第一步,调用 GCC 驱动获取两个关键信息——目标三元组(`-dumpmachine`)和安装路径(`-print-search-dirs`)。第二步,将这两个信息注入 Clang 的驱动(`--target=` 和 `--gcc-install-dir=`),让 Clang 驱动模拟 GCC 的行为,获取 cc1 参数。这样 Clang 前端就能正确找到 GCC 的标准库和系统头文件。 + +**Clang / Zig**:调用驱动的 `-###` 选项,它会打印出完整的 cc1 命令行而不实际执行编译。解析输出中的第一行 cc1 命令即可。对于 Zig,驱动路径包含两部分(`zig cc` 或 `zig c++`),探测时需要特殊处理。 + +> 外部驱动的版本可能比 clice 内嵌的 LLVM 版本更新,输出的 cc1 参数中可能包含 clice 不认识的选项。解析时会将未知选项静默丢弃,保证兼容性。 + +**MSVC / ClangCL**:通过 `--driver-mode=cl` 指令切换 Clang 驱动到 MSVC 兼容模式,然后用 Clang 驱动获取 cc1 参数。 + +探测完成后,会从结果中移除临时探测文件的路径和模块输出相关的选项(它们引用了已删除的临时文件)。如果 clice 的 resource dir 与探测结果中的不一致,还会替换所有相关路径,确保前端使用匹配版本的内置头文件。 + +**启动预热。** 在服务器启动的依赖扫描阶段,会收集所有唯一的工具链缓存键并发起并行探测。探测子进程的管道读取和进程等待也是并发执行的,避免管道阻塞导致的死锁。预热完成后,后续所有的工具链查询都命中缓存。 + +### 搜索路径提取 + +工具链探测得到 cc1 参数后,搜索路径提取遍历这些参数,将 include 路径选项按类型分入四个段。所有路径被解析为绝对路径并规范化(消除 `.` 和 `..`)。 + +`-iprefix` / `-iwithprefix` / `-iwithprefixbefore` 三个选项需要按出现顺序配合处理:`-iprefix` 设置前缀,后续的 `-iwithprefix` 将前缀拼接到路径前再放入 After 段,`-iwithprefixbefore` 则放入 Angled 段。 + +四段拼接完成后,从 Angled 段开始去重。Quoted 段不参与去重——同一路径同时出现在 Quoted 和 Angled 中是合法的,两者都会被保留。这复制了 Clang 内部的行为。 + +## FAQ + +- **为什么按选项 ID 分类,而不是字符串匹配?** + + clice 的参数解析基于 Clang 自身的选项表(从 `Options.inc` 生成),通过选项 ID 进行分类判断。字符串匹配容易遗漏边界情况:Clang 的选项语法有多种形式——`-std=c++20` 是 joined 形式,`-I /path` 是 separate 形式,`-Wall` 是 flag 形式,有些选项还有 `/` 前缀(MSVC 风格)。按 ID 分类意味着所有这些语法变体都被 Clang 的解析器正确处理,分类逻辑只需要关心"这个选项是什么",不需要关心"它是怎么拼写的"。 + +- **为什么用户内容选项不参与工具链缓存键?** + + 这是两级分离设计的核心收益。`-I` 和 `-D` 不会改变编译器驱动输出的系统路径、目标三元组或语言默认设置。将它们从缓存键中排除,使得缓存键的数量从"每文件一个"降低到"每配置一个"。一个上万文件的项目通常只有几十个不同的缓存键,启动时只需要几十次子进程调用。 + +- **为什么搜索路径去重要精确匹配 Clang 的行为?** + + 如果语言服务器的头文件搜索顺序与实际编译器不一致,可能导致 `#include` 解析到不同的文件——同名头文件在不同目录中存在时,搜索顺序决定了使用哪一个。`#include_next` 的语义更是直接依赖于搜索目录的去重结果。严格匹配确保语言服务器看到的代码与编译器看到的完全一致。 + +- **为什么 GCC 探测要分两步,而不是直接用 GCC 的 `-v` 输出?** + + Clang 前端需要知道 GCC 的目标三元组和安装路径,才能正确找到 GCC 的标准库头文件。直接解析 GCC 的 `-v` 输出虽然可行,但会引入额外的文本解析逻辑且容易因不同 GCC 版本的输出格式变化而出错。通过 `--target=` 和 `--gcc-install-dir=` 将信息注入 Clang 驱动,让 Clang 自身完成搜索路径的组装,更加可靠。 + +- **为什么编译器家族识别基于可执行文件名而不是路径解析?** + + 编译器驱动的行为受调用名称影响。例如,`/usr/bin/clang++` 通常是 `/usr/lib/llvm-20/bin/clang` 的符号链接,但以 `clang++` 名义调用时会自动启用 C++ 模式并链接 C++ 库。如果使用 `realpath` 解析到真实路径后再判断,会丢失调用名称携带的语义信息。同样,`arm-none-eabi-g++` 如果被解析成某个通用 GCC 二进制的路径,交叉编译的上下文就会丢失。 + +- **CDB 中没有条目的文件怎么处理?** + + 合成一个默认命令。根据文件扩展名选择 `clang` 或 `clang++ -std=c++20`,并注入 resource dir。这确保即使文件不在 CDB 中,基本的语义分析仍然可用。对于头文件,还会尝试通过依赖图找到包含它的源文件,使用该源文件的编译命令作为上下文(详见[编译上下文](compilation-context.md))。 + +## 已知局限 + +- **部分编译器家族支持不完整。** NVCC 和 Intel 编译器(`icc`、`icx`、`dpcpp`)虽然被识别,但目前回退到通用的 Clang 驱动路径,没有专门的探测逻辑。这意味着这些编译器的特殊系统路径可能无法被正确发现。 + +- **SearchConfig 不支持部分搜索路径选项。** `-cxx-isystem`(仅 C++ 模式生效的系统目录)、`-iwithsysroot`(拼接 sysroot 前缀)和 HeaderMap 支持尚未实现。这些选项在实际项目中不常见,但可能在特定的 Apple 或交叉编译工具链中出现。 + +- **配置规则的全局影响。** `clice.toml` 中的 `[[rules]]` 可以向编译命令追加或移除选项。如果用户修改了影响所有文件的规则(例如追加一个全局的 `-I`),所有文件的编译配置都会改变,可能触发全量重索引。目前没有机制检测哪些规则变更实际影响了哪些文件。 + +- **MSVC 兼容模式的选项解析。** 在非 Windows 系统上,需要特别处理 MSVC 风格的选项前缀(`/U`、`/D`、`/I`),避免 Unix 绝对路径(如 `/Users/...`)被误解析为 MSVC 选项。目前通过根据驱动名称动态调整选项可见性来解决,但边界情况仍可能存在。 diff --git a/docs/zh/design/command.md b/docs/zh/design/command.md deleted file mode 100644 index accc5fa89..000000000 --- a/docs/zh/design/command.md +++ /dev/null @@ -1,97 +0,0 @@ -# 命令处理 - -## 背景 - -语言服务器需要知道如何编译每一个文件。这些信息来自编译数据库(`compile_commands.json`),由构建系统(CMake、Bazel、Meson 等)生成。然而,CDB 中的原始编译命令不能直接交给 Clang 前端使用,存在以下问题: - -**编译命令是驱动级别的,不是前端级别的。** CDB 记录的是构建系统调用编译器的命令(如 `g++ -std=c++17 -O2 -c foo.cpp`),但语言服务器需要的是 Clang 前端(cc1)的参数。从驱动命令到 cc1 参数的转换需要查询编译器工具链,获取目标三元组、系统头文件搜索路径等信息。 - -**编译命令包含语义无关的选项。** 构建系统会传入仅用于代码生成的选项(如 `-fPIC`、`-fomit-frame-pointer`),这些不影响语义分析但增加复杂度。同时也包含构建产物相关的选项(如 `-o foo.o`、`-emit-pch`),这些对语言服务器完全无意义。 - -**编译命令缺少隐式信息。** 系统头文件路径、默认的语言标准、编译器内置的宏定义等,在 CDB 中不会显式出现——它们由编译器工具链隐式提供。语言服务器需要通过工具链探测来补全这些信息。 - -**大型项目的编译命令高度冗余。** 一个拥有上万文件的项目中,绝大多数文件使用相同的编译器和语义选项(`-std=c++17`、`-Wall` 等),只有 include 路径和宏定义不同。不做去重会浪费大量内存和启动时间。在 clangd 的 issue 中,有用户报告 Bazel 生成的 `compile_commands.json` 超过 10GB,单条条目可达 250KB。 - -## 设计方案 - -### 整体流程 - -命令处理是一个多阶段管线: - -```text -compile_commands.json(原始命令) - ↓ 加载与解析 -参数分类(codegen-only / discarded / user-content / semantic) - ↓ 两级分离 -CanonicalCommand(共享的语义选项) + Patch(每文件的用户内容) - ↓ 工具链探测 -cc1 参数(前端可消费的完整编译命令) - ↓ 搜索路径提取 -SearchConfig(四段式头文件搜索路径) -``` - -### 参数分类 - -加载 CDB 时,每个编译选项被分为四类: - -**丢弃(Discarded)**:与构建产物相关的选项,语言服务器不需要。例如 `-o`(输出文件)、`-c`(编译模式)、`-M`(依赖扫描)、`-emit-pch`(PCH 构建)等。这些直接丢弃。 - -**仅代码生成(Codegen-only)**:只影响代码生成后端,不影响语义分析的选项。例如 `-fPIC`、`-fomit-frame-pointer`、`-funwind-tables`、调试信息选项组(`-g*`)等。这些不会改变 AST 或诊断结果,直接丢弃。注意 `-O` 和 `-fsanitize=` 等选项虽然看起来是代码生成相关的,但它们会定义宏(如 `__OPTIMIZE__`、`__has_feature(address_sanitizer)`),因此被保留。 - -**用户内容(User-content)**:每个文件可能不同的选项,但不影响工具链探测结果。主要是 include 路径(`-I`、`-isystem`、`-iquote`、`-idirafter`)和宏定义(`-D`、`-U`)。这些被提取为 patch,在工具链探测之后重新附加。 - -**语义选项(Semantic)**:剩余的所有选项。这些影响编译语义,且在工具链探测中起作用。例如 `-std=c++17`、`-Wall`、`-target`、`-march=` 等。 - -### 两级分离:Canonical 与 Patch - -分类完成后,编译命令被分为两部分: - -- **CanonicalCommand**:驱动程序名 + 所有语义选项。代表"如何调用编译器的语义方面"。 -- **Patch**:所有用户内容选项。代表"这个特定文件额外需要什么"。 - -**为什么要做这个分离?** - -核心原因是**工具链探测的缓存效率**。工具链探测需要实际调用编译器驱动(如 `g++ -v`),这是一个昂贵的操作(通常需要 100ms+)。探测结果只取决于驱动程序和语义选项——用户内容选项(`-I`、`-D`)不会影响驱动输出的 cc1 参数。因此,只要语义选项相同,无论文件有什么不同的 `-I` 路径,都可以共享同一个探测结果。 - -在实际项目中,上万个文件可能只有 5-50 个不同的 CanonicalCommand。这意味着工具链只需要探测 5-50 次,而不是上万次。 - -同时,这也使得编译命令的内存去重成为可能。相同的 CanonicalCommand 和相同的 Patch 组合会被合并为同一个 CompilationInfo 实例,通过指针共享。 - -### 路径绝对化 - -在加载阶段,所有用户内容中的 include 路径选项会被绝对化——如果路径是相对的,基于 CDB 条目的 `directory` 字段解析为绝对路径。这确保了后续使用时路径的一致性,无论当前工作目录是什么。 - -### 工具链探测 - -工具链探测将驱动级别的命令转换为 cc1 参数: - -**探测过程**:对于给定的 CanonicalCommand,调用对应的编译器驱动(GCC、Clang、MSVC 等),获取完整的 cc1 参数——包括目标三元组、系统头文件搜索路径、默认宏定义等所有隐式信息。 - -**缓存策略**:探测结果以 (driver, file extension, non-user-content flags) 为键缓存。文件扩展名参与缓存键是因为 `.c` 和 `.cpp` 文件可能触发不同的驱动规则。用户内容选项不参与缓存键,因为它们不影响驱动输出。 - -**负缓存**:对于探测失败的工具链(如不存在的 GCC 版本),也会缓存失败结果,避免对每个文件重复尝试。 - -**启动预热**:在服务器启动时,会对所有唯一的缓存键并行发起探测,预先填充缓存。这样当 LSP 请求开始到达时,所有工具链探测结果已经就绪。 - -**编译器适配**:不同编译器家族(GCC、Clang、MSVC/ClangCL 等)有不同的驱动调用方式和输出解析逻辑。部分编译器家族(NVCC、Intel、Zig)虽被识别,但目前回退到通用的 Clang 驱动路径,尚无专门的适配处理。 - -### 搜索路径提取 - -工具链探测完成后,cc1 参数中包含了完整的头文件搜索路径。SearchConfig 将这些路径提取并组织为四段式结构: - -1. **Quoted**(`-iquote` 目录):`#include "foo.h"` 的搜索路径 -2. **Angled**(`-I` 目录):`#include ` 的搜索路径 -3. **System**(`-isystem` 目录):系统头文件路径,诊断被抑制 -4. **After**(`-idirafter` 目录):在系统目录之后搜索 - -这个四段式模型与 Clang 内部的搜索逻辑基本一致(部分较少见的选项如 `-cxx-isystem`、`-iwithsysroot` 和 Framework 搜索路径尚未支持)。段内的路径会去重(从 Angled 段开始,与 Clang 的 `RemoveDuplicates` 算法一致),确保搜索行为与实际编译器行为匹配。 - -SearchConfig 是 include 路径解析、include 路径补全、依赖图构建等功能的基础输入。 - -## 设计决策与权衡 - -**为什么按选项 ID 分类,而不是字符串匹配?** 使用 Clang 自身的选项表(`OptTable`)进行分类,确保所有选项都被正确识别,包括带有复杂语法的选项(如 `-Wno-error=deprecated`、`-isystem=/usr/include`)。字符串匹配容易遗漏边界情况。 - -**为什么用户内容选项不参与工具链缓存键?** 这是两级分离设计的核心。`-I` 和 `-D` 不会改变驱动程序输出的系统路径或目标三元组等信息。分离它们使得缓存键的数量从"每文件一个"降低到"每配置一个",在大型项目中带来数量级的改善。 - -**为什么搜索路径去重要匹配 Clang 的算法?** 如果语言服务器的搜索路径顺序与实际编译器不一致,可能导致 include 解析到不同的文件,产生令人困惑的不一致行为。严格匹配确保行为一致性。 diff --git a/docs/zh/design/compilation-context.md b/docs/zh/design/compilation-context.md index 1db515a7b..f688501a5 100644 --- a/docs/zh/design/compilation-context.md +++ b/docs/zh/design/compilation-context.md @@ -102,9 +102,9 @@ const char* error_message(int code) { 头文件上下文在概念上很清楚(宿主源文件 + include 位置),但给定一个头文件上下文后,如何让 Clang 在正确的预处理器状态下编译这个头文件,是一个需要选择方案的工程问题。 -> 这里讨论的前缀合成只针对用户打开的头文件。磁盘上未被打开的头文件不需要单独处理——它们在各个源文件编译时会被正常包含和处理。索引系统在索引每个源文件时会顺带收集其中头文件的符号信息,再由 MergedIndex 将同一个头文件在不同源文件中产生的索引数据合并起来(合并机制见 [索引设计](index-design.md))。 +> 这里讨论的前缀合成只针对用户打开的头文件。磁盘上未被打开的头文件不需要单独处理——它们在各个源文件编译时会被正常包含和处理。索引系统在索引每个源文件时会顺带收集其中头文件的符号信息,再由 MergedIndex 将同一个头文件在不同源文件中产生的索引数据合并起来(合并机制见 [索引设计](symbol-index.md))。 -clice 采用的方案是**前缀合成 + `-include` 注入**:根据头文件上下文中的宿主源文件和 include 位置,沿 include 链提取目标头文件之前的所有内容,合成为前缀代码并写入磁盘上的前缀文件,然后通过 Clang 的 `-include` 标志将该文件注入到编译命令中。这个方案的核心优势是与 PCH 优化的天然配合——前缀文件的内容就是目标头文件的 preamble,可以编译为 PCH 缓存起来,用户后续编辑头文件主体时不必每次都重新处理前缀中大量的头文件包含。方案选择的详细论证见下方 FAQ 一节。PCH 的构建、缓存和失效机制见 [增量编译设计](incremental.md)。 +clice 采用的方案是**前缀合成 + `-include` 注入**:根据头文件上下文中的宿主源文件和 include 位置,沿 include 链提取目标头文件之前的所有内容,合成为前缀代码并写入磁盘上的前缀文件,然后通过 Clang 的 `-include` 标志将该文件注入到编译命令中。这个方案的核心优势是与 PCH 优化的天然配合——前缀文件的内容就是目标头文件的 preamble,可以编译为 PCH 缓存起来,用户后续编辑头文件主体时不必每次都重新处理前缀中大量的头文件包含。方案选择的详细论证见下方 FAQ 一节。PCH 的构建、缓存和失效机制见 [增量编译设计](incremental-parse.md)。 合成过程分为四个阶段。下面用一个具体的例子来说明。假设项目中有如下文件: @@ -168,7 +168,7 @@ Clang 会先处理 `-include` 指定的前缀文件,再编译 `math.h`。效 #endif ``` -如果索引只记录一个上下文的结果,用户做 go-to-definition 或 find-references 时就会丢失另一个上下文中的信息。MergedIndex 会将同一个文件在不同编译上下文下产生的索引数据合并存储,查询时返回所有上下文的并集——对 `config.h` 来说,find-references 能同时找到 `AsyncHandler` 和 `SyncHandler` 的引用。合并和去重的具体机制见 [索引设计](index-design.md)。 +如果索引只记录一个上下文的结果,用户做 go-to-definition 或 find-references 时就会丢失另一个上下文中的信息。MergedIndex 会将同一个文件在不同编译上下文下产生的索引数据合并存储,查询时返回所有上下文的并集——对 `config.h` 来说,find-references 能同时找到 `AsyncHandler` 和 `SyncHandler` 的引用。合并和去重的具体机制见 [索引设计](symbol-index.md)。 ## FAQ diff --git a/docs/zh/design/dependency-scanning.md b/docs/zh/design/dependency-scanning.md index 2d76e4804..91283d0f8 100644 --- a/docs/zh/design/dependency-scanning.md +++ b/docs/zh/design/dependency-scanning.md @@ -2,139 +2,171 @@ ## 背景 -C++ 语言服务器需要知道文件之间的包含关系。这个需求来自多个方面: +C++ 的 `#include` 指令在文件之间建立了依赖关系。语言服务器需要知道这些依赖关系,原因有三: -**头文件的编译上下文**:头文件不在编译数据库(CDB)中,语言服务器需要为它选择一个宿主源文件来提供编译命令。这要求知道"哪些源文件包含了这个头文件"——即反向的 include 关系。如果没有这个信息,用户打开头文件时将无法获得正确的诊断、补全和导航功能。 +**为头文件选择编译上下文。** 头文件不在编译数据库(CDB)中,无法直接编译。语言服务器需要找到一个包含该头文件的源文件,借用它的编译命令来编译头文件。这要求知道"哪些源文件(直接或间接)包含了这个头文件"——即 include 关系的反向查询。如果没有这个信息,用户打开头文件时将无法获得正确的诊断、补全和导航。关于编译上下文的完整讨论见 [编译上下文](compilation-context.md)。 -**模块依赖解析**:C++20 模块的编译需要知道模块之间的依赖关系。模块名到文件的映射以及模块间的 import 关系都需要通过扫描文件内容来发现。 +**解析 C++20 模块依赖。** 编译一个导入了模块的文件之前,必须先编译模块接口单元的 PCM。语言服务器需要知道"哪个文件提供了模块 `foo`"以及"这个文件导入了哪些模块"。这些信息只能通过扫描文件内容来发现,因为 CDB 不记录模块之间的依赖关系。关于模块编译的完整讨论见 [模块编译](module-graph.md)。 -**优先级调度**:后台索引需要决定索引文件的顺序。与用户当前打开的文件存在 include 关系的文件应该优先索引。通过 include 图可以追踪这种关联性。 +**优化后台索引顺序。** 与用户当前打开的文件存在 include 关系的文件应该优先索引,这样用户能更快获得跨文件导航功能。include 图提供了文件之间的关联度信息。 -问题在于,精确地获取 include 关系需要运行完整的 C 预处理器——展开所有宏、求值所有条件编译指令。对于大型 C++ 项目(上万个文件),在启动阶段对每个文件运行预处理器需要数分钟,这是不可接受的。 +问题在于:精确地获取 include 关系需要运行完整的 C 预处理器。C++ 的 `#include` 是文本替换——头文件的展开结果取决于包含它之前的预处理器状态。同一个头文件被不同源文件包含时,由于前面的 `#define` 和 `#include` 不同,可能走不同的条件编译分支,展开出不同的 include 列表: -clangd 的做法是在后台索引过程中逐步构建 include 图——每编译一个文件,就记录它的 include 关系。这意味着在后台索引完成之前(可能需要数十分钟),include 图是不完整的。用户在此期间打开头文件可能找不到正确的宿主源文件,导致错误的诊断信息。在 clangd 的社区中,用户反复报告启动后头文件长时间显示错误诊断、需要等待后台索引完成才能正常工作的问题。 - -clice 通过一个专门设计的快速依赖扫描器解决这个问题——用不精确但足够好的结果换取启动阶段数秒内扫完上万个文件的能力。 - -## 设计方案 +```cpp +// config.h +#ifdef USE_OPENSSL +#include // 源文件 A 看到 +#else +#include // 源文件 B 看到 +#endif +``` -### 核心思想 +这意味着每个 (源文件, 头文件) 的组合都必须独立预处理,无法共享结果。对于一个拥有上万个源文件、每个源文件平均包含数百个头文件的大型项目,工作量是 O(源文件数 × 平均头文件深度)——可能需要数百万次预处理操作,耗时数分钟,作为启动开销是不可接受的。 -依赖扫描的核心权衡是**用精确性换速度**。完整的预处理器能精确地确定哪些 include 生效、哪些被条件编译排除,但它有一个根本性的性能问题:同一个头文件在不同源文件的上下文中展开结果可能不同,因此每个 (源文件, 头文件) 的组合都必须独立预处理,无法共享结果。快速扫描器放弃预处理,使得扫描结果不依赖于包含上下文——每个文件只需扫描一次,结果在所有上下文间共享。代价是将所有条件分支中的 include 都视为有效,得到的是一个**包含所有可能 include 关系的超集**。这个超集虽然比实际的 include 关系更大,但对于依赖扫描的主要用途来说是安全的: +clangd 选择在后台索引过程中逐步构建 include 图——每编译一个翻译单元,就记录它的 include 关系。这个策略的问题在于,后台索引可能需要数十分钟才能覆盖整个项目。在此期间 include 图是不完整的,导致头文件的编译上下文选择只能依赖文件名匹配这类启发式方法。clangd [#123](https://github.com/clangd/clangd/issues/123) 讨论了这个问题:后台索引积累的 include 信息可以改善头文件的编译命令选择,但这些信息直到索引完成才可用。在此之前,用户打开头文件时可能看到错误的诊断,或者根本无法跳转定义。 -- 查找宿主源文件时,超集意味着可能找到更多候选宿主,但不会遗漏任何正确的宿主 -- 判断文件关联性时,超集意味着可能认为一些实际不相关的文件相关,但不会错过真正相关的文件 +clice 通过在启动阶段运行一个快速依赖扫描器来解决这个问题。扫描器放弃预处理,使用纯词法分析来提取 include 信息——不展开宏、不求值条件表达式。这样做的结果是得到一个包含所有可能 include 关系的**超集**,但换取了数量级的速度提升:每个文件只需扫描一次(结果不依赖包含上下文),上万个文件在数秒内完成。 -### 为什么能这么快 +## 设计 -快速扫描器的速度来自一个根本性的设计选择和几个关键优化: +### 超集扫描 -**根本性优势:跨上下文的结果共享** +依赖扫描的核心概念是**超集扫描**——放弃预处理以获得上下文无关的结果。 -这是快速扫描能比完整预处理快数量级的根本原因。在完整预处理模式下,同一个头文件被不同源文件包含时,由于前面的 `#define` 和 `#include` 不同,预处理展开的结果可能完全不同。这意味着每个 (源文件, 头文件) 组合都需要独立预处理——工作量是 O(源文件数 × 平均头文件深度)。对于一个包含上万个源文件、每个源文件平均包含数百个头文件的大型项目,这是数百万次预处理操作。 +完整的预处理器在展开 include 时,会求值所有宏和条件表达式,因此只输出当前上下文下实际生效的 include。结果是精确的,但依赖于包含上下文——同一个头文件被不同源文件包含时,输出可能不同。这迫使每个 (源文件, 头文件) 组合都独立处理,工作量是 O(源文件数 × 平均头文件深度)。 -快速扫描器放弃了预处理,因此扫描结果**不依赖于包含上下文**——同一个头文件无论被哪个源文件包含,扫描出的 `#include` 列表都是相同的。这使得每个头文件只需要扫描一次,结果在所有源文件间共享。工作量降为 O(唯一文件数)——通常只有几千到几万个文件,而非数百万次操作。 +快速扫描器跳过预处理,直接在词法层面识别 `#include` 指令和模块声明。它不知道哪些条件分支会生效,因此把所有分支中的 include 都记录下来。结果是一个超集——包含了所有可能被包含的头文件,可能多于任何单个编译上下文下的实际 include。但由于不依赖包含上下文,每个文件只需扫描一次,所有源文件共享同一个扫描结果,工作量降为 O(唯一文件数)。 -**轻量词法扫描** +这个超集对于依赖扫描的主要用途是安全的:查找宿主源文件时,超集意味着可能找到更多候选宿主,但不会遗漏正确的宿主;判断文件关联性时,超集意味着可能认为一些实际不相关的文件有关联,但不会错过真正相关的文件。 -扫描器使用 Clang 提供的依赖指令扫描器(scanSourceForDependencyDirectives),它只识别以 `#` 开头的预处理指令行,不做宏展开、不求值条件表达式、不展开 include。每个文件的扫描时间在微秒级。 +### 三种扫描模式 -扫描结果是原始的 include 名称(如 `"foo.h"` 或 ``),需要后续的路径解析阶段将它们映射到实际文件路径。 +系统提供三种扫描模式,面向不同的使用场景: -对于条件编译中的 include,扫描器通过跟踪 `#if`/`#ifdef`/`#ifndef` 的嵌套深度来标记:深度大于零时遇到的 include 被标记为"条件性"的。这个标记虽然不知道条件的具体求值结果,但提供了有用的元信息。 +**快速词法扫描** 是启动阶段的默认模式。使用 Clang 内置的依赖指令扫描器(`scanSourceForDependencyDirectives`),只识别以 `#` 开头的预处理指令行和模块声明。输出原始的 include 名称(如 `"foo.h"` 或 ``),需要后续的路径解析步骤。每个文件的扫描时间在微秒级。 -**分离的路径解析阶段** +**精确预处理扫描** 运行完整的 Clang 预处理器。输出的是已解析的文件路径(不是原始 include 名称),准确反映特定编译配置下的实际 include 关系。用于需要精确依赖信息的场景——例如 `CompileGraph` 在惰性解析模块依赖时,需要知道一个模块文件实际导入了哪些模块。由于需要完整的编译参数和预处理器实例,成本远高于快速扫描。 -include 名称到实际文件路径的解析是独立的阶段。解析器使用目录列表缓存——预先读取搜索路径中每个目录的文件列表,后续通过内存中的字符串集合检查文件是否存在,避免了大量的 stat 系统调用。 +**轻量模块声明扫描** 是一种介于快速和精确之间的回退模式。当快速扫描发现模块声明位于条件编译指令内部时——例如 `#ifdef _WIN32` 后面的 `export module platform;`——它无法确定哪个模块声明实际生效。此时触发轻量扫描:启动预处理器,但只词法分析到模块声明为止就停止,不处理整个文件。代价远低于完整的精确扫描,只应用于极少数文件(绝大多数模块声明在文件顶层,不在条件编译中)。 -尖括号 include(如 ``)的解析结果可以跨文件缓存——相同的编译配置下,同一个头文件名总是解析到同一个路径。双引号 include 依赖于包含者的目录,无法跨文件缓存,但每个文件的双引号 include 数量通常很少。 +### DependencyGraph -**波前式 BFS 发现** +`DependencyGraph` 是依赖关系的存储结构,在启动阶段由扫描器构建,之后在整个服务器生命周期内被多个模块查询。 -扫描不是一次性处理所有文件,而是按波次展开: +**正向 include 边** 记录"一个文件直接包含了哪些文件"。以 (文件, 编译配置) 为键——同一个文件在不同的编译配置下,由于搜索路径不同,可能解析出不同的 include 目标。每条 include 边附带一个标记位,标识该 include 是否位于条件编译指令内部。 -- **第 0 波**:扫描编译数据库中的所有源文件(并行 I/O + 词法扫描) -- **路径解析**:将发现的 include 名称解析为文件路径,识别出新发现的头文件 -- **第 1 波**:扫描新发现的头文件,发现它们的 include... -- 重复直到没有新文件 +**反向 include 映射** 记录"一个文件被哪些文件直接包含"。在所有正向边建立完成后批量构建。主要用于宿主源文件查找:从目标头文件出发,沿反向边向上 BFS,直到找到没有 includer 的根节点(即 CDB 中的源文件)。 -关键优化:当前波次在做路径解析时(串行),下一波次的文件已经可以开始预取和扫描(并行)。这种流水线式的重叠隐藏了大部分 I/O 延迟。 +**模块名映射** 记录模块名到模块接口单元文件的映射关系。只有接口单元(`export module` 声明的文件)被注册。一个模块名可能对应多个文件——当不同的编译配置中扫描到同名的模块接口单元时。 -### DependencyGraph 的存储结构 +### 搜索配置 -DependencyGraph 存储文件间的 include 关系,支持正向和反向查询: +将原始 include 名称(如 `` 或 `"foo.h"`)解析为实际文件路径,需要知道头文件的搜索目录。`SearchConfig` 封装了从编译参数中提取的搜索配置。 -**正向 include**:给定一个文件和编译配置,返回它直接包含的所有文件。每条边记录了是否为条件性 include(通过位标记区分)。一个文件在不同的编译配置下可能有不同的 include 集合(因为搜索路径不同),因此以 (文件, 配置) 对为键。 +搜索目录按照 Clang 的 `InitHeaderSearch::Realize` 布局分为四段: -**反向 include**:给定一个文件,返回所有直接包含它的文件。这是在所有正向 include 建立完成后批量构建的反向索引。用于"查找宿主源文件"——从目标头文件沿反向边向上 BFS,直到找到在 CDB 中有条目的源文件。 +``` +[Quoted (-iquote)] [Angled (-I)] [System (-isystem)] [After (-idirafter)] +``` -**模块映射**:模块名到模块接口单元文件路径的映射,用于 C++20 模块的依赖解析。只有接口单元(`export module`)被注册到映射中。一个模块名可能对应多个路径(在不同编译配置中发现相同模块)。 +这个分段决定了 include 解析的搜索顺序。`#include "foo.h"`(双引号)从 Quoted 段开始搜索——先在包含者所在目录查找,然后依次搜索 Quoted、Angled、System、After 段。`#include `(尖括号)跳过 Quoted 段,直接从 Angled 段开始。`#include_next` 从当前文件被找到的搜索目录的下一个位置开始。这些规则与 Clang 的头文件搜索行为一致。关于搜索配置的提取过程见 [命令解析](command-resolve.md)。 -### 条件 include 的处理 +### 条件 include 标记 -扫描器不求值条件表达式,但通过嵌套深度跟踪将 include 标记为"条件性"或"无条件": +快速扫描虽然不求值条件表达式,但通过跟踪 `#if`/`#ifdef`/`#ifndef` 的嵌套深度,为每条 include 标记是"无条件的"还是"条件性的": ```cpp -#include // 无条件(深度 0) +#include // 无条件——嵌套深度 0 #ifdef USE_BOOST -#include // 条件性(深度 1) -#ifdef BOOST_HAS_X -#include // 条件性(深度 2) -#endif +#include // 条件性——嵌套深度 1 #endif -#include // 无条件(深度 0) +#include // 无条件——嵌套深度 0 +``` + +当合并一个文件在所有编译配置下的 include 结果时,如果同一个目标在某个配置下是无条件的、在另一个配置下是条件性的,无条件优先。这反映了"至少在一个配置下,它一定会被包含"的语义。这个信息对宿主选择有指导意义:通过无条件 include 链到达目标头文件的源文件,是比通过条件 include 链到达的源文件更可靠的宿主候选。 + +## 实现 + +### 波前 BFS + +依赖扫描的核心流程是一个波前式 BFS。CDB 中只有源文件,头文件需要通过解析源文件的 include 来发现。每一波处理当前已知但尚未扫描的文件,发现新的头文件作为下一波的输入。 + ``` +Wave 0: CDB 中的源文件 ──→ 扫描 ──→ 解析 include ──→ 发现头文件 +Wave 1: 上一波发现的头文件 ──→ 扫描 ──→ 解析 include ──→ 发现更深层的头文件 +Wave 2: ... + ↓ +直到没有新文件被发现 +``` + +每一波分为两个阶段: + +**Phase 1(并行):读取 + 词法扫描。** 当前波次的所有文件被提交到线程池,并行执行文件读取和快速词法扫描。每个文件的输出是一个 `ScanResult`,包含原始 include 名称列表和模块声明信息。 + +**Phase 2(串行):include 路径解析 + 图构建。** 遍历 Phase 1 的扫描结果,将每条原始 include 名称通过搜索配置解析为实际文件路径,检查文件是否存在,将解析成功的 include 记录为 `DependencyGraph` 中的边。同时收集新发现的文件(之前未见过的路径)作为下一波的输入。 + +两个阶段之间有两处流水线优化,使相邻波次的工作互相重叠: + +1. Wave 0 的 Phase 1(文件扫描)和目录列表缓存的预填充在线程池上并行执行。目录列表缓存只在 Phase 2 的 include 解析中才被使用,因此两者可以完全重叠。 +2. Phase 2 在发现新文件时,立刻将其提交到线程池预取和扫描。当下一波的 Phase 1 启动时,这些预取任务通常已经完成或正在运行,减少了等待时间。 + +### 编译配置分组 + +一个项目中上万个源文件可能有不同的编译参数,但大部分文件共享相同的 include 搜索路径和编译选项。`CompilationDatabase` 将共享相同 `CompilationInfo`(目录、规范选项、用户内容选项均相同)的文件归为一个 `ConfigGroup`。依赖扫描为每个 `ConfigGroup` 提取一份 `SearchConfig`,同一组内的所有文件共用同一份搜索配置。 -当查询一个文件在所有配置下的 include 并集时,如果同一个头文件在某个配置下是无条件的、在另一个配置下是条件性的,无条件优先——这反映了"至少在一个配置下它一定会被包含"的语义。 +这种分组还服务于工具链探测的去重:不同 `ConfigGroup` 可能使用相同的编译器和语义选项,只有 include 路径不同。工具链的 `warm()` 方法在内部进一步按编译器标识去重,使得上万个源文件可能只需要一两次编译器子进程调用。 -### 模块声明的处理 +### Include 路径解析 -C++20 模块声明(`export module foo;`、`module foo:bar;`)通常出现在文件开头,快速扫描器可以直接识别。但如果模块声明出现在条件编译块中(`#ifdef` 内),快速扫描器无法确定哪个声明生效。 +将原始 include 名称解析为文件路径时,需要在搜索目录中查找文件是否存在。传统做法是对每个候选路径调用 `stat()` 系统调用。clice 使用目录列表缓存(`DirListingCache`)来替代——对搜索路径中的每个目录执行一次 `readdir()`,将结果缓存在内存中的字符串集合里,后续的文件存在性检查通过集合查找完成。 -此时扫描器设置 `need_preprocess` 标志,触发一个精确的回退——只对该文件的头部运行预处理器(到模块声明为止),求值条件表达式以确定实际的模块名。这个回退只影响极少数文件(绝大多数模块声明在文件顶层),不影响整体扫描速度。 +这种方式将 N 次 `stat()` 调用替换为 1 次 `readdir()` + N 次内存查找。在 Windows 上效果尤其显著,因为 Windows 的 `stat()` 调用开销约为 Linux 的 10 倍。 -### 缓存与增量更新 +对于多级路径的 include(如 ``),解析器使用快速拒绝优化:先检查搜索目录中是否存在第一级目录名(如 `llvm`),大多数搜索目录不包含这个子目录,因此可以跳过后续的完整路径构建和子目录解析。 -扫描结果在多个层面缓存: +尖括号 include(`<...>`)的解析结果可以跨文件缓存——相同编译配置下的相同头文件名总是解析到相同路径,包括解析失败的负缓存。双引号 include(`"..."`)依赖包含者所在目录,无法跨文件缓存。 -- **扫描结果缓存**:每个文件的扫描结果(include 列表、模块声明)按 path_id 缓存,单次扫描内复用避免重复读取。 -- **目录列表缓存**:搜索路径中每个目录的文件列表缓存在内存中,首次扫描时通过并发 readdir 任务填充。 -- **include 解析缓存**:尖括号 include 的解析结果按 (配置, 头文件名) 缓存,包括解析失败的负缓存。 +### 与其他模块的协作 -这些缓存目前在单次扫描调用中有效。缓存基础设施的设计支持跨扫描持久化(热启动),但尚未在服务器集成中启用。 +**编译上下文选择。** 当用户打开一个头文件时,`Compiler` 通过 `DependencyGraph` 查找宿主源文件。先调用 `find_host_sources` 沿反向 include 边 BFS 到达 CDB 中有编译命令的根源文件,然后调用 `find_include_chain` 沿正向边 BFS 找到从宿主到目标头文件的最短 include 链。这条链用于合成头文件的前缀代码——即还原它在宿主源文件中被包含时的预处理器状态。完整讨论见 [编译上下文](compilation-context.md)。 -### 与宿主源文件查找的协作 +**模块依赖解析。** `CompileGraph` 在惰性解析模块依赖时,通过 `DependencyGraph` 的模块名映射查找模块接口单元的文件路径。快速扫描在启动时就建立了模块名到文件的映射,`CompileGraph` 可以立即使用,无需等待任何编译完成。完整讨论见 [模块编译](module-graph.md)。 -依赖扫描的核心用途之一是为头文件查找宿主源文件。流程如下: +**Agent 接口。** `DependencyGraph` 的正向和反向查询通过 Agent 协议暴露给外部工具。Agent 可以查询一个文件的直接和传递 include/includer 关系,也可以请求影响分析——给定一个文件,返回所有直接和间接依赖它的文件。 -1. 从目标头文件出发,通过反向 include 索引向上遍历 -2. 找到所有传递包含该头文件的根文件(没有被其他文件包含的文件) -3. 筛选出在 CDB 中有编译命令的源文件作为候选宿主 -4. 选择第一个有有效 include 链的候选宿主作为默认宿主 +## FAQ -include 链的查找使用 BFS 保证找到的是最短路径。编译上下文系统随后利用这个 include 链合成头文件的编译环境。 +- **为什么选择超集而不是精确结果?** 精确的 include 关系需要运行完整的预处理器,对上万个文件来说需要数分钟。而依赖扫描的主要消费者——宿主源文件查找和文件关联性判断——都可以容忍多余的候选结果。超集在数秒内完成,对这些场景是安全的:多找到一些候选宿主比遗漏正确的宿主好得多。需要精确结果的场景(如模块依赖解析)使用精确预处理扫描单独处理。 -### 精确扫描与后台索引的补充 +- **为什么不完全依赖后台索引来构建 include 图?** 这是 clangd 的做法:后台索引编译每个翻译单元时顺带记录 include 关系。问题在于后台索引可能需要数十分钟才能覆盖整个项目,在此期间 include 图是不完整的,头文件体验很差。clice 用快速扫描保证启动后数秒内就有完整的 include 图,后台索引提供的精确信息作为补充。 -快速扫描提供的是启动阶段的近似结果。随着服务器运行,后台索引系统会逐步编译项目中的每个翻译单元,在编译过程中获取精确的 include 关系(经过完整预处理,求值了所有宏和条件编译)。这些精确的 include 信息被记录在索引数据(TUIndex/MergedIndex)中。需要注意的是,后台索引目前不会更新启动阶段构建的 DependencyGraph——快速扫描的 include 图在整个服务器生命周期内持续用于宿主源文件查找和文件依赖查询。 +- **快速扫描和精确扫描的关系是什么?** 两者不是替代关系,而是服务于不同场景。快速扫描在启动阶段对所有文件运行,构建全局的 `DependencyGraph`,为宿主查找和关联性判断提供即时可用的 include 图。精确扫描按需运行,用于需要准确依赖信息的特定操作(如模块编译时的依赖解析)。快速扫描是超集,精确扫描是子集——快速扫描可能包含实际不存在的 include 边,精确扫描的每条边都是真实的。 -此外,系统也提供精确扫描模式(scan_precise)——运行完整的 Clang 预处理器,用于需要准确依赖信息的特定场景,如模块编译时的惰性依赖解析(通过 CompileGraph 的 resolve_fn)。精确扫描的结果是已解析的文件路径(而非原始 include 名称),准确反映了特定编译配置下的实际 include 关系。 +- **为什么用目录列表缓存而不是 stat 调用?** 因为 include 路径解析需要在多个搜索目录中查找文件是否存在。如果对每个候选路径调用 `stat()`,一个文件的 include 解析可能触发数十次系统调用——乘以上万个文件就是数十万次。目录列表缓存将这些替换为每个目录一次 `readdir()` 加上内存中的集合查找。在 Windows 上(`stat()` 开销远高于 Linux),这个优化带来的改善更加显著。 -快速扫描和后台索引是互补的:快速扫描保证启动后几秒内就有可用的 include 图(用户立即获得头文件支持),后台索引在后续数分钟内逐步补充精确信息。两者不是替代关系——即使后台索引完成,快速扫描的结果仍然用于首次发现文件和建立初始依赖图。 +- **为什么用波前 BFS 而不是一次性扫描所有文件?** CDB 中只列出了源文件。头文件不在 CDB 中——它们的存在只能通过解析源文件的 include 指令来发现。因此无法一次性确定需要扫描哪些文件,只能通过 BFS 逐层发现。波前 BFS 是这种惰性发现的自然表达,同时通过流水线重叠最大化了并行度。 -## 设计决策与权衡 +## 已知局限 -**为什么用超集而不是精确结果?** 精确结果需要运行预处理器,对上万个文件来说需要数分钟。超集在数秒内完成,对于宿主查找和关联性判断来说是安全的——多找到一些候选宿主比漏掉正确的宿主要好得多。 +- **宏化的 include** 无法被快速扫描器识别。`#include MACRO_NAME` 需要宏展开才能得到实际的头文件名,而快速扫描不做宏展开。这种模式在实际项目中不常见,受影响的文件在精确编译时会被正确处理,但在 include 图中会缺失对应的边。 -**为什么不只靠后台索引构建精确的 include 图?** clangd 完全依赖后台索引来构建 include 关系,但后台索引可能需要几十分钟才能完成。在此期间头文件的体验是不完整的——用户打开头文件后看到错误的诊断,要等很久才会自动修复。clice 用快速扫描保证启动后几秒内就有可用的 include 图,后台索引随后逐步补充精确信息。两者结合既保证了即时可用,又保证了最终精确。 + ```cpp + #define PLATFORM_HEADER "platform_linux.h" + #include PLATFORM_HEADER // 快速扫描无法识别 + ``` -**为什么按波次展开而不是一次性扫描所有文件?** 不知道有哪些头文件需要扫描——它们不在 CDB 中,只能通过解析源文件的 include 来发现。波前式展开是惰性发现的自然模式,同时通过流水线重叠最大化了并行度。 +- **条件编译的精确性丢失。** 快速扫描记录了所有条件分支中的 include,导致 include 图比实际更大。对于宿主源文件查找这是安全的(不会遗漏),但可能导致关联性判断过于宽泛——将实际不相关的文件误认为相关。 -**为什么条件 include 的标记有用?** 虽然不知道条件的具体结果,但区分"一定会包含"和"可能包含"对宿主选择有指导意义。如果一个源文件通过无条件 include 链包含了目标头文件,它是比通过条件 include 链的源文件更可靠的宿主候选。 + ```cpp + #ifdef _WIN32 + #include // 仅 Windows 上生效 + #endif + #ifdef __linux__ + #include // 仅 Linux 上生效 + #endif + // 快速扫描会同时记录 windows.h 和 unistd.h + ``` -## 已知限制 +- **大小写敏感性。** 在大小写不敏感的文件系统(macOS HFS+/APFS、Windows NTFS)上,`#include` 中的大小写可能与磁盘上的文件名不一致。目录列表缓存的文件名匹配是大小写敏感的,可能导致解析失败。 -- **宏化的 include**:`#include MACRO_NAME` 形式的 include 无法被快速扫描器解析——它需要宏展开才能得到实际的头文件名。这种模式在实际项目中较少见,受影响的文件在精确编译时会被正确处理。 -- **条件编译的精确性**:所有条件分支中的 include 都被记录,导致 include 图比实际更大。这对宿主查找是安全的(不会漏掉),但可能导致关联性判断过于宽泛。 -- **大小写敏感性**:在大小写不敏感的文件系统(macOS、Windows)上,如果 include 的大小写与文件名不一致,可能导致解析失败。 -- **Framework 搜索路径**:macOS 的 Framework 目录(`-F`、`-iframework`)尚未支持,`` 形式的 Framework include 无法正确解析。 +- **macOS Framework 搜索路径。** 尚未实现 macOS 的 Framework 目录搜索(`-F`、`-iframework`)。`` 形式的 Framework include 应该在 `Foo.framework/Headers/` 下查找 `Bar.h`,但目前的解析器不识别这种特殊的目录结构。 diff --git a/docs/zh/design/incremental-parse.md b/docs/zh/design/incremental-parse.md new file mode 100644 index 000000000..9bbdf38d8 --- /dev/null +++ b/docs/zh/design/incremental-parse.md @@ -0,0 +1,199 @@ +# 增量编译 + +## 背景 + +C++ 的 `#include` 是文本替换——预处理器将被包含的头文件内容原样插入到源文件中。一个只有几十行用户代码的源文件,经过 `#include` 展开后,可能膨胀到数万行甚至更多。例如,仅 `#include ` 一条指令就会引入数千行标准库代码,如果再加上项目自身的头文件,展开后的代码量可以轻松超过十万行。 + +语言服务器需要在每次用户编辑后重新编译文件,以提供最新的诊断、补全和语义信息。如果每次都完整编译这十万行代码,延迟会达到数秒,显然不可接受。 + +但观察用户的实际编辑行为,会发现一个关键特征:文件头部的 `#include` 区域很少变化,真正频繁变化的只是底部的用户代码。在一次典型的编辑会话中,用户可能修改了上百次代码,但 `#include` 区域一次也没动过。 + +```cpp +#include +#include +#include +#include "project/config.h" +#include "project/logging.h" +// ── preamble ↑ 变化慢,编译结果可以缓存 ── +// ── 用户代码 ↓ 变化快,每次需要重新编译 ── + +void process(const std::vector& data) { + // ... +} +``` + +基于这个观察,C++ 语言服务器普遍采用 **preamble 分离**策略:将文件分为头部的 preamble(预处理指令区域)和剩余的用户代码。Preamble 编译为预编译头文件(PCH)并缓存,后续编译直接加载 PCH,只重新处理用户代码。这样每次编辑后的重编译只涉及几十行到几百行代码,延迟通常可以控制在一秒以内。 + +clangd 也采用了这一策略,但在失效检测和生命周期管理方面存在一些设计层面的不足: + +- **内存占用高**。clangd 将 preamble 的编译产物保留在进程内存中。对于大型项目,多个打开文件的 preamble AST 会消耗大量内存,长时间运行后内存占用持续增长(clangd [#251](https://github.com/clangd/clangd/issues/251)、[#115](https://github.com/clangd/clangd/issues/115))。 + +- **崩溃后 PCH 文件泄漏**。clangd 使用临时文件存储磁盘上的 PCH。崩溃时 RAII 清理无法执行,临时文件残留在 `/tmp` 中。在多用户服务器上,累积的泄漏文件会耗尽 `/tmp` 空间(clangd [#209](https://github.com/clangd/clangd/issues/209)、[#255](https://github.com/clangd/clangd/issues/255))。 + +- **失效检测不够精确**。当头文件被 touch 但内容未变时(常见于构建工具的依赖扫描、版本控制的分支切换),仅靠修改时间(mtime)判断会导致不必要的 PCH 重建。在大型项目中,这种误报造成的重复构建会明显影响编辑体验。 + +- **重启后需要冷启动**。PCH 不跨会话持久化,服务器重启后需要为所有打开的文件重新构建 PCH。 + +clice 针对这些问题重新设计了增量编译机制:磁盘持久化的内容寻址 PCH 存储、两层失效检测、拉取式编译模型和 preamble 完整性检测。 + +## 设计 + +增量编译围绕四个核心概念组织:preamble 分离定义了"缓存什么",两层失效检测定义了"何时重建",拉取式编译定义了"何时触发",内容寻址存储定义了"怎么存"。 + +### Preamble 分离 + +Preamble 是源文件头部由预处理指令(`#include`、`#define`、`#pragma` 等)和模块声明(`module;`)组成的区域。Preamble 的结束位置由一个字节偏移量(bound)标识——第一行非预处理内容之前的字节位置。 + +Preamble 编译为 PCH 文件后缓存在磁盘上。后续编译加载 PCH 后,只需要处理 bound 之后的用户代码。如果 preamble 为空(bound 为零),则不需要 PCH,直接跳过整个 PCH 流程。 + +`PCHState` 是 PCH 在缓存中的条目,包含: + +- PCH 文件在磁盘上的路径 +- preamble 内容的哈希值 +- preamble 的字节边界(bound) +- 依赖快照(`DepsSnapshot`,见下文) +- PCH 中提取的 DocumentLink 信息(`#include` 指令的位置和目标,供编辑器显示可点击链接) + +### 两层失效检测 + +PCH 缓存了 preamble 中所有被 `#include` 的头文件的预处理结果。当任何一个依赖头文件的内容发生变化时,PCH 就已过时,需要重建。问题在于如何精确判断"内容是否真的变了"。 + +最直接的方案是检查文件修改时间(mtime):如果所有依赖文件的 mtime 都不晚于 PCH 的构建时间,说明没有文件被修改过。这种检查只需 `stat` 系统调用,非常快。但 mtime 检查会产生误报:构建工具的依赖扫描、版本控制的分支切换、编辑器的自动保存等操作都会更新 mtime 而不改变文件内容。 + +另一个方案是直接比较文件内容的哈希值:每次检查时重新计算所有依赖文件的哈希,与构建时记录的哈希比较。这种方案完全精确,但需要读取和计算所有依赖文件的内容。一个典型的 C++ 文件可能依赖数百个头文件,每次检查都全量哈希的 I/O 开销不可忽视。 + +clice 将两者结合为两层检测策略: + +- **第一层(mtime 快速筛选)**:遍历所有依赖文件,比较每个文件的 mtime 与 PCH 的构建时间戳。如果所有 mtime 都不晚于构建时间戳,PCH 有效,直接复用。 +- **第二层(内容哈希精确验证)**:对第一层标记为"可能已修改"的文件(mtime 晚于构建时间戳),重新计算其 xxh3 内容哈希,与构建时记录的哈希比较。只有哈希不一致时才触发重建。 + +第一层过滤掉了绝大多数未变化的文件(常态路径),第二层消除了 mtime 误报(构建工具 touch、VCS checkout 等场景)。两层结合的效果是:只有依赖文件的内容真正发生变化时才重建 PCH。 + +`DepsSnapshot` 是两层检测的基础数据结构,在 PCH 构建完成时捕获。它记录所有依赖文件的路径标识、内容哈希和构建时间戳。 + +### 拉取式编译 + +clice 采用拉取式(pull-based)编译模型:编译不在文件变更时立即触发,而是在功能请求(hover、补全、语义高亮等)需要最新 AST 时按需触发。 + +当用户编辑文件时(`didChange`),主进程只更新内存中的文件内容并标记 AST 为脏(`ast_dirty`),不启动任何编译。当功能请求到达时,编译服务检查 AST 是否脏或是否因外部文件变化而过时,如果需要重编译,则先确保 PCH 和模块依赖就绪,再将编译任务发送到工作进程。 + +这种模型的好处是避免了用户快速连续输入时的无效编译。用户每秒可能触发十几次 `didChange`,但只有当鼠标悬停、请求补全等实际需要编译结果时,才执行一次编译。 + +> 注意"外部文件变化"和"用户编辑"是两条独立的脏标记路径。用户编辑通过 `didChange` 标记 `ast_dirty`;外部文件变化(如依赖的头文件被修改)通过两层失效检测在编译前动态发现。 + +### 内容寻址 PCH 存储 + +PCH 文件在磁盘上以 preamble 内容的哈希值命名(如 `a3f7e8c1d2b4f6e9.pch`),实现内容寻址。这带来两个好处: + +- **磁盘共享**:具有相同 preamble 内容的不同文件自然共享同一个 PCH 磁盘文件,无需额外的去重逻辑。 +- **跨会话持久化**:PCH 缓存的元数据(路径、哈希、边界、依赖快照)序列化到磁盘上的 `cache.json` 文件。服务器重启时加载这些元数据,通过两层失效检测验证 PCH 是否仍然有效,避免冷启动时重建所有 PCH。 + +当 preamble 内容变化时,新的 PCH 使用不同的哈希命名,旧文件成为孤立文件。清理机制定期回收超过一定期限未使用的孤立 PCH 文件。 + +## 实现 + +### Preamble 边界计算 + +Preamble 的边界通过词法扫描确定:使用项目的 `Lexer` 从文件开头逐行扫描,识别以 `#` 开头的预处理指令和 `module;` 全局模块片段声明。遇到第一行非预处理内容时停止,返回此时的字节偏移作为边界。 + +这种基于词法的检测不需要启动完整的预处理器,速度非常快。 + +### Preamble 完整性检查 + +在触发 PCH 重建之前,需要检查 preamble 是否在语法上完整。两种典型的不完整状态: + +```cpp +#include "lib // 引号未闭合,用户正在输入文件名 +import std.core // 缺少分号,用户正在输入模块声明 +``` + +如果在不完整的 preamble 上构建 PCH,会产生一个包含错误预处理状态的 PCH 文件。后续加载这个错误 PCH 的编译会看到大量虚假错误。因此,当检测到 preamble 不完整时,延迟重建,继续使用旧的 PCH(如果存在)。 + +### PCH 构建流水线 + +当编译请求需要 PCH 时,按以下流程处理: + +``` +计算 preamble 边界和哈希 + │ + ▼ + ┌─ 缓存命中?──── 是 → 复用缓存的 PCH + │ │ + │ 否 + │ │ + │ ▼ + │ preamble 完整?── 否 → 延迟重建,沿用旧 PCH + │ │ + │ 是 + │ │ + │ ▼ + │ 有其他协程正在构建?── 是 → 等待构建完成,使用结果 + │ │ + │ 否 + │ │ + │ ▼ + │ 发送到无状态工作进程构建 PCH + │ │ + │ ▼ + └─ 更新缓存,捕获依赖快照 +``` + +缓存命中的条件有两个:preamble 哈希与缓存一致(preamble 内容没变),且两层失效检测通过(依赖文件内容没变)。两个条件必须同时满足。 + +PCH 构建由无状态工作进程执行(详见[多进程架构](multi-process.md))。工作进程使用 Clang 的 Preamble 编译模式,只处理 bound 之前的 preamble 部分。构建完成后返回 PCH 文件路径和依赖文件列表。 + +### 并发构建序列化 + +多个功能请求可能同时触发同一文件的 PCH 构建。`PCHState` 中包含一个共享事件(`building`):第一个发起构建的协程设置这个事件,后续协程发现事件存在时等待其完成,然后使用构建结果。这确保同一文件的 PCH 只构建一次。 + +### 依赖快照的时序保证 + +`DepsSnapshot` 的构建时间戳(`build_at`)在计算文件哈希**之前**获取。这个顺序确保了不存在遗漏修改的时间窗口: + +如果一个文件在哈希计算过程中被修改,它的 mtime 会晚于 `build_at`。下次两层检测时,第一层会将这个文件标记为"可能已修改",第二层会重新计算哈希并发现变化。 + +如果反过来先计算哈希再获取时间戳,就可能出现这样的窗口:文件在哈希计算和时间戳获取之间被修改,但 mtime 不晚于 `build_at`,导致修改被遗漏。 + +### 整体编译流程 + +当功能请求到达时,编译流水线按以下顺序执行: + +1. 检查 AST 是否已缓存且未过时——是则直接复用 +2. 如果使用了 C++20 模块,先确保模块依赖就绪(详见[模块编译](module-graph.md)) +3. 确保 PCH 就绪 +4. 将编译任务发送到有状态工作进程,附带 PCH 路径和模块文件路径 +5. 工作进程加载 PCH 后只编译 preamble 之后的用户代码 + +AST 的依赖文件(preamble 之后的 `#include`)同样通过 `DepsSnapshot` 跟踪,使用相同的两层失效检测。即使用户没有编辑当前文件,如果其依赖的头文件在磁盘上被修改,下次功能请求时也会触发重编译。 + +### 与编译上下文的交互 + +对于非自包含头文件,编译上下文系统会合成一个前缀文件,通过 `-include` 注入到编译参数中(详见[编译上下文](compilation-context.md))。这个注入的前缀文件会被 Clang 在 preamble 编译阶段处理,因此自然地被 PCH 缓存覆盖。PCH 的构建流水线对源文件和头文件的处理是统一的。 + +### 缓存持久化 + +PCH 和 PCM 的缓存元数据通过 `cache.json` 文件持久化到磁盘。每次成功构建后更新,服务器启动时加载。写入采用先写临时文件再原子重命名的模式,避免写入过程中崩溃导致文件损坏。 + +启动时加载缓存后,所有 PCH 条目通过两层失效检测验证有效性。过时的条目会在下次编译时自动重建,无需特殊的缓存一致性恢复逻辑。 + +## FAQ + +- **为什么用两层检测,而不是只用内容哈希?** 内容哈希虽然精确,但需要读取所有依赖文件的内容。一个典型的 C++ 文件可能依赖数百个头文件,每次检查都全量哈希的 I/O 开销不可忽视。mtime 快速筛选将需要哈希的文件数量缩减到"自上次构建以来被 touch 过的文件",在常态路径下通常为零。 + +- **为什么每次都完全重建 PCH?能否增量更新?** Clang 支持链式 PCH(chained PCH):将 preamble 中的每条 `#include` 构建为独立的 PCH 链节,每个链节依赖前一个链节的编译产物。当用户在 preamble 末尾新增一条 `#include` 时,只需要在已有链的末尾追加一个链节,而不是重建整个 preamble。基准测试表明(PR [#405](https://github.com/ykiko/clice/pull/405)),对于 70 个 C++ 标准库头文件的 preamble,增量追加一条 `#include` 只需约 36ms,而整体重建需要约 1230ms(35 倍加速)。链式 PCH 的 AST 加载延迟几乎不受影响(+2% ~ +6%)。clice 计划引入链式 PCH 来优化增量重建性能,目前尚在实验阶段。 + +- **为什么采用拉取式编译而不是推送式?** 关键原因在于 clice 将所有文件的 PCH 持久化到磁盘上,缓存的文件数量远多于 clangd 的内存模型(clangd 通过 LRU 策略只保留少量活跃文件的 preamble)。当一个头文件被修改时,可能有大量文件的 PCH 受到影响。如果采用推送式编译,就需要在头文件修改时立即重建所有受影响的 PCH,这个数量级是不可接受的。拉取式编译将重建延迟到功能请求到达时,只重建用户当前需要的那个文件的 PCH。由于 PCH 加载本身很快(见下一条),这种按需重建引入的延迟很小。 + +- **磁盘 PCH 比内存 PCH 慢吗?** 实际影响很小。Clang 的 PCH 加载使用 mmap 将文件映射到内存,避免了完整的读取拷贝。更重要的是,Clang 对 PCH 中的 AST 节点采用惰性反序列化——只有实际被引用的节点才会被反序列化,大部分 PCH 内容在编译过程中不会被访问。因此 PCH 的加载性能主要取决于二进制文件的映射方式,磁盘文件和内存缓冲区之间没有本质差异。而磁盘 PCH 带来的好处——跨重启持久化、不占用进程常驻内存、内容寻址共享——使得这个 trade-off 是值得的。 + +## 已知局限 + +- **每文件独立缓存**。PCH 缓存以文件的路径标识为键,即使两个文件具有完全相同的 preamble 内容,它们在缓存中也是独立的条目——各自执行失效检测和构建。虽然磁盘上的 PCH 文件通过内容寻址命名实现了共享,但缓存的元数据(依赖快照、构建状态等)没有共享。改进方向是以 preamble 内容哈希加编译标志为键,实现跨文件的缓存元数据共享。 + +- **完全重建**。任何一个依赖文件的内容变化都触发 PCH 的完全重建,无法做到只重建受影响的部分。改进方向是引入链式 PCH(见 FAQ),将重建范围限制在变化点之后的链节。 + +- **编译标志不参与缓存键**。PCH 的磁盘文件名和缓存查找都不考虑编译标志。当两个文件具有相同的 preamble 文本但不同的编译标志(如 `-D`)时,可能错误地共享 PCH。改进方向是将影响预处理的编译标志纳入缓存键。 + +- **Preamble 完整性检查不完整**。当前的完整性检查只覆盖了 `#include`/`import` 指令中的未闭合引号和缺失分号。其他类型的不完整编辑(如正在输入 `#define` 的值)不会被检测到。如果这类不完整的 preamble 被构建为 PCH,对后续编译的影响尚未充分测试。需要进一步研究 Clang 在处理不完整预处理指令时的行为,以确定是否需要扩展完整性检查的范围。 + +- **头文件保存后不主动推送诊断更新**。当前的纯拉取式模型下,用户保存一个头文件后,依赖该头文件的已打开源文件不会立即更新诊断——需要用户在源文件上触发操作(如 hover、编辑)才会通过两层失效检测发现变化并重编译。改进方向是在 `didSave` 时检查已打开的 session 是否受影响,为它们主动触发一次编译(混合推/拉模型)。 diff --git a/docs/zh/design/incremental.md b/docs/zh/design/incremental.md deleted file mode 100644 index 8957c5c34..000000000 --- a/docs/zh/design/incremental.md +++ /dev/null @@ -1,139 +0,0 @@ -# 增量编译 - -## 背景 - -语言服务器每次用户编辑代码都需要重新编译文件以提供最新的诊断、补全和语义信息。C++ 文件的完整编译可能需要数秒——一个典型的源文件通过 `#include` 展开后可能包含数万甚至数十万行代码。如果每次编辑都完整重编译,用户体验是不可接受的。 - -clice 通过 **preamble 分离**实现增量编译:将源文件分为两部分——头部的 `#include` 区域(preamble)和剩余的用户代码。preamble 编译为预编译头文件(PCH)缓存在磁盘上,后续编译直接加载 PCH,只重新处理用户代码部分。例如,一个包含 `#include ` 的文件,展开后约有 2 万行代码;构建 PCH 后,后续重编译的代码量只剩用户自己写的几十行。 - -这一机制的核心挑战在于 PCH 的**失效检测**和**生命周期管理**: - -**失效检测的精确性**:PCH 缓存了所有被 `#include` 的头文件的预处理结果。当任何一个头文件被修改时,PCH 就已过时。但简单地检查文件修改时间(mtime)会导致大量误报——许多构建工具会 touch 文件但不改变内容(如 `cmake --build` 的依赖扫描、版本控制系统的分支切换),触发不必要的 PCH 重建。 - -**重建开销控制**:PCH 的构建本身需要完整的预处理和序列化,通常需要数百毫秒到数秒。如果每次保存都触发 PCH 重建,用户会感受到明显的卡顿。需要精确判断"PCH 是否真的过时了",只在必要时重建。 - -**并发协调**:多个文件可能共享相同的 preamble。当 PCH 需要重建时,多个编译请求可能同时触发重建。需要确保只执行一次构建,其他请求等待结果。 - -**不完整编辑**:用户正在输入 `#include "` 时,preamble 处于不完整状态。此时触发 PCH 重建会产生一个错误的 PCH,导致后续编译出现大量虚假错误。 - -在 clangd 中,增量编译相关的问题长期困扰用户:头文件修改后诊断信息不更新,需要重新打开文件或重启服务器;PCH 重建过于频繁,在大型项目中导致编辑体验卡顿;多文件同时编辑时偶尔出现 PCH 损坏,导致整个项目无法正确编译。 - -## 设计方案 - -### 核心思想 - -增量编译的基本策略是将源文件分为**变化慢的部分**(preamble)和**变化快的部分**(用户代码),分别处理。preamble 编译为 PCH 缓存在磁盘上,只在依赖的头文件真正发生内容变化时才重建;用户代码部分在每次编辑后加载 PCH 重新编译,成本很低。 - -PCH 管理的核心挑战是**精确的失效检测**——在"过于激进地重建"和"使用过时的缓存"之间找到平衡。clice 通过两层失效检测(mtime 快速筛选 + 内容哈希精确验证)实现精确判断,结合 preamble 完整性检测和并发构建序列化,确保 PCH 始终处于正确且高效的状态。 - -### Preamble 边界检测 - -PCH 的核心是 preamble——源文件头部由预处理指令组成的区域。语言服务器需要精确确定 preamble 的结束位置(字节偏移),以便只将这部分内容编译为 PCH。 - -preamble 边界检测基于词法扫描:从文件开头逐行扫描,识别以 `#` 开头的预处理指令(`#include`、`#define`、`#pragma` 等)和全局模块片段声明(`module;`)。遇到第一行非预处理内容时停止。返回的字节偏移就是 preamble 的边界。 - -如果边界为零(文件没有任何预处理指令),则不需要 PCH——直接跳过 PCH 构建流程,并清除该文件已有的 PCH 缓存条目。 - -### 两层失效检测 - -PCH 的失效检测需要回答:"自上次构建以来,是否有任何被 PCH 包含的头文件内容发生了变化?" - -clice 采用两层检测策略,在精确性和性能之间取得平衡: - -**第一层:修改时间(mtime)快速筛选** - -遍历 PCH 的所有依赖文件,获取每个文件的 mtime。如果所有文件的 mtime 都不晚于 PCH 的构建时间戳(build_at),说明没有文件被修改过,PCH 仍然有效——直接复用。这一层只需要 stat 系统调用,成本很低。 - -**第二层:内容哈希(content hash)精确验证** - -对于 mtime 晚于 build_at 的文件(被第一层标记为"可能已修改"),重新计算其内容的 xxh3 哈希值,与构建时记录的哈希值比较。如果哈希一致,说明文件内容实际上没有变化(只是被 touch 了),PCH 仍然有效。只有哈希不一致时才真正触发重建。 - -这种两层策略的核心价值在于**消除误报**:构建工具的 touch 操作、版本控制的 checkout 操作等都会更新 mtime 但不改变内容。仅靠 mtime 会导致大量不必要的 PCH 重建。第二层的内容哈希以较小的额外开销(只对"嫌疑"文件计算哈希)过滤掉这些误报。 - -### DepsSnapshot:依赖快照 - -两层检测的基础是 DepsSnapshot——在 PCH 构建完成时捕获的依赖状态快照。它记录: - -- **path_ids**:所有依赖文件的路径标识 -- **hashes**:每个依赖文件在构建时的内容哈希值(xxh3_64bits) -- **build_at**:快照捕获的时间戳 - -**时序正确性**:build_at 在计算文件哈希**之前**获取。这确保了:如果一个文件在哈希计算过程中被修改(mtime 晚于 build_at),下次检测时它会被第一层标记为"可能已修改",进入第二层验证。不会出现修改被遗漏的时间窗口(TOCTOU)。 - -### PCH 构建流程 - -当编译请求需要 PCH 时,通过 ensure_pch() 触发以下流程: - -**1. 计算 preamble 边界和哈希** - -提取文件内容的 preamble 部分,计算其 xxh3 哈希值。这个哈希值用于 PCH 文件的磁盘命名,同时作为缓存有效性的验证依据。 - -**2. 查找缓存** - -在 Workspace 的 pch_cache 中查找该文件是否已有缓存条目。如果存在且: - -- preamble 哈希一致(preamble 内容没变) -- 两层失效检测通过(依赖文件内容没变) - -则直接复用缓存的 PCH,无需重建。 - -**3. 检测 preamble 完整性** - -如果缓存未命中或已失效,在触发重建之前检查 preamble 是否完整——是否存在未闭合的引号(如 `#include "lib`)或未终结的模块声明(如 `import foo` 缺少分号)。如果不完整,说明用户正在编辑 preamble 区域,此时延迟重建,继续使用旧的 PCH(如果存在)。 - -**4. 序列化并发构建** - -检查是否有其他协程正在为该文件构建 PCH。如果有,等待其完成事件,然后使用构建结果。这通过 PCHState 中的共享事件(kota::event)实现——第一个发起构建的协程设置事件,后续协程等待同一个事件。 - -**5. 分派到无状态工作进程** - -将 PCH 构建作为高优先级任务发送到无状态工作进程。工作进程使用 Clang 的 Preamble 编译模式(CompilationKind::Preamble),通过文件重映射将完整文件内容和 preamble 边界传递给 Clang,Clang 只处理边界之前的 preamble 部分。构建完成后返回 PCH 文件路径和依赖文件列表。 - -**6. 更新缓存** - -将构建结果存入 pch_cache:PCH 文件路径、preamble 哈希、preamble 边界、依赖快照。同时捕获 PCH 中的 DocumentLink 信息(`#include` 指令的位置和目标),供编辑器显示可点击的头文件链接。 - -### PCH 文件的磁盘命名 - -PCH 文件在磁盘上以 preamble 内容的哈希值命名,实现内容寻址。这使得具有相同 preamble 内容的不同文件**在磁盘上共享同一个 PCH 文件**。preamble 内容变化时,新的 PCH 文件使用不同的哈希值命名,旧文件成为孤立文件。清理机制定期回收超过一定期限未使用的孤立 PCH 文件。 - -### PCH 缓存的存储模型 - -PCH 缓存(pch_cache)当前以源文件的 path_id 为键。这意味着:即使两个文件具有完全相同的 preamble,它们在缓存中也是独立的条目——各自执行失效检测和重建。虽然磁盘上的 PCH 文件因内容寻址命名而共享,但缓存的元数据(依赖快照、构建时间戳等)没有共享。 - -PCH 缓存的元数据(路径、哈希、边界、依赖快照)持久化到磁盘上的 cache.json 文件中,服务器重启时可以恢复,避免冷启动时重建所有 PCH。 - -### 编译触发与整体流程 - -clice 采用**拉取式**编译模型——编译不是在文件变更时立即触发,而是在功能请求(hover、补全、语义高亮等)需要最新 AST 时按需触发。这避免了用户快速连续输入时的无效编译。 - -当功能请求需要最新的 AST 时,编译流水线依次执行: - -1. 检查 AST 是否已缓存且不脏——如果是,直接复用 -2. 检查 AST 的依赖是否过时(复用同样的两层失效检测) -3. 如果需要重新编译,先确保 C++20 模块依赖就绪(通过 CompileGraph,详见模块编译文档) -4. 然后确保 PCH 就绪(通过 ensure_pch) -5. 最后将编译任务发送到有状态工作进程,附带 PCH 路径 - -有状态工作进程加载 PCH 后只需编译 preamble 之后的用户代码,这就是增量编译的核心收益——将编译时间从数秒降低到通常不到一秒。 - -### PCH 与头文件上下文 - -对于非自包含头文件,编译上下文系统会合成一个前缀文件(详见编译上下文文档),通过 `-include` 注入来模拟头文件在宿主源文件中的包含环境。这个注入的前缀文件会出现在头文件编译的 preamble 中,因此也会被 PCH 缓存覆盖。 - -ensure_pch() 对头文件和源文件的处理逻辑是统一的:无论 preamble 包含的是文件自身的 `#include` 区域还是注入的前缀内容,都经过相同的哈希、缓存、失效检测和构建流程。 - -## 设计决策与权衡 - -**为什么用两层检测而不是只用内容哈希?** 内容哈希虽然精确,但需要读取和哈希每个依赖文件的全部内容。一个典型的 C++ 文件可能依赖数百个头文件,全部哈希的 I/O 成本不可忽视。mtime 快速筛选将需要哈希的文件数量缩减到"自上次构建以来被 touch 过的文件",通常远小于总依赖数。 - -**为什么不用增量 PCH 修补?** 理想情况下,当只有一个头文件发生了小的变更时,应该可以增量修补 PCH 而不是完全重建。但 Clang 的 PCH 格式不支持增量修补——它是一个整体序列化的 AST 和预处理器状态快照。增量修补需要深入修改 Clang 的序列化层,收益与复杂度不成比例。 - -**为什么 PCH 哈希不包含编译标志?** 当前 PCH 的哈希只基于 preamble 文本内容,不包含编译标志(`-D`、`-std` 等)。这是一个有意识的折中——它允许使用相同 preamble 文本但不同编译标志的文件共享 PCH 磁盘文件。在实践中,同一个 preamble 在不同编译标志下的预处理结果通常是相同的(差异主要来自宏定义,而宏定义差异通常会反映在 preamble 文本的 `#define` 中)。但这确实存在不正确的边界情况——例如通过命令行 `-D` 定义影响头文件行为的宏时,不同的编译标志应该使用不同的 PCH。 - -**为什么 preamble 完整性检测很重要?** 在不完整的 preamble 上构建 PCH 会产生一个包含错误预处理状态的 PCH 文件。后续使用这个错误 PCH 的编译会看到大量虚假错误。延迟重建直到 preamble 完整,避免了这种连锁错误。 - -## 已知限制 - -- **每文件独立缓存**:pch_cache 以 path_id 为键,无法在具有相同 preamble 的不同文件之间共享缓存元数据(依赖快照、构建状态等)。改进方向是以 preamble 内容哈希加编译标志为键,实现真正的跨文件缓存共享,同时解决编译标志不参与缓存键的问题,并引入 LRU 淘汰策略。 -- **完全重建**:任何依赖文件的内容变化都触发 PCH 的完全重建。无法做到只重建受影响的部分。这是 Clang PCH 格式的固有限制。 diff --git a/docs/zh/design/index-design.md b/docs/zh/design/index-design.md deleted file mode 100644 index b756e7eb6..000000000 --- a/docs/zh/design/index-design.md +++ /dev/null @@ -1,157 +0,0 @@ -# 索引设计 - -## 背景 - -语言服务器的跨文件功能——跳转到定义、查找引用、调用层次、符号搜索等——都依赖于符号索引。索引系统需要解决以下核心问题: - -**大规模项目的性能**:C++ 项目可能有上万个文件,每个文件编译后产生大量符号。索引系统必须在合理的时间和内存约束下完成索引构建、持久化和查询。 - -**增量更新**:用户修改文件后,相关的索引应该能够增量更新,而不是重新索引整个项目。 - -**编译上下文感知**:同一个头文件在不同的编译上下文下可能产生不同的符号。clangd 等现有方案只存储最后一次编译的结果,导致用户在切换上下文后看到过时的索引数据。 - -**打开文件的实时反馈**:用户正在编辑的文件应该立即反映在查询结果中,而不是等待后台索引完成。 - -在 clangd 的 issue 中,以下问题反复出现: - -- 跳转到定义跳到了错误的位置(当多个编译单元定义了同名符号时) -- 后台索引速度慢,大型项目需要数十分钟才能完成首次索引 -- 符号搜索结果不完整或包含过时条目 -- 头文件中的模板代码在不同实例化上下文中有不同的引用关系 - -## 设计方案 - -### 三级索引层次 - -clice 采用三级索引结构,每级服务于不同的目的: - -```text -TUIndex(编译产物) - ↓ 合并 -ProjectIndex(全局符号目录)+ MergedIndex(按文件分片的关系数据) - ↑ 叠加 -OpenFileIndex(打开文件的实时覆盖) -``` - -### TUIndex:单次编译产出 - -TUIndex 是一次编译产生的原始索引数据。当一个翻译单元被编译时,SemanticVisitor 遍历 AST,记录所有符号的出现和关系,产出一个 TUIndex。 - -TUIndex 包含: - -- **按文件组织的索引数据**:编译涉及的每个文件(主文件 + 所有包含的头文件)各有一份 FileIndex,记录该文件中的符号出现(Occurrence)和符号关系(Relation) -- **符号表(SymbolTable)**:将符号哈希映射到符号名称和种类 -- **include 图**:记录文件间的 include 关系 - -TUIndex 是临时数据——它在合并到 ProjectIndex 和 MergedIndex 后就不再需要。在后台索引场景下,TUIndex 在工作进程中生成并序列化,传输到主进程后合并。 - -### ProjectIndex:全局符号目录 - -ProjectIndex 是全局的符号目录,汇聚所有已索引文件的符号信息。它的作用类似于搜索引擎的倒排索引——给定一个符号哈希,可以快速查到它的名称、种类、以及它出现在哪些文件中。 - -ProjectIndex 存储: - -- **符号表**:符号哈希 → 符号名称、符号种类、引用文件位图(Bitmap) -- **路径池**:文件路径的内部化映射 - -**关键设计**:ProjectIndex 不存储符号的具体位置信息(哪一行哪一列)。位置信息存储在 MergedIndex 的按文件分片中。ProjectIndex 的角色是"目录"——告诉你符号存在于哪些文件中,然后你去对应的 MergedIndex 分片中查找具体位置。 - -这种分离使得 ProjectIndex 保持紧凑。引用文件位图使用 Roaring Bitmap 压缩存储,内存效率很高。 - -### MergedIndex:按文件分片的索引 - -MergedIndex 是索引系统的核心存储层。它为项目中的每个文件维护一个分片,存储该文件中所有符号的具体出现位置和关系信息。 - -MergedIndex 的核心特性是**编译上下文合并**。同一个头文件可能被多个源文件包含,每次编译可能产生不同的符号关系(例如由于条件编译、模板实例化等)。MergedIndex 将这些来自不同编译上下文的索引数据合并存储在同一个分片中。 - -合并使用内容寻址的去重:对每次编译产出的 FileIndex 计算内容哈希,相同内容的不同编译上下文共享同一份数据。这避免了头文件被多次索引时的重复存储。 - -MergedIndex 分片同时存储文件内容,用于字节偏移到 LSP 位置(行号/列号)的转换。 - -### 符号标识:SymbolHash - -系统中的每个符号通过 SymbolHash(`uint64_t`)标识。它基于 Clang 的 USR(Unified Symbol Resolution)生成——USR 是符号身份的规范字符串表示,编码了命名空间、类名、函数签名、模板参数等信息。 - -SymbolHash 的关键属性: - -- **跨文件一致**:同一个符号在不同文件中的 SymbolHash 相同,这是跨文件导航的基础 -- **紧凑高效**:64 位整数比可变长度的 USR 字符串更适合作为 DenseMap 的键 -- **确定性**:相同的声明总是产生相同的哈希值 - -### 符号关系模型 - -索引的核心数据是符号之间的**关系(Relation)**。每条关系包含: - -- **关系类型(RelationKind)**:定义、声明、引用、弱引用、读、写、接口、实现、类型定义、基类、派生类、构造函数、析构函数、调用者、被调用者等 -- **位置**:关系发生的源码范围 -- **目标符号**:关系的另一端(用于类型关系,如继承和调用) - -与之对应的是**出现(Occurrence)**——记录符号在某个位置的简单出现,不带关系类型。 - -这两者的区分使得不同的查询可以高效完成: - -- **跳转到定义**:查找 RelationKind::Definition -- **查找引用**:查找 RelationKind::Reference -- **调用层次**:查找 RelationKind::Caller / Callee -- **类型层次**:查找 RelationKind::Base / Derived -- **光标下的符号**:查找 Occurrence - -## 查询流程 - -### 跨文件查询 - -以"跳转到定义"为例,查询流程为: - -1. **定位光标符号**:在当前文件中,通过字节偏移查找光标位置的 Occurrence,获取 SymbolHash -2. **查询符号目录**:在 ProjectIndex 中查找该 SymbolHash 引用的所有文件 -3. **分文件查询关系**:对于每个引用文件,判断它是否处于打开状态: - - 如果该文件**已打开**(有 Session),使用其 OpenFileIndex 查找关系数据,**跳过**对应的 MergedIndex 分片 - - 如果该文件**未打开**,使用 MergedIndex 分片查找关系数据 - -关键点:对于同一个文件,OpenFileIndex **优先于** MergedIndex——打开的文件优先查 OpenFileIndex,当 Session AST 为脏或不可用时回退到 MergedIndex。未打开的文件始终查 MergedIndex。这确保了打开文件的查询结果尽可能反映最新的缓冲区内容。 - -### 打开文件的实时覆盖 - -每个打开的文件在 Session 中维护一个 OpenFileIndex,存储最近一次编译的索引数据。它包含与 MergedIndex 分片相同类型的数据(关系和出现),但来自内存中的 AST 编译结果而非后台索引。 - -设计原则:打开文件的索引不写入 Workspace(不影响全局状态),只有文件保存后才会通过后台索引更新 MergedIndex。这保证了全局索引的稳定性——不会因为用户正在输入一半的代码而污染其他文件的查询结果。 - -在查询时,如果某个引用文件处于打开状态,查询会优先使用该文件的 OpenFileIndex,当 AST 为脏或不可用时回退到 MergedIndex 分片。这确保了用户尽可能看到最新的缓冲区内容对应的结果。 - -## 后台索引 - -### 调度策略 - -后台索引由 Indexer 管理。它维护一个待索引文件队列,通过以下策略调度: - -- **空闲超时批处理**:文件加入队列后不会立即索引,而是设置一个空闲定时器。当定时器触发时,批量处理队列中的所有文件。这避免了频繁的小量索引。 -- **暂停/恢复**:当用户发起交互式请求(hover、completion 等)时,后台索引可以暂停,优先处理用户请求。支持嵌套暂停语义。 - -### 索引合并 - -后台索引完成后,TUIndex 的合并过程为: - -1. 将 TUIndex 中的符号合并到 ProjectIndex 的全局符号表 -2. 将 TUIndex 中每个文件的 FileIndex 合并到对应的 MergedIndex 分片 -3. 更新引用文件位图 - -合并是增量的——只处理新数据,不需要重建整个索引。 - -## 序列化与持久化 - -索引使用 FlatBuffers 序列化,支持高效的磁盘持久化和惰性加载。FlatBuffers 的零拷贝特性使得索引可以通过内存映射直接访问,避免了完整的反序列化开销。 - -每个 MergedIndex 分片以独立文件存储在磁盘上,按需加载。ProjectIndex 作为全局数据整体序列化。 - -## 设计决策与权衡 - -**为什么是三级而不是两级?** ProjectIndex 和 MergedIndex 的分离是关键。如果将位置信息也存入 ProjectIndex,它的体积会急剧膨胀,无法作为轻量的"目录"使用。而如果没有 ProjectIndex,每次跨文件查询都需要扫描所有 MergedIndex 分片来定位符号所在的文件,这在大型项目中不可接受。 - -**为什么打开文件的索引不写入全局状态?** 这是两层状态模型(Workspace vs Session)的一部分。未保存的编辑可能是不完整的、有语法错误的代码。将其写入全局索引会导致其他文件的查询结果受到污染。只有保存到磁盘的稳定状态才应该影响全局索引。 - -**为什么使用内容寻址去重?** 头文件经常被多个源文件包含,每次编译都产生相同的索引数据。内容寻址避免了存储 N 份相同的数据。当头文件内容变化时,旧的数据自然被新的替代。 - -## 已知限制与改进方向 - -- **符号表的局部性**:目前所有符号(包括函数内的局部变量等)都被合并到 ProjectIndex 的全局符号表中。这导致大量只在单个文件内部有意义的符号被全局存储,浪费内存和合并时间。改进方向是引入文件级的 SymbolTable——局部符号留在 TUIndex / MergedIndex 分片中,只有真正跨文件引用的符号才进入 ProjectIndex。 -- **模糊符号搜索**:目前的符号搜索是简单的子串匹配。对于 agent 和 workspace symbol 等场景,需要更高效的模糊搜索索引(分词、前缀树或 n-gram 索引等)。 diff --git a/docs/zh/design/module-graph.md b/docs/zh/design/module-graph.md new file mode 100644 index 000000000..c1dd4c52e --- /dev/null +++ b/docs/zh/design/module-graph.md @@ -0,0 +1,170 @@ +# 模块编译 + +## 背景 + +C++20 引入了模块(Modules),改变了 C++ 自诞生以来的独立编译模型。传统的 C++ 编译中,每个源文件独立编译为目标文件,文件之间通过头文件共享声明。模块打破了这种独立性: + +```cpp +// math.cppm — 模块接口单元 +export module math; +export int add(int a, int b) { return a + b; } + +// main.cpp — 导入模块 +import math; +int main() { return add(1, 2); } +``` + +编译 `main.cpp` 之前,必须先编译 `math.cppm` 的模块接口单元,产出预编译模块文件(PCM)。如果项目中有多层模块依赖——A 导入 B,B 导入 C——编译顺序必须是 C → B → A。这构成了一个有向无环图(DAG),节点是模块文件,边是 `import` 关系。 + +构建系统(CMake、Ninja 等)天然适合处理这种 DAG 调度:扫描所有文件、构建完整依赖图、按拓扑序编译。但语言服务器面临的挑战不同: + +**实时性**。用户打开一个导入了模块的文件后,期望在数百毫秒内获得编辑反馈。等待整个模块图编译完成是不可接受的。语言服务器需要一种惰性的、按需的编译策略——只编译当前文件实际需要的那些模块。 + +**文件变更的级联影响**。当用户修改一个模块接口文件并保存时,所有直接或间接依赖它的模块的 PCM 都已过时。语言服务器需要检测这种情况,取消正在进行的编译,标记受影响的模块为脏,在下次需要时重新编译。 + +**并发与取消**。多个文件可能同时需要同一个模块的 PCM。语言服务器需要避免重复编译,让后到达的请求等待先发起的编译完成。同时,当用户关闭了触发编译的文件、不再需要某个模块时,正在进行的编译应该能被取消以释放资源。 + +**临时的循环依赖**。虽然 C++ 模块不允许循环 `import`,但用户在编辑过程中可能临时引入循环。语言服务器需要检测并报错,而不是死锁。 + +在 clangd 中,C++20 模块支持长期处于实验阶段,缺乏专门的设计。[clangd/clangd#1293](https://github.com/clangd/clangd/issues/1293) 是 clangd 项目方对模块支持问题的总结——结论是"这不是能靠修 bug 解决的问题,需要全新的设计和基础设施",并且当时"没有人有计划或时间来做这件事"。多年过去,社区中相关问题持续出现:[clangd/clangd#2569](https://github.com/clangd/clangd/issues/2569) 列举了模块中缺失的关键功能(重命名、查找引用等);[clangd/clangd#2292](https://github.com/clangd/clangd/issues/2292) 和 [clangd/clangd#2497](https://github.com/clangd/clangd/issues/2497) 反映了 clangd 与构建系统共用 PCM 文件导致的文件锁定冲突——clangd 打开 PCM 文件后,构建系统无法覆写更新。 + +这些问题的共同根源是缺乏一个为语言服务器场景专门设计的模块编译调度系统。clice 通过 CompileGraph 解决这一问题——一个基于引用计数的编译 DAG,支持惰性构建、按需编译、实时取消和依赖级联。 + +## 设计 + +CompileGraph 是一个运行时构建的编译依赖图。每个节点(CompileUnit)代表一个模块文件,边代表 `import` 依赖关系。 + +### 惰性构建 + +与构建系统在编译前扫描所有文件并构建完整 DAG 不同,CompileGraph 是惰性构建的——节点只在第一次被编译请求触及时才解析其依赖并加入图中。用户打开一个文件时,只有该文件实际需要的模块链被扫描和编译,而不是整个项目的模块图。 + +CompileGraph 有两种编译入口: + +- **编译模块**:编译指定模块及其所有传递依赖,产出 PCM 文件。用于模块接口单元自身的编译。 +- **编译依赖**:编译指定文件的所有模块依赖,但不编译文件本身。用于普通源文件——它们不属于模块 DAG,但可能通过 `import` 使用模块。 + +### 引用计数 + +引用计数(interest counting)是 CompileGraph 的核心调度机制,追踪"当前有多少活跃请求关心某个编译单元",回答两个问题:是否需要启动编译?是否可以取消编译? + +当某个文件需要一个模块时,该模块的引用计数加一(acquire);当请求结束或被取消时,引用计数减一(release)。引用计数归零意味着当前没有任何活跃请求关心这个模块。 + +引用计数的管理通过 RAII 守卫完成:编译请求创建守卫时自动 acquire,守卫析构时自动 release。这与协程取消的语义天然契合——取消导致协程帧销毁,析构器释放引用计数,不需要额外的清理代码。 + +引用计数有两个层级:请求级和编译任务级。请求级引用表示"这个请求在等待某个模块的 PCM";编译任务级引用表示"这个模块的编译任务正在等待其直接依赖完成"。后者确保了一个模块在被其他模块的编译任务依赖时不会因为请求取消而被过早终止。 + +### 编译轮次 + +每次编译尝试构成一个编译轮次(Round)。等待者通过完成事件得知某轮编译结束,通过结果决定下一步动作。一个轮次有三种可能的结果: + +- **成功(Success)**:PCM 产出成功,清除脏标志 +- **失败(Failed)**:编译出错(依赖失败、循环依赖等),保留脏标志 +- **过时(Stale)**:编译被取消(文件在编译期间被修改,或引用计数归零),等待者自动驱动新一轮 + +失败不是粘性的——保留脏标志意味着用户修复错误后,下次请求会自然触发重试,不需要手动重启服务器。 + +### 脏状态与世代计数器 + +CompileUnit 使用脏标志表示需要编译。世代计数器是一个单调递增的值,在每次文件更新时递增,用于检测异步竞态:编译任务启动时记录当前世代,完成时比较——如果不一致,说明编译期间文件被修改过,结果已过时,这轮编译的结果为 Stale。 + +### 依赖关系 + +每个 CompileUnit 维护正向依赖(它导入的模块)和反向依赖(导入它的模块)。正向依赖通过惰性解析获得;解析时同时填充反向依赖。反向依赖是文件变更时级联通知的基础——从被修改的模块出发,沿反向依赖边找到所有受影响的模块。 + +模块名到文件的映射由 DependencyGraph 维护(见[依赖扫描](dependency-scanning.md))——启动时的快速扫描会发现所有模块声明,建立模块名 → 文件路径的注册表。CompileGraph 使用这个注册表将 `import` 语句解析为具体的文件路径。 + +## 实现 + +### 编译流程 + +一个模块编译请求的完整流程: + +``` +请求进入 + │ + ├─ RAII 守卫对目标模块 acquire + │ + ├─ 目标模块不脏? ──→ 直接返回(PCM 已可用) + │ + ├─ 没有正在进行的编译? ──→ 启动编译任务 + │ │ + │ ├─ 惰性解析依赖(扫描 import 声明) + │ ├─ 检测自环 + │ ├─ 对直接依赖 acquire + │ ├─ 并行等待所有依赖编译完成 + │ ├─ 分派到工作进程(产出 PCM) + │ └─ 检查世代计数器 ──→ 不一致则结果为 Stale + │ + ├─ 等待编译轮次完成 + │ + └─ 根据结果:Success 返回 / Failed 报错 / Stale 重试 +``` + +惰性解析使用 Clang 预处理器进行精确扫描,与启动阶段[依赖扫描](dependency-scanning.md)使用的快速词法扫描不同。快速扫描不展开宏和条件编译,适合构建全局的包含关系概览;精确扫描展开所有预处理指令,得到该文件在当前编译命令下实际的模块依赖列表。解析结果被缓存,后续编译复用缓存,直到文件更新时重置。 + +> 模块实现单元(没有 `export` 的 `module X;`)隐式依赖对应的模块接口单元。精确扫描会检测这种情况并自动添加依赖。 + +### 延迟零引用取消 + +当引用计数降为零时,不立即取消编译,而是延迟一个事件循环 tick 后再检查——如果引用计数仍为零,才真正取消。 + +这是为了处理编译请求被取代的场景。当用户在编译期间编辑了文件,新的编译请求取代旧的:旧请求释放引用(计数瞬间降为零),紧接着新请求在同一 tick 内建立引用(计数回升)。立即取消会导致共享依赖的编译被不必要地终止和重启。延迟一个 tick 让这种引用交接平滑完成,避免了对正在进行的共享编译的干扰。 + +### 级联更新 + +当模块文件被保存(didSave)时,CompileGraph 执行级联更新: + +1. 重置被修改文件的已解析标志——下次编译时重新扫描依赖,因为文件可能增删了 `import` +2. 清除旧的正向依赖边 +3. 标记为脏,递增世代计数器 +4. 沿反向依赖边遍历所有传递依赖方,对每个受影响的模块:取消其编译轮次、标记为脏、递增世代 +5. 返回所有被标记为脏的文件列表,调用方据此清除对应的 PCM 缓存 + +级联更新不修改引用计数——现有的等待者保持它们的引用。它们在观察到编译结果为 Stale 后,会自动驱动新一轮编译。 + +### 循环依赖检测 + +在等待某个依赖的编译完成之前,CompileGraph 检查是否存在等待环:从目标节点出发,沿依赖链搜索,只跟随正在编译中的节点,检查是否会回到当前等待者。如果检测到环,立即返回失败,避免死锁。 + +### RAII 守卫与结构化并发 + +协程取消在 kotatsu 中的语义是销毁协程帧——挂起点之后的代码不会执行,只有已构造对象的析构器保证执行。CompileGraph 利用这一点,将所有清理逻辑放在两层 RAII 守卫中: + +- **RefGuard**(请求级):持有请求对模块的根引用。请求完成或被取消时析构,释放引用计数。 +- **UnitGuard**(编译轮次级):管理一轮编译的所有状态。析构时发布结果、清除编译中标志、释放所有已获取的依赖引用、触发完成事件通知等待者。 + +所有编译任务通过 kota::task_group 管理,提供结构化并发保证:关闭时先取消所有任务,再等待它们的帧退出,确保没有悬空的编译任务。 + +### PCM 缓存 + +PCM 文件采用内容寻址的路径命名——由模块名和编译参数的哈希值决定文件名,存储在专用的缓存目录中。这与构建系统的产物完全隔离,避免了文件锁定冲突。 + +PCM 缓存使用两层新旧检测:先比较依赖文件的修改时间(mtime),时间变化时再比对内容哈希。只有依赖文件的内容实际发生变化时才重新编译,避免"touch 但未修改"导致的不必要重编译。缓存元数据持久化到磁盘上的 `cache.json`,服务器重启后可以恢复。 + +### 与编译流程的集成 + +CompileGraph 在服务器启动阶段初始化。Compiler 在每次编译文件之前,通过 CompileGraph 确保所有模块依赖就绪——这是编译准备阶段的第一步,在 PCH 构建之前执行。 + +对于用户在编辑器中新增的 `import` 语句(尚未保存,因此 CompileGraph 不知道),Compiler 会额外扫描缓冲区内容,识别出新增的模块依赖并尝试构建对应的 PCM。这是一个补偿机制,覆盖了用户编辑但未保存时的体验。 + +## FAQ + +- **为什么用引用计数而不是任务队列?** 任务队列无法表达"这个编译不再被任何人需要"的语义。当用户关闭了触发编译的文件时,继续编译浪费资源。引用计数精确追踪需求,使取消决策有据可依——不是基于超时或启发式判断,而是基于"当前还有没有请求在等这个结果"。 + +- **为什么延迟一个 tick 而不是立即取消?** 在事件循环模型中,一个操作的多个步骤在同一 tick 内完成。当编译请求被新版本取代时,旧引用先释放、新引用随后建立,中间存在瞬间的零引用。立即取消会导致共享依赖被不必要地终止——延迟一个 tick 确保同一 tick 内的引用交接不触发取消。 + +- **为什么用世代计数器而不是锁?** 主进程是单线程事件循环,没有数据竞争。需要检测的"竞态"来自异步操作的时序——"编译期间文件是否被更新过"。世代计数器用最小的开销回答这个问题,不需要引入锁。 + +- **为什么惰性解析依赖?** 依赖解析需要对每个模块文件运行 Clang 预处理器(精确扫描),成本不低。如果启动时解析所有模块的依赖,会增加启动时间。惰性解析确保只扫描实际需要编译的模块——用户打开一个文件不会触发整个模块图的扫描。 + +- **失败为什么不是粘性的?** 用户正在编辑代码,语法错误是常态。如果把失败标记为持久状态,用户修复错误后仍然无法获得正确结果,必须重启服务器。保留脏标志让下次请求自然触发重试。 + +- **为什么 PCM 存储在独立的缓存目录?** clangd 共用构建系统的 PCM 文件会导致文件锁定冲突——语言服务器读取 PCM 时,构建系统无法覆写更新(见 [clangd/clangd#2292](https://github.com/clangd/clangd/issues/2292))。clice 使用内容寻址的独立缓存,避免了这个问题。代价是磁盘空间的额外占用和首次编译的额外时间。 + +## 已知局限 + +- **编译任务的内存积累**。每轮编译在 task_group 中创建一个任务。已完成的任务帧在 task_group 析构时才回收,而非完成时立即回收。长期运行的服务器中,已完成的任务帧会持续积累。 + +- **依赖解析的确定性**。解析依赖和实际编译之间存在时间窗口。如果文件在此期间被修改,解析出的依赖可能与编译时的实际依赖不一致。世代计数器可以检测到这种情况并触发重试,但会产生额外的编译开销。 + +- **未保存文件的模块依赖覆盖有限**。CompileGraph 基于磁盘文件构建依赖关系。用户在编辑器中新增 `import` 但未保存时,Compiler 通过缓冲区扫描来补偿,但如果导入的模块本身还未被 CompileGraph 管理(例如一个尚未保存的新模块文件),对应的 PCM 无法被构建。用户需要先保存模块接口文件,再保存导入方文件。 diff --git a/docs/zh/design/module.md b/docs/zh/design/module.md deleted file mode 100644 index 6a5857556..000000000 --- a/docs/zh/design/module.md +++ /dev/null @@ -1,143 +0,0 @@ -# 模块编译 - -## 背景 - -C++20 引入了模块(Modules),这是 C++ 编译模型自诞生以来最大的变革。传统的 C++ 编译是"按文件独立编译"——每个源文件独立编译为目标文件,文件之间通过头文件共享声明。模块打破了这一独立性:一个文件 `import` 另一个模块时,必须先编译被导入模块的**模块接口单元**,产出**预编译模块文件(PCM)**,然后才能编译导入方。 - -这引入了**编译时依赖关系**——一个有向无环图(DAG),其中节点是模块文件,边是 `import` 关系。构建系统(CMake、Ninja 等)天然支持这种 DAG 调度,但语言服务器面临的挑战截然不同: - -**实时性要求**:构建系统可以一次性扫描所有文件、构建完整的依赖图、按拓扑序编译。语言服务器不能——用户打开一个文件后期望在毫秒级获得编辑反馈。等待整个模块图编译完成是不可接受的。 - -**按需编译**:用户只打开了项目中的几个文件,没有必要编译整个模块图。语言服务器需要的是"只编译当前文件需要的那些模块"——一种惰性的、增量的编译策略。 - -**文件变更的级联更新**:当用户修改了一个模块接口文件并保存时,所有直接或间接依赖它的模块的 PCM 都已过时。语言服务器需要在用户不感知的情况下取消正在进行的编译、标记受影响的模块为脏、并在下次需要时重新编译。 - -**并发与竞态**:多个文件可能同时需要同一个模块的 PCM。如果两个请求同时触发同一模块的编译,需要避免重复编译,同时确保后到达的请求能等待先发起的编译完成。 - -**循环依赖**:虽然 C++ 模块不允许循环 `import`,但在编辑过程中用户可能临时引入循环依赖。语言服务器需要检测这种情况并优雅地报错,而不是死锁。 - -在 clangd 中,C++20 模块支持长期处于实验阶段。社区中反复出现的问题包括:模块编译的依赖解析不完整,导致找不到 PCM 文件;模块文件修改后依赖方没有重新编译,导致诊断信息过时;缺乏对模块依赖的实时跟踪,需要手动重启服务器才能获得正确的编译结果。 - -这些问题的根源在于缺乏一个专门的模块编译调度系统。clice 通过 CompileGraph——一个基于引用计数的编译 DAG——从根本上解决这些问题。 - -## 设计方案 - -### 核心思想 - -CompileGraph 是一个运行时构建的编译依赖 DAG。每个节点(CompileUnit)代表一个模块文件,边代表 `import` 依赖关系。与构建系统的静态 DAG 不同,CompileGraph 是**惰性构建**的——只有当某个模块被实际需要时,才会解析它的依赖并加入图中。 - -CompileGraph 的调度基于**引用计数(interest counting)**:当某个文件需要一个模块时,该模块的引用计数加一;当不再需要时减一。引用计数归零意味着当前没有任何活跃请求关心这个模块,可以取消它正在进行的编译以释放资源。 - -### CompileUnit 的关键概念 - -每个 CompileUnit 围绕以下几个核心概念组织: - -**依赖关系**:每个单元维护正向依赖(它 `import` 的模块)和反向依赖(依赖它的模块)。正向依赖通过惰性解析获得,反向依赖在解析时自动填充。反向依赖是文件变更时级联通知的基础。 - -**脏状态与世代**:dirty 标志表示该单元需要重新编译。世代计数器(generation)在每次文件更新时递增,用于检测异步编译结果是否过时——编译任务启动时捕获世代值,完成时比较是否一致。 - -**编译轮次**:每次编译产生一个轮次(round),包含完成事件和结果。等待者通过完成事件得知编译完成,通过结果决定下一步动作。取消源(cancellation_source)用于取消当前轮次的编译任务。 - -**引用计数**:追踪有多少活跃请求关心该单元。归零时可以取消编译以释放资源。 - -### 引用计数与 RAII 守卫 - -引用计数是 CompileGraph 调度的核心机制。它追踪"当前有多少活跃请求关心某个编译单元",决定编译任务的启动和取消。 - -CompileGraph 提供两种 RAII 守卫来管理引用计数: - -**RefGuard(请求级守卫)**:一个编译请求的入口会创建 RefGuard,对所有需要的编译单元执行 acquire(引用计数加一)。RefGuard 析构时自动 release(引用计数减一)。这确保即使请求被取消(协程帧被销毁),引用计数也会正确释放。 - -**UnitGuard(编译轮次守卫)**:每个编译任务内部使用 UnitGuard 维护编译轮次的状态。它负责:在依赖解析后 acquire 所有直接依赖的引用计数;在编译完成或取消时发布结果、清除 compiling 标志、release 所有已获取的依赖引用、触发完成事件通知等待者。 - -使用析构器而非协程 finally 块来执行清理是一个关键设计选择——取消通过协程帧销毁实现,因此清理逻辑必须放在析构器中才能保证执行。 - -### 惰性依赖解析 - -依赖解析是昂贵的操作——需要扫描模块文件的内容,解析 `import` 声明。CompileGraph 不在节点创建时立即解析依赖,而是在首次需要编译该节点时才调用 resolve_fn 进行解析。 - -resolve_fn 由外部提供(通常基于模块文件的精确扫描),返回该模块直接依赖的文件列表。解析结果缓存在 CompileUnit 的 dependencies 中,同时反向填充 dependents 列表。后续编译复用缓存的依赖关系,除非文件更新重置了 resolved 标志。 - -### 编译流程 - -CompileGraph 提供两个入口点: - -**compile(path_id)**——编译指定模块及其所有传递依赖: - -1. RefGuard 对目标模块执行 acquire -2. 如果目标模块不脏,直接返回 -3. 如果没有正在进行的编译,启动编译任务 -4. 等待编译轮次完成 -5. 根据结果决定返回成功、失败或重试(过时则重试) - -**compile_deps(path_id)**——编译指定文件的所有传递模块依赖,但不编译文件本身。这用于普通源文件(非模块文件)——它们本身不是模块 DAG 的一部分,但可能 `import` 模块。 - -编译任务(unit_body)的内部流程: - -1. 解析依赖(ensure_resolved) -2. 检测自环(自己 import 自己) -3. 对所有直接依赖执行 acquire -4. 递归等待所有依赖编译完成 -5. 将实际编译工作分派到无状态工作进程(dispatch_fn) -6. 检查世代计数器——如果编译期间文件被更新,结果已过时,不清除 dirty 标志 - -### 零引用延迟取消 - -当一个编译单元的引用计数降为零时,意味着当前没有活跃请求关心它。但直接取消可能过于激进——在依赖切换场景中,引用计数可能在同一个事件循环轮次内先降为零再重新增加。 - -CompileGraph 采用延迟一个事件循环 tick 的策略:引用计数降为零时,不立即取消,而是调度一个延迟检查。延迟检查在下一个事件循环 tick 执行——如果此时引用计数仍为零,才真正取消编译。这避免了瞬态的零引用触发不必要的取消和重启。 - -### 世代计数器与竞态检测 - -在异步系统中,编译任务被分派到工作进程后需要等待结果返回。在这段时间内,用户可能修改了文件,导致文件内容已经更新。如果此时将旧的编译结果标记为"干净",会导致使用过时的 PCM。 - -世代计数器解决这个问题:编译任务启动时捕获当前世代值,编译完成后比较当前世代值是否与捕获的一致。如果不一致,说明期间发生了文件更新,编译结果已过时——不清除 dirty 标志,后续请求会触发重新编译。 - -### 文件更新与级联传播 - -当一个模块文件被修改(通过 didSave)时,CompileGraph 的 update() 方法执行级联更新: - -1. 重置被修改文件的 resolved 标志(强制下次编译重新解析依赖) -2. 清除旧的依赖边 -3. 标记为脏,递增世代计数器 -4. 沿 dependents 反向边遍历所有传递依赖方 -5. 对每个受影响的依赖方:取消其编译轮次、标记为脏、递增世代计数器 - -update() 返回所有被标记为脏的模块列表,供调用方清除对应的 PCM 缓存。 - -注意 update() 不修改引用计数——现有的等待者保持它们的引用。它们在观察到编译结果为"过时"后会自动重试,驱动新一轮编译。 - -### 循环依赖检测 - -虽然合法的 C++20 模块不允许循环依赖,但编辑过程中可能临时引入。CompileGraph 在等待依赖编译完成之前检测是否存在等待环:沿目标的依赖链搜索,只跟随正在编译的节点,检查是否会回到当前等待者。如果检测到环,立即返回失败,避免死锁。 - -### 编译结果与重试语义 - -编译轮次有三种结果: - -- **Success**:编译成功,清除 dirty 标志 -- **Failed**:编译失败(依赖失败、分派失败、或循环依赖)。不清除 dirty 标志——下次显式请求可以重试 -- **Stale**:编译被取消(因为文件更新或引用归零)。等待者自动重试新一轮编译 - -Stale 的重试是有界的——每次 update() 只递增一次世代,不会累积。没有持续的更新流时,重试会自然终止。 - -### 结构化并发与关闭 - -所有编译任务通过 kota::task_group 管理。task_group 提供结构化并发保证——析构时等待所有子任务完成。CompileGraph 的 shutdown() 方法通过取消 task_group 中的所有任务并等待它们退出来实现优雅关闭。 - -## 设计决策与权衡 - -**为什么用引用计数而不是简单的任务队列?** 任务队列无法表达"这个编译不再被任何人需要"的语义。在模块编译场景中,用户可能在一个模块编译过程中关闭了触发编译的文件,此时继续编译是浪费资源的。引用计数精确追踪需求,使得取消决策有据可依。 - -**为什么延迟一个 tick 而不是立即取消?** 在事件循环模型中,一个操作的多个步骤可能在同一个 tick 内完成。如果旧的引用在新的引用建立之前被释放(瞬态零引用),立即取消会导致不必要的工作重做。延迟一个 tick 确保同一 tick 内的引用交接不会触发取消。 - -**为什么用世代计数器而不是锁?** 主进程是单线程的事件循环,没有数据竞争问题——竞态来自于异步操作的时序。世代计数器用最小的开销检测"编译期间是否发生了文件更新",不需要引入锁的复杂性。 - -**为什么依赖解析是惰性的?** 依赖解析需要读取和扫描文件内容,成本不低。如果在节点创建时就解析所有依赖,用户只打开一个文件就可能触发整个模块图的扫描。惰性解析确保只扫描实际需要编译的模块。 - -**为什么失败不是粘性的?** 模块编译失败通常是暂时的——用户正在编辑代码,语法错误是常态。将失败标记为持久状态会阻止用户修复错误后获得正确的结果。不清除 dirty 标志意味着下次请求会自然重试。 - -## 已知限制 - -- **编译图的内存积累**:每个编译轮次会在 task_group 中创建一个任务。已完成的任务帧在 task_group 析构时才被回收,而非完成时立即回收。长期运行的服务器中可能积累大量已完成的任务帧。 -- **依赖解析的确定性**:resolve_fn 的结果取决于文件当前的内容。如果文件在 resolve 和实际编译之间被修改,解析出的依赖可能与编译时的不一致。世代计数器可以检测到这种情况,但会导致额外的重试。 diff --git a/docs/zh/design/overview.md b/docs/zh/design/overview.md index ff07d5118..946b99681 100644 --- a/docs/zh/design/overview.md +++ b/docs/zh/design/overview.md @@ -4,163 +4,107 @@ ## 项目定位 -clice 是一个全新的 C++ 语言服务器,从架构层面重新设计,解决以往 C++ 语言服务器中长期存在的根本性问题。核心创新包括: +clice 是一个全新的 C++ 语言服务器,从架构层面重新设计,解决以往 C++ 语言服务器中长期存在的问题。主要特点: -- **编译上下文(Compilation Context)**:同一个文件在不同的编译上下文下可能产生不同的结果。clice 将这一概念作为一等公民贯穿整个设计——从编译、索引到查询响应。现有的所有语言服务器(不仅是 C++)都没有正式处理这一概念。 +- **编译上下文**:clice 是第一个将编译上下文作为正式概念引入的语言服务器。编译、索引、查询的每一步都明确区分当前使用的编译上下文,用户可以查询和切换。详见 [编译上下文](compilation-context.md)。 -- **多进程架构**:通过 master + worker 的进程模型隔离 Clang 的内存泄漏和崩溃问题,同时实现优先级调度和实时内存监控。 +- **多进程架构**:通过 master + worker 的进程模型隔离 Clang 的内存泄漏和崩溃问题,同时实现优先级调度和实时内存监控。详见 [多进程架构](multi-process.md)。 -- **协程异步模型**:基于 C++20 协程和 kotatsu 库,取代传统的回调式异步,使业务逻辑更加清晰。 +- **协程异步模型**:基于 C++20 协程和 kotatsu 库,取代传统的回调式异步,使业务逻辑更清晰。 -- **实时模块编译系统**:基于引用计数的 C++20 模块编译 DAG,支持实时取消和依赖级联,是目前已知的最优实时模块编译调度方案。 +- **实时模块编译**:基于引用计数的 C++20 模块编译 DAG,支持实时取消和依赖级联。详见 [模块编译图](module-graph.md)。 -clice 的远期目标不仅是语言服务器,还将集成 clang-tidy 等工具的并发调度,实现跨翻译单元的优化(如头文件去重检查),成为 C++ 工具生态中统一的高级平台。 +## 模块概览 -## 基础层 +### `src/support/` — 基础工具库 -### `src/support/` +通用工具和基础设施,被其他所有模块共享。 -通用工具库。提供日志、文件系统抽象、字符串操作、模式匹配等基础设施。 +- `PathPool`:将文件路径内部化为 `uint32_t` 标识符,在整个系统中用作文件的稳定标识 +- `FuzzyMatcher`:分词感知的模糊匹配,用于代码补全和符号搜索 +- Markup / Doxygen:文档注释的解析与格式化 +- 日志、文件系统抽象、字符串工具等 -关键组件: +### `src/command/` — 编译命令处理 -- **路径池(PathPool)**:将文件路径内部化为紧凑的 `uint32_t` 标识符,在整个系统中用作文件的稳定标识。这是全局共享的基础设施,几乎所有涉及文件路径的模块都依赖它。 -- **日志系统**:基于 spdlog 的结构化日志。 -- **模糊匹配器(FuzzyMatcher)**:用于代码补全等场景的分词感知模糊匹配。 -- **标记文档(Markup)/ Doxygen**:文档注释的解析与格式化。 -- **去重集合(StringSet / ObjectSet)**:字符串和对象的去重内部化,通过 bump allocator 保证指针稳定性。 +从编译数据库(CDB)读取的命令是构建系统生成的原始命令,不能直接交给 Clang 前端使用——可能包含仅用于代码生成的选项、缺少系统头文件搜索路径、或包含语言服务器不需要的参数。这个模块负责对原始命令进行分类、过滤、工具链探测和去重,转化为语言服务器可消费的编译参数。 -### `src/command/` +- `CompilationDatabase`:加载 `compile_commands.json`,将编译命令分为影响语义的标志(canonical)和仅影响用户内容的标志(patch,如 `-I`、`-D`),实现跨文件的命令去重与共享 +- `Toolchain`:通过系统编译器查询完整的编译参数(如系统头文件搜索路径),按 (driver, file extension, non-user-content flags) 缓存查询结果 +- `SearchConfig`:头文件搜索路径的四段式模型(Quoted / Angled / System / After),与 Clang 内部的搜索逻辑一致 -CLI 解析与编译命令处理。从编译数据库(CDB)读取的命令是构建系统生成的原始命令,不能直接交给 Clang 前端使用——它们可能包含仅用于代码生成的选项、缺少系统头文件搜索路径、或者包含语言服务器不需要的参数。这个模块负责对原始命令进行分类、过滤、工具链探测和去重,最终转化为语言服务器可消费的编译参数。 +详见 [编译命令解析](command-resolve.md)。 -关键组件: +### `src/compile/` — 编译抽象 -- **编译数据库(CompilationDatabase)**:加载 `compile_commands.json`,对编译命令进行两级分离——将影响语义的标志(canonical)与仅影响用户内容的标志(patch,如 `-I`、`-D`)分开,实现跨文件的命令去重与共享。 -- **工具链(Toolchain)**:通过系统编译器查询完整的编译参数(如系统头文件搜索路径),按 (driver, file extension, non-user-content flags) 缓存查询结果,避免对同一工具链的重复探测。 -- **参数分类器(ArgumentParser)**:将 Clang 编译选项分类为 codegen-only(丢弃)、discarded(丢弃)、user-content(提取为 patch)三类,确保只保留对语义分析有意义的参数。 -- **搜索配置(SearchConfig)**:头文件搜索路径的四段式模型(Quoted / Angled / System / After),与 Clang 内部的搜索逻辑一致。 +对 Clang 编译器的封装,将 Clang API 抽象为安全、统一的编译接口。这一层是纯粹的编译抽象,不涉及任何服务器逻辑。 -## 编译抽象层 +- `CompilationUnit` / `CompilationUnitRef`:对 Clang AST 上下文的 RAII 封装。`CompilationUnitRef` 提供统一的只读视图,用于访问源码位置映射、预处理指令、AST 节点等,是 `src/feature/` 和 `src/semantic/` 的主要输入 +- `CompilationParams`:描述一次编译的完整配置,包括编译类型(Preamble / Content / Completion / Indexing 等)、文件重映射、PCH/PCM 复用等 -### `src/compile/` +### `src/syntax/` — 轻量级语法处理 -对 Clang 编译器的封装。将原始的 Clang API 抽象为安全、统一的编译接口,屏蔽底层细节。 +不依赖完整 AST 的语法层处理,运行在编译之前,用于快速获取文件的结构信息和依赖关系。 -这一层是纯粹的编译抽象,不涉及任何服务器逻辑。它接收编译参数,驱动 Clang 完成编译,产出编译结果——包括 AST、预处理器状态、源码位置映射等。 +- `Lexer`:基于 Clang raw lexer 的 token 级别工具,不经过预处理器,用于指令扫描、include 路径解析等 +- `DependencyGraph`:全局的 include / module 依赖关系图,支持正向查询、反向查询、宿主源文件搜索、include 链查找等 +- 依赖扫描:封装 Clang 的 `DependencyDirectivesScanner`,快速提取 include 和 module 依赖 +- `IncludeResolver`:根据搜索路径配置解析 include 路径到实际文件 -关键组件: +详见 [依赖扫描](dependency-scanning.md)。 -- **编译单元(CompilationUnit / CompilationUnitRef)**:对 Clang AST 上下文的 RAII 封装。`CompilationUnitRef` 提供统一的只读视图,用于访问源码位置映射、预处理指令、文件内容、AST 节点等。这是 `src/feature/` 和 `src/semantic/` 的主要输入。 -- **编译参数(CompilationParams)**:描述一次编译操作的完整配置,包括编译类型(Preprocess / Content / Preamble / ModuleInterface / Completion / Indexing)、文件重映射、PCH/PCM 复用、取消标志等。 -- **编译类型**:不同的编译类型产出不同的结果。Preamble 产出 PCH,ModuleInterface 产出 PCM,Content 产出完整 AST,Completion 产出补全候选项。它们共享同一套编译参数框架,但语义不同。 +### `src/semantic/` — 语义分析 -### `src/syntax/` +超越 Clang 原生 API 的语义分析能力。接收 `CompilationUnitRef`,提取更高层次的语义信息。 -轻量级语法处理,不依赖完整 AST。 +- `SemanticVisitor`:AST 遍历器,记录每个符号的出现位置和关系(定义、引用、继承、调用等),是索引数据的主要生产者 +- `TemplateResolver`:通过伪实例化技术解析依赖名称(dependent names),使语义分析能穿透模板上下文。详见 [模板解析器](template-resolver.md) +- `SymbolKind` / `RelationKind`:细粒度的符号种类和关系类型 -这一层处理那些不需要完整编译就能完成的任务:词法分析、依赖扫描、include 路径解析等。它运行在编译之前,用于快速获取文件的结构信息和依赖关系。 +### `src/index/` — 符号索引 -关键组件: +符号索引系统,支持跨翻译单元的查询。采用三级结构: -- **词法分析器(Lexer)**:基于 Clang 原始词法分析器(raw lexer)的 token 级别工具,提供流式接口。不经过预处理器,不展开宏。用于指令扫描、include 路径解析等场景。 -- **依赖图(DependencyGraph)**:全局的 include / module 依赖关系图。记录每个文件(在每个 SearchConfig 下)的 include 关系,支持正向查询(文件包含了谁)、反向查询(谁包含了这个文件)、以及模块名到文件的映射。还提供 BFS 向上搜索宿主源文件、最短 include 链查找等导航功能。 -- **依赖扫描(Scan)**:封装 Clang 的 `DependencyDirectivesScanner`,快速提取文件的 include 和 module 依赖而不需要完整预处理。 -- **Include 解析器(IncludeResolver)**:根据搜索路径配置解析 include 路径到实际文件。 -- **补全(Completion)**:include 路径补全和 module 导入补全的候选项生成。 +- `TUIndex`:单次编译产出的索引数据,由 `SemanticVisitor` 生成 +- `ProjectIndex`:全局符号表,汇聚所有已索引文件的符号信息,支持按符号哈希查询 +- `MergedIndex`:按文件分片的索引存储,将同一文件在不同编译上下文下产生的索引数据合并 -## 语义分析层 +详见 [符号索引](symbol-index.md)。 -### `src/semantic/` +### `src/feature/` — LSP 功能实现 -超越 Clang 原生 API 的语义分析能力。 +LSP 功能的具体实现。每个功能接收一个 `CompilationUnitRef`,返回对应的 LSP 响应数据。这一层是纯粹的计算层,不涉及网络通信、状态管理或进程调度。 -这一层接收 `CompilationUnitRef`,从中提取更高层次的语义信息——符号关系、符号分类、模板解析等。它的产出是索引系统和 LSP 功能的核心数据来源。 +包括:代码补全、悬停信息、签名帮助、语义高亮、内嵌提示、文档符号、文档链接、折叠范围、格式化、诊断等。 -关键组件: +> `feature/` 只涵盖单文件、基于 AST 的特性实现。跨文件的导航功能(go to definition、find references 等)由 `Indexer` 基于索引数据完成。部分功能存在多阶段处理——例如代码补全中的 include 路径补全在语法层就能完成,不需要完整编译。 -- **语义访问器(SemanticVisitor)**:AST 遍历器,记录每个符号的出现位置和关系(定义、引用、读、写、继承、调用等)。这是索引数据的主要生产者。 -- **模板解析器(TemplateResolver)**:通过伪实例化技术解析依赖名称(dependent names),使语义分析能穿透模板上下文。这是 clice 的一项关键创新,clangd 在此领域能力有限。 -- **符号分类(SymbolKind / RelationKind)**:细粒度的符号种类和关系类型,远比 LSP 协议定义的更加精细。 -- **目标解析(find_target)**:将 AST 节点解析到其声明目标,处理各种间接引用和隐式转换。 +### `src/server/` — 服务器运行时 -## 索引层 +语言服务器的核心运行时,负责将上述所有层组装成一个可运行的服务。 -### `src/index/` +**`protocol/`** — 协议定义。描述主进程与工作进程之间、以及与客户端之间的通信消息格式。包括 Worker 协议(编译/查询/构建请求)、LSP 扩展协议(编译上下文切换等)、以及面向 AI agent 的 agentic 协议。 -符号索引系统,支持跨翻译单元的查询。 +**`workspace/`** — 项目级全局状态。`Workspace` 持有编译数据库、工具链、路径池、依赖图、PCH/PCM 缓存、项目索引等全部项目级状态。核心不变量:打开文件的未保存缓冲区内容不会修改 `Workspace`,它只反映磁盘上的状态。 -索引系统采用三级结构,对应不同的查询场景和生命周期: +**`compiler/`** — 编译调度与索引管理。 -- **TU 索引(TUIndex)**:单次编译产出的索引数据。记录编译涉及的每个文件中的符号出现、关系和 include 图。这是索引的原始数据来源,由 `SemanticVisitor` 在编译过程中生成。 -- **项目索引(ProjectIndex)**:全局符号表。汇聚所有已索引文件的符号信息,支持按符号哈希查询名称、种类和引用文件。用于跨文件的符号搜索和导航。 -- **合并索引(MergedIndex)**:按文件分片的索引存储。将同一文件在不同编译上下文下产生的索引数据合并存储。合并索引是编译上下文感知的核心体现——一个头文件可能被多个源文件包含,每次包含产生不同的符号关系,合并索引将它们统一起来。 +- `Compiler`:编译生命周期的调度器,协调 PCH 构建、模块依赖解析、AST 编译的先后顺序,将编译任务分发给工作进程 +- `CompileGraph`:C++20 模块编译的 DAG 调度器,基于引用计数实现兴趣追踪,支持依赖级联取消 +- `Indexer`:后台索引调度和跨文件查询的入口,综合 `ProjectIndex`、`MergedIndex` 和打开文件的内存索引提供查询结果 -索引使用 FlatBuffers 序列化,支持磁盘持久化和惰性加载。 +**`service/`** — 服务入口与会话管理。 -## 功能层 +- `MasterServer`:顶层协调器,持有 `Workspace`、`Session` 映射、`WorkerPool`、`Compiler`、`Indexer`,将 LSP 请求路由到对应的处理逻辑 +- `Session`:每个打开文件的编辑状态(缓冲区内容、编译版本号、内存索引、PCH 引用等),在 didOpen 时创建,didClose 时销毁 +- `LSPClient` / `AgentClient`:LSP 协议和 agentic 协议的请求处理器 -### `src/feature/` +**`worker/`** — 工作进程管理。 -LSP 功能的具体实现。每个功能接收一个 `CompilationUnitRef`(或 `CompilationParams`),返回对应的 LSP 响应数据。 +- `WorkerPool`:管理工作进程的生命周期和调度。有状态工作进程(`StatefulWorker`)持有 AST 服务查询请求,无状态工作进程(`StatelessWorker`)执行一次性任务(PCH/PCM 构建、补全、索引等) +- 进程容错:工作进程崩溃时自动重启,有状态工作进程崩溃时自动重新分配文档 -这一层是纯粹的计算层——不涉及网络通信、状态管理或进程调度。它只关心"给定一个编译结果,如何产出正确的 LSP 响应"。包括: - -- 代码补全、悬停信息、签名帮助 -- 语义高亮、内嵌提示 -- 文档符号、文档链接 -- 折叠范围、格式化、诊断 - -每个功能实现为一个独立函数,遵循统一的模式。新增 LSP 功能时,参考已有实现即可。 - -> **注意**:`feature/` 只涵盖单文件、基于 AST 的特性实现。跨文件的语言特性(如 go to definition、find references、call hierarchy 等)由 `src/server/compiler/` 中的 Indexer 基于索引数据完成。此外,部分特性存在多阶段处理流程,并不一定走到 AST 层——例如代码补全在某些场景下可以在语法层(include 路径补全、module 导入补全)直接完成,document links 也可能在预处理阶段就能提取。`feature/` 中的实现仅对应"需要完整编译结果"的那个阶段。 - -## 服务器层 - -### `src/server/` - -语言服务器的核心运行时。这是最复杂的一层,负责将上述所有层组装成一个可运行的服务。 - -#### `src/server/protocol/` - -协议定义。描述主进程与工作进程之间、以及与客户端之间的通信消息格式。 - -- **Worker 协议**:定义有状态和无状态工作进程的请求/响应消息——编译请求、查询请求、构建请求、文档更新通知等。使用 bincode 序列化。 -- **扩展协议(Extension)**:LSP 标准之外的扩展请求,如编译上下文查询与切换。 -- **Agentic 协议**:面向 AI agent 的高级查询接口,提供编译命令查询、项目文件列表、文件依赖分析、影响分析、符号搜索、符号详情读取等功能。这些接口通过命令行直接访问,不需要 MCP 等中间层。 - -#### `src/server/workspace/` - -项目级全局状态——来自磁盘的唯一真相源。 - -- **工作区(Workspace)**:持有编译数据库、工具链、路径池、依赖图、PCH/PCM 缓存、项目索引等全部项目级状态。核心不变量:打开文件的未保存缓冲区内容不会修改 Workspace。Workspace 的状态变更来自三个路径:初始化加载、文件保存(didSave)触发的级联更新、以及后台索引完成后的索引合并。 -- **配置(Config)**:TOML/JSON 配置文件的加载与校验。 - -#### `src/server/compiler/` - -编译调度与索引管理。这一层是 `src/compile/`(底层编译抽象)和 `src/server/service/`(服务层)之间的桥梁。 - -- **编译器(Compiler)**:编译生命周期的调度器。协调 PCH 构建、模块依赖解析、AST 编译的先后顺序,将编译任务分发给工作进程。它本身不持有持久数据——所有数据都存在 Workspace 和 Session 中,Compiler 只负责编排执行流程。 -- **编译图(CompileGraph)**:C++20 模块编译的 DAG 调度器。基于引用计数实现兴趣追踪——当没有任何请求者关心某个模块单元时,自动取消其编译。支持依赖级联取消和文件更新触发的重编译。 -- **索引器(Indexer)**:后台索引的调度和跨文件查询的入口。管理索引队列、去重、空闲超时批处理。查询时综合三个数据源:ProjectIndex 用作目录(定位符号所在文件),MergedIndex 提供按文件分片的关系数据,打开文件的内存索引提供未保存的实时状态覆盖。 - -#### `src/server/service/` - -服务入口与会话管理。 - -- **主服务器(MasterServer)**:顶层协调器。持有 Workspace、Session 映射、WorkerPool、Compiler、Indexer,将 LSP 请求路由到对应的处理逻辑。管理服务器生命周期(Uninitialized → Initialized → Ready → ShuttingDown → Exited),启动文件监视任务。 -- **会话(Session)**:每个打开文件的编辑状态。持有当前缓冲区内容、编译版本号、ABA 防护的 generation 计数器、内存符号索引、PCH 引用、依赖快照等。Session 在 didOpen 时创建,didClose 时销毁。Session 的修改不影响其他文件——所有跨文件依赖指向磁盘文件。 -- **LSP 客户端(LSPClient)**:LSP 协议的请求处理器,注册各 LSP 方法的 handler。 -- **Agent 客户端(AgentClient)**:Agentic 协议的请求处理器。 - -#### `src/server/worker/` - -工作进程管理。 - -- **工作池(WorkerPool)**:管理工作进程的生命周期、路由和调度。 - - **有状态工作进程(StatefulWorker)**:在内存中持有 AST,服务查询请求(hover、semantic tokens 等)。每个打开的文件通过 path_id 亲和性绑定到一个有状态工作进程,新文件分配给当前负载最轻的工作进程。 - - **无状态工作进程(StatelessWorker)**:执行一次性任务(PCH/PCM 构建、补全、索引等)。采用优先级感知调度——交互式请求(补全、签名帮助)优先于后台任务(索引)。并发数量根据内存压力和工作进程崩溃情况动态调整。 -- **进程容错**:工作进程崩溃时自动重启(有最大重启次数限制),有状态工作进程崩溃时自动重新分配其拥有的文档。 +详见 [多进程架构](multi-process.md)。 ## 模块间关系 @@ -175,7 +119,7 @@ semantic(提取语义信息) ──→ index(构建索引) ↓ feature(产出 LSP 响应) -server 层将以上所有层通过 Workspace/Session/WorkerPool 组装成可运行的服务 +server 层将以上各层通过 Workspace / Session / WorkerPool 组装成可运行的服务 ``` -`support` 和 `syntax` 是横切层,被多个模块共享。`server/protocol` 定义了进程间和客户端间的消息契约。 +`support` 和 `syntax` 是横切层,被多个模块共享。 diff --git a/docs/zh/design/symbol-index.md b/docs/zh/design/symbol-index.md new file mode 100644 index 000000000..c2b5b3d5e --- /dev/null +++ b/docs/zh/design/symbol-index.md @@ -0,0 +1,246 @@ +# 符号索引 + +## 背景 + +语言服务器的许多功能需要跨文件工作。用户在一个文件中触发"跳转到定义",目标可能在项目的任意另一个文件里;"查找引用"需要扫描整个项目中所有可能提及该符号的文件;调用层次和类型层次更是涉及多个文件之间的符号关系链。要实现这些功能,语言服务器必须维护一份覆盖整个项目的符号索引,记录每个符号在哪些文件中出现、它们之间的语义关系(定义、引用、调用、继承等)。 + +C++ 让索引的构建和维护变得特别困难。 + +第一个层面是头文件的编译上下文问题。C++ 的 `#include` 是文本替换——头文件的内容在编译时被逐字插入到包含它的源文件中。这意味着同一个头文件在不同的编译上下文下可以产生完全不同的符号: + +```cpp +// crypto.h +#ifdef USE_OPENSSL + using TLSContext = OpenSSLContext; +#else + using TLSContext = BoringSSLContext; +#endif +``` + +当 `crypto.h` 被一个定义了 `USE_OPENSSL` 的源文件包含时,`TLSContext` 是 `OpenSSLContext`;被另一个源文件包含时,它是 `BoringSSLContext`。条件编译是最直观的例子,但 include 顺序和模板实例化也会导致头文件在不同上下文下产生不同的符号关系。如果索引只记录一个上下文的结果,用户切换上下文后就会看到错误的跳转目标或不完整的引用列表。 + +第二个层面是规模。一个中等规模的 C++ 项目(数千个源文件)编译后产生数十万个符号,大型项目(LLVM、Chromium)达到百万量级。索引系统需要在这个规模下保持合理的内存占用、构建时间和查询延迟。 + +clangd 在这两个层面的处理都有不足。在编译上下文方面,clangd 的后台索引对每个头文件只保留最后一次编译的结果——后编译的源文件会覆盖先前的索引数据。如果一个符号的引用只在某个特定的编译上下文下存在,而该上下文不是最后被索引的,这条引用就丢失了。 + +在跨文件查找方面,clangd 的后台索引将符号信息存储在符号声明所在文件对应的索引分片中。这种以声明文件为中心的存储方式导致引用计数无法正确跨文件累加(clangd [#23](https://github.com/clangd/clangd/issues/23))。用户也经常遇到"查找引用"结果不完整的情况——某些引用直到手动打开对应文件后才出现,因为动态索引此时才补上了缺失的数据(clangd [#516](https://github.com/clangd/clangd/issues/516)、[#802](https://github.com/clangd/clangd/issues/802))。此外,编译命令发生变化时,clangd 的过期检测不会触发重新索引(clangd [#199](https://github.com/clangd/clangd/issues/199)),导致索引数据与实际编译状态长期不一致。 + +clice 的索引系统针对这些问题做了重新设计:采用三级索引结构将全局符号目录与按文件分片的关系数据分离,通过内容寻址去重来合并不同编译上下文的索引结果,用 FlatBuffers 实现按需惰性加载来控制内存占用,并用打开文件的实时覆盖层保证编辑过程中查询结果的时效性。 + +## 设计 + +### 符号标识 + +索引系统需要一种跨文件、跨编译单元的方式来标识同一个符号。clice 使用 `SymbolHash`(一个 64 位整数)作为符号的唯一标识。 + +`SymbolHash` 基于 Clang 的 USR(Unified Symbol Resolution)生成。USR 是符号身份的规范字符串表示,编码了符号的完全限定名:命名空间、类名、函数签名、模板参数等信息。例如 `std::vector::push_back` 和 `std::vector::push_back` 产生不同的 USR。`SymbolHash` 是对 USR 字符串计算的哈希值。 + +`SymbolHash` 有两个关键特性。第一,跨文件一致:同一个符号无论在哪个文件中出现,`SymbolHash` 始终相同。文件 A 中的 `std::string` 和文件 B 中的 `std::string` 拥有相同的哈希值,通过它可以关联所有的定义和引用——这是跨文件导航的基础。第二,紧凑高效:64 位整数比可变长度的 USR 字符串更适合作为哈希表的键和序列化存储。 + +### 符号出现与符号关系 + +索引存储两种核心数据:符号出现(`Occurrence`)和符号关系(`Relation`)。 + +`Occurrence` 记录一个符号在源码中的出现位置,只包含源码范围和目标符号的 `SymbolHash`。它回答的问题是"光标位置下是什么符号"。 + +`Relation` 记录更丰富的语义信息,包含三个要素:关系类型(`RelationKind`)、源码位置、以及目标符号。关系类型涵盖了常见的符号间语义: + +- 定义与声明(Definition、Declaration) +- 引用(Reference、WeakReference) +- 继承(Base、Derived) +- 调用(Caller、Callee) +- 类型关系(Interface、Implementation、TypeDefinition) +- 构造与析构(Constructor、Destructor) + +两者分开存储是因为查询模式不同。`Occurrence` 按位置索引——给定一个字节偏移量,用二分查找快速定位光标下的符号。`Relation` 按 `SymbolHash` 索引——给定一个符号,查找它的所有定义、引用、调用关系。这两种查询对数据排序的要求相互矛盾,分开存储使两种查询都能高效执行。 + +### 三级索引层次 + +clice 的索引分为三个层次,每层有不同的生命周期和职责: + +``` +TUIndex 一次编译的原始产物,合并后丢弃 + ↓ 合并 +ProjectIndex 全局符号目录(符号出现在哪些文件中),常驻内存 +MergedIndex 按文件分片的关系数据(符号在文件中的具体位置和关系),按需加载 + ↑ 叠加 +FileIndex 打开文件的实时覆盖层(来自内存中的 AST) +``` + +**TUIndex** 是编译一个翻译单元时产生的原始索引数据。`SemanticVisitor` 遍历 AST,对每个符号出现和关系生成记录,按文件分组组装成 `TUIndex`。由于一次编译涉及主文件和所有被 include 的头文件,`TUIndex` 内部为每个涉及的文件各维护一份独立的 `FileIndex`。`TUIndex` 还包含一份 `SymbolTable`(符号哈希到名称和种类的映射)和 `IncludeGraph`(这次编译中的 include 关系)。`TUIndex` 是临时数据,合并到持久索引后即被丢弃。 + +**ProjectIndex** 是全局的符号目录。它汇聚所有已索引翻译单元的符号信息,维护一张全局符号表:`SymbolHash` → 符号名称、符号种类、引用文件位图。其中引用文件位图记录了该符号出现在哪些文件中,使用 Roaring Bitmap 压缩存储。 + +`ProjectIndex` 不存储符号的具体位置(偏移量、行号)。它的角色是"目录"——告诉你一个符号存在于哪些文件中,然后你去对应文件的 `MergedIndex` 分片中查找具体位置。这种分离使 `ProjectIndex` 保持紧凑,可以常驻内存。 + +**MergedIndex** 是按文件分片的索引存储层。项目中的每个文件对应一个 `MergedIndex` 分片,存储该文件中所有符号的出现位置和关系信息。这是索引系统中体积最大的部分,也是实际承载查询的层。它支持从磁盘惰性加载——未被查询的分片不会加载到内存中。 + +`MergedIndex` 的核心能力是将同一个文件在不同编译上下文下产生的索引数据合并去重存储,具体机制在实现部分详述。 + +**FileIndex**(打开文件的覆盖层)存在于每个打开文件的 `Session` 中,来自最近一次内存编译的结果。它和 `MergedIndex` 分片存储相同类型的数据(`Occurrence` 和 `Relation`),但不写入全局状态——只在查询时作为叠加层使用,用当前编辑缓冲区的编译结果覆盖磁盘索引的数据。 + +### 符号表 + +`SymbolTable` 将 `SymbolHash` 映射到符号的元信息——名称和种类(Class、Function、Variable 等)。它出现在两个位置:`ProjectIndex` 中的全局符号表和 `Session` 中的局部符号表。查询符号名称时先查 `Session`(内容更新),找不到再查 `ProjectIndex`。 + +### IncludeGraph + +`IncludeGraph` 记录一次编译中所有文件的 include 关系。它包含两部分:路径列表(该编译涉及的所有文件路径)和 `IncludeLocation` 记录(每条记录表示一个文件在某一行被 include,以及 include 的来源文件)。 + +`IncludeGraph` 有两个用途。在 `TUIndex` 合并时,它提供编译单元内部的文件 ID 到项目全局路径 ID 的映射。在 `MergedIndex` 中,它作为编译上下文的一部分被存储,include 链用于过期检测。 + +## 实现 + +### 索引构建 + +`TUIndex` 的构建由 `SemanticVisitor` 完成:给定一个编译单元,遍历 AST,为每个命名声明和宏生成 `Occurrence` 和 `Relation` 记录。遍历完成后,对每个文件的数据进行去重和排序——`Occurrence` 按位置排序以支持二分查找,`Relation` 按类型和位置排序以支持过滤。 + +构建过程中,主文件(源文件)的 `FileIndex` 被单独提取出来。这使得合并阶段可以区分处理——主文件作为源文件上下文合并,其余文件作为头文件上下文合并。 + +### 索引合并 + +`TUIndex` 合并到持久索引分两步进行。 + +第一步,将符号信息合并到 `ProjectIndex`。`TUIndex` 中的所有符号插入全局符号表,同时更新每个符号的引用文件位图——将本次编译涉及的文件加入位图。这一步同时完成路径映射:`TUIndex` 内部使用的路径 ID 转换为 `ProjectIndex` 的全局路径 ID。 + +第二步,将各文件的 `FileIndex` 合并到对应的 `MergedIndex` 分片。对于主文件,附带编译上下文信息(编译时间戳、include 链);对于头文件,附带头文件上下文信息(include 位置标识)。 + +### 编译上下文去重 + +`MergedIndex` 面对的核心问题是:同一个头文件被 N 个源文件包含,会产生 N 份 `FileIndex`。如果每份都完整存储,空间会随编译单元数量线性增长。但实际上,绝大多数头文件在不同编译上下文下产生的索引是完全一样的——相同的符号出现在相同的位置,产生相同的关系。只有像上面 `crypto.h` 那样受条件编译影响的头文件,才会在不同上下文下产生不同的索引内容。 + +`MergedIndex` 用内容寻址去重来解决这个问题。每份 `FileIndex` 在合并前计算 SHA-256 内容哈希。哈希相同的 `FileIndex` 内容一定相同,它们共享同一个 canonical ID(一个自增的整数标识)。 + +具体来说,`MergedIndex` 中的 `Occurrence` 和 `Relation` 不是简单的列表,而是每条记录关联一个 Roaring Bitmap,标记该记录属于哪些 canonical ID。合并一份新的 `FileIndex` 时: + +1. 计算其 SHA-256 哈希 +2. 查缓存:如果这个哈希已存在,说明数据完全相同,直接复用对应的 canonical ID,递增引用计数 +3. 如果是新哈希,分配新的 canonical ID,将所有 `Occurrence` 和 `Relation` 记录插入,并关联到这个新 ID + +当编译上下文被移除时(例如源文件从项目中删除),对应 canonical ID 的引用计数递减。引用计数归零的 canonical ID 被标记进"已移除"集合。查询时,属于已移除集合的数据会被过滤掉。 + +这种设计使得存储量取决于索引内容的种类数而非编译上下文的数量。对于大多数头文件,无论被多少源文件包含,只存储一份数据。 + +### 编译上下文类型 + +`MergedIndex` 内部区分两类编译上下文: + +- **CompilationContext**:文件作为源文件被直接编译时产生。记录编译时间戳和 include 链(用于过期检测),以及对应的 canonical ID。一个文件可以有多个 `CompilationContext`,对应编译数据库中不同的编译命令。 +- **HeaderContext**:文件作为头文件被其他源文件包含时产生。记录包含它的源文件和 include 位置,以及对应的 canonical ID。 + +这两类上下文配合[编译上下文](compilation-context.md)系统工作。查询时不需要区分上下文类型——所有上下文的数据已经通过 canonical ID 的 Bitmap 统一管理。过期检测时,`CompilationContext` 的 include 链用于判断是否需要重新索引。 + +### 惰性加载 + +`MergedIndex` 使用 FlatBuffers 序列化。FlatBuffers 的设计允许直接在序列化数据上执行查询,无需反序列化到内存结构。`MergedIndex` 利用这一特性实现两层访问模式: + +- **只读路径**:从磁盘加载后,`MergedIndex` 保持为原始的内存映射缓冲区,查询操作直接在 FlatBuffers 数据上执行,没有反序列化开销。 +- **读写路径**:需要修改时(合并新数据或移除旧上下文),先将 FlatBuffers 数据反序列化到内存结构,后续操作在内存结构上进行。修改后的分片保存时重新序列化。 + +启动时只需加载 `ProjectIndex`(体积较小),`MergedIndex` 分片按需加载,且大多数分片在一次会话中不会被访问。 + +### 查询流程 + +以"查找引用"为例,说明跨文件查询的完整流程: + +1. 在当前文件中,用光标的字节偏移量在 `Occurrence` 列表中二分查找,得到光标下符号的 `SymbolHash` +2. 在 `ProjectIndex` 中查找该 `SymbolHash` 的引用文件位图,得到所有包含该符号的文件列表 +3. 对列表中的每个文件分别查询: + - 如果该文件已打开(有活跃的 `Session`),使用 `Session` 的 `FileIndex`,跳过对应的 `MergedIndex` 分片 + - 如果该文件未打开,加载对应的 `MergedIndex` 分片并在其中查找 +4. 汇总所有文件中找到的 `Relation`(按目标 `RelationKind` 过滤),转换为 LSP 位置返回给客户端 + +第 3 步中,打开文件优先使用 `Session` 的 `FileIndex` 而非 `MergedIndex`,因为缓冲区内容可能与磁盘不一致。`Session` 的 `FileIndex` 来自内存中的编译结果,更准确地反映用户当前看到的代码。当 `Session` 的 AST 处于脏状态(用户编辑后尚未重新编译)时,才回退到 `MergedIndex`。 + +> 偏移量到 LSP 位置的转换需要文件内容和行首偏移表。`MergedIndex` 分片中存储了对应文件的内容和行首偏移表,因此即使文件未打开也能完成转换。 + +### 过期检测 + +过期检测决定一个文件是否需要重新索引。`MergedIndex` 分片中存储了编译时间戳和 include 链。检测时,遍历 include 链中每个文件的最后修改时间(mtime),如果任何一个文件的 mtime 晚于编译时间戳,说明依赖已更新,需要重新索引。 + +这种检测是保守的——mtime 变化不一定意味着内容变化(例如 `touch` 操作、分支切换)。但误判只会导致多执行一次索引,不会遗漏需要更新的文件。 + +### 后台索引调度 + +后台索引的调度需要平衡索引的及时性和对用户交互的干扰。索引模块采用以下策略: + +- **队列与空闲延迟**:需要索引的文件加入队列,在编辑器空闲一段时间后才开始处理。这避免了用户快速编辑时频繁触发索引任务。 +- **并发控制与内存监控**:同时运行的索引任务数有上限。索引过程中动态监控系统内存占用——内存紧张时自动降低并发数,内存恢复后逐步提升回基线值。 +- **优先级管理**:用户主动触发的操作(如编译打开的文件)会暂停后台索引。操作完成后恢复,确保用户请求的响应延迟不受后台索引影响。 +- **结果合并与持久化**:每个索引任务在无状态子进程中编译文件并构建 `TUIndex`,结果序列化后传回主进程,由主进程合并到 `ProjectIndex` 和 `MergedIndex` 中。索引完成后,修改过的分片被写回磁盘,下次启动时可以直接加载。 + +## FAQ + +- **为什么将 `ProjectIndex` 和 `MergedIndex` 分开,而不是用一个统一的索引?** + + 如果将位置信息也存入 `ProjectIndex`,它的体积会急剧膨胀,无法常驻内存。而如果没有 `ProjectIndex`,每次跨文件查询都需要遍历所有 `MergedIndex` 分片来定位符号所在的文件——在一个上万文件的项目中,加载上万个分片是不可接受的。`ProjectIndex` 作为轻量的目录层,先缩小搜索范围到具体的几个文件,再到对应分片中精确查找。 + +- **为什么打开文件的索引不写入全局状态?** + + 用户正在编辑的缓冲区内容可能是不完整的、有语法错误的代码。如果将这些临时状态写入全局索引,会污染其他文件的查询结果。例如,一个正在编辑的头文件中某个符号的定义临时消失,会导致所有引用该符号的文件的查找引用结果受到影响。全局索引只接受保存到磁盘的稳定状态,通过后台索引从磁盘文件构建。 + +- **内容寻址去重是否有哈希冲突的风险?** + + 理论上 SHA-256 存在冲突可能,但概率可以忽略(2^-128 量级)。实践中将 SHA-256 冲突视为"不会发生"是标准做法。即使发生冲突,后果也只是两份不同的 `FileIndex` 共享了数据,不会导致崩溃或数据损坏。 + +- **为什么使用 FlatBuffers 而不是 Protocol Buffers 或自定义格式?** + + FlatBuffers 允许直接在序列化数据上查询,不需要反序列化。对于 `MergedIndex` 这样可能有数千个分片的数据,大多数分片在一次会话中不会被访问。FlatBuffers 的零拷贝特性使得加载一个分片的开销接近于零——只需要内存映射文件,实际访问的数据才被读入内存。Protocol Buffers 需要完整的反序列化步骤,不适合这种按需加载模式。 + +- **为什么 `Occurrence` 和 `Relation` 要分开存储?** + + `Occurrence` 按位置索引——给定偏移量,二分查找定位光标下的符号,需要按位置排序。`Relation` 按 `SymbolHash` 索引——给定符号,查找它的所有关系,需要按符号分组。如果合并为一种数据结构,无法同时满足两种排序需求,必然在其中一种查询上损失效率。 + +- **为什么跨文件查询在主进程而非子进程中完成?** + + clice 采用多进程架构,打开的文件各自在有状态子进程中编译。跨文件查询(如查找引用)需要遍历所有打开文件的 `Session`,汇总它们的 `FileIndex` 结果。如果这些 Session 分散在不同的子进程中,每次查询都需要跨多个进程通信再聚合结果,延迟和复杂度都不可接受。因此,子进程编译完成后会将 `FileIndex` 传回主进程,由主进程统一完成查询——既可以访问所有打开文件的 `FileIndex`,也可以直接访问 `ProjectIndex` 和 `MergedIndex`,在一个进程内完成完整的查询流程。 + + 这个设计还有一个语义上的考虑:打开文件的查询结果应该反映编辑器中的缓冲区状态,而非磁盘状态。即使磁盘上的文件已被外部工具修改,只要编辑器没有发送 `didChange`,查询结果仍应基于编辑器持有的版本。将所有 `FileIndex` 集中在主进程中管理使得这一语义约束更容易维护。 + +- **为什么不用数据库存储索引?** + + clice 需要持久化多种缓存文件:索引分片、PCH、PCM 等。PCH 和 PCM 文件体积大(可达上百 MB)但数量少(大致与打开文件数或模块数成正比),生命周期也很简单——创建、读取、过期后删除,没有复杂的查询或事务需求。数据库擅长的能力(事务、索引、复杂查询)对这类文件没有意义,用文件系统管理足够了。 + + 索引分片是唯一可能从数据库中获益的部分:数量多(等于项目文件数)、体积小,可能受益于原子写入和自动 LRU 淘汰。但目前基于文件系统的方案已经能满足这些需求。为了索引分片一个场景引入数据库依赖,增加的复杂度是否值得,需要等实际遇到瓶颈再评估。对于大型项目(数万个文件),同一目录下存放大量索引分片可能带来文件系统层面的压力,未来可以考虑分层存储或引入轻量数据库来缓解。 + +## 已知局限 + +- **符号表的局部性**。目前所有符号(包括函数内的局部变量)都被合并到 `ProjectIndex` 的全局符号表中。这导致大量只在单个文件内部有意义的符号被全局存储,增加了合并时的哈希表插入开销和内存占用。 + + 改进方向是引入多级符号表——不只 `ProjectIndex` 有 `SymbolTable`,`MergedIndex` 分片也应该有自己的 `SymbolTable`。判断一个符号属于哪一级的规则是:符号定义在哪个文件,就属于哪个文件的 `SymbolTable`,前提是该符号是内部的(不会被其他文件引用)。例如: + + ```cpp + // utils.h + inline int helper(int x) { + auto temp = x * 2; // temp 是 utils.h 的局部符号 + return temp + 1; + } + ``` + + ```cpp + // main.cpp + #include "utils.h" + static int counter = 0; // counter 是 main.cpp 的局部符号 + + int main() { + counter = helper(42); + } + ``` + + `temp` 定义在 `utils.h` 中,不会被任何其他文件引用,它应该在 `utils.h` 对应的 `MergedIndex` 分片的 `SymbolTable` 中,而不是 `ProjectIndex` 的全局符号表中。`counter` 是 `main.cpp` 的静态变量,同理应该在 `main.cpp` 的 `MergedIndex` 分片中。注意 `temp` 虽然出现在头文件中,但它属于头文件的 `SymbolTable` 而非包含它的源文件的 `SymbolTable`,因为它定义在头文件中。只有 `helper`、`main` 这样可能被跨文件引用的符号才需要进入 `ProjectIndex`。 + + 这样做的目的是最小化 `ProjectIndex` 的体积和合并开销,同时避免同一个局部符号因出现在多个编译单元中而被重复存储。 + +- **过期检测的精度**。当前的过期检测只使用 mtime——只要依赖文件的 mtime 晚于编译时间戳就触发重新索引。这在 `touch`、分支切换、CI 还原等场景下会产生不必要的重新索引(文件 mtime 变了但内容没变)。改进方向是 mtime + 内容哈希双层检测:第一层用 mtime 快速判断,mtime 未变则跳过(零 I/O);第二层对 mtime 变化的文件计算内容哈希,哈希不变说明内容未改,同样跳过。这种方案已在编译产物(PCH、AST)的过期检测中使用,索引的过期检测应当对齐。 + +- **模糊符号搜索**。当前的全局符号搜索(workspace/symbol)是简单的子串匹配,对 `ProjectIndex` 中所有符号做线性扫描。在大型项目中效率不够,且不支持模糊匹配。 + + C++ 符号名有结构性:`getSymbolHash` 是 camelCase,`get_symbol_hash` 是 snake_case,`std::vector::push_back` 带命名空间限定。用户搜索时通常只输入缩写或片段(如 `symhash`、`gSH`、`vec_pb`),期望匹配到完整符号名。子串匹配无法处理这类查询。 + + 改进方向是为符号名建立专门的搜索索引。需要一个分词器将符号名按命名约定拆分为词元(`getSymbolHash` → `[get, Symbol, Hash]`,`push_back` → `[push, back]`),然后基于词元建立倒排索引。例如使用 trigram(三字符组)作为索引键,查询时取 trigram 交集得到候选集,再精确评分排序。clangd 的 Dex 索引采用了这种 trigram posting list 方案,是一个可参考的实现。另一个方向是引入成熟的全文搜索库,但需要评估引入外部依赖的代价。 + +- **PCH 导致的索引分裂**。使用 PCH(预编译头)优化时,一个文件的编译实际上被分成两个阶段:先编译 preamble 部分(文件顶部的 `#include` 指令)生成 PCH,再用 PCH 编译文件的其余部分。PCH 本身也是一个编译单元,会产生独立的索引数据。 + + 这种分裂对索引产生了影响。以 document links(编辑器中可点击的 `#include` 指令)为例:preamble 中的 `#include` 指令属于 PCH 编译阶段,主文件的编译看不到它们。而索引系统目前不存储 document links 信息,因此无法通过索引来补全 PCH 部分的结果。当前的做法是在 PCH 构建时将其 document links 预序列化为 JSON 存储在 PCH 元数据中,查询时手动拼接到主文件的结果中。这种做法可以工作但不够干净。 + + 更合理的方案是将 document links 纳入 PCH 的元数据体系(目前 PCH 元数据已经存储了依赖文件列表等信息),或者利用索引中已有的 include 关系信息来重建 document links。后者的问题在于索引构建需要时间,而 PCH 编译完成后应尽快投入使用,等待索引完成会增加延迟。目前两种方案都尚未完整实现。 diff --git a/docs/zh/design/template-resolver.md b/docs/zh/design/template-resolver.md index d0c996ebe..dacac2ea9 100644 --- a/docs/zh/design/template-resolver.md +++ b/docs/zh/design/template-resolver.md @@ -1,149 +1,234 @@ -# 模板解析 +# 模板解析器 ## 背景 -C++ 模板的核心设计是**延迟实例化**——模板代码在被具体类型参数实例化之前,编译器不会(也无法)解析依赖于模板参数的名称。这些名称被称为**依赖名称(dependent names)**,它们的类型在模板定义阶段是未知的。 +在 C++ 中,模板代码的类型信息要到实例化时才会确定。编译器在解析模板定义时,会将依赖于模板参数的名称标记为"依赖名称(dependent name)",推迟到实例化阶段再做解析。这是 C++ 两阶段名称查找(two-phase name lookup)的核心规则。 -对于语言服务器来说,这意味着模板代码内部是一个"盲区": +对于语言服务器来说,这意味着模板函数体内的大量名称是不可用的。考虑以下代码: ```cpp template -void foo(std::vector vec) { - vec. // 光标在这里——vec 的类型已知是 vector,可以提供补全 +void process(std::vector& vec) { + auto it = vec.begin(); + it-> // 这里应该补全什么? } ``` -对于简单的情况,假设使用主模板的定义即可提供基本的补全。但考虑更复杂的场景: +`it` 的类型是 `std::vector::iterator`,这是一个依赖名称。编译器只知道它依赖于 `T`,但不知道它具体是什么类型。在 AST 中,它被表示为一个 `DependentNameType` 节点,语言服务器看到的只是"某个和 `T` 有关的类型"——无法提供补全、跳转定义、悬停信息等任何语言功能。 + +问题在简单场景下还可以接受。`vec.push_back()` 这样的调用,语言服务器可以通过查看 `vector` 主模板的成员来提供基本的补全。但实际的 C++ 代码,尤其是涉及标准库的代码,远比这复杂: ```cpp template -void foo(std::vector> vec2) { - vec2[0]. // vec2[0] 的类型是什么? +void process(std::vector>& vec) { + vec[0].push_back(T{}); // vec[0] 的类型是什么? } ``` -`vec2[0]` 的类型是 `std::vector>::reference`——一个依赖名称。在标准库实现中,解析它需要追踪数十层嵌套的模板 typedef 链:`reference` → `allocator_traits::value_type` → `__alloc_traits::reference` → ... 每一层都涉及偏特化匹配、默认模板参数、typedef 展开等复杂操作。 +`vec[0]` 返回 `std::vector>::reference`。在标准库的实现中,要知道这个类型实际上是 `std::vector&`,需要追踪一条很长的 typedef 链:`reference` → `allocator_traits::value_type` → `__alloc_traits::reference` → ...每一步都涉及偏特化匹配、默认模板参数展开、typedef 解析等操作。 -在 clangd 的社区中,模板代码的补全和悬停问题长期存在。用户反复报告在模板函数体内无法获得补全建议——光标处显示 ``,所有 LSP 功能失效。问题的根源在于 clangd 的解析策略存在几个根本性限制: +clangd 在 Clang 中实现了一个 `HeuristicResolver` 来处理依赖名称(该类最初在 clangd 内部,后来上游到了 Clang 库中)。它的做法是:遇到依赖名称时,尝试在主模板的定义中查找对应的成员。这个策略能覆盖一部分简单场景,但存在根本性的限制: -**不处理偏特化**:clangd 假设总是使用主模板的定义来查找成员。但标准库大量使用偏特化(如 `allocator_traits` 对不同分配器的特化),主模板可能根本没有要查找的成员——成员只存在于特定的偏特化中。 +- **只查找主模板**。标准库大量使用偏特化来定义成员类型。例如 `allocator_traits` 的成员类型定义在对 `allocator` 的偏特化中,主模板中根本没有这些成员。`HeuristicResolver` 在这些场景下完全失效。 -**缺少参数映射**:即使通过名称查找找到了成员的类型,这个类型仍然是用被查找模板的参数表达的(如 `allocator_traits::value_type` 中的 `Alloc`),而不是用调用方的参数(如 `T`)。clangd 不做实例化,因此无法建立参数之间的映射关系。 +- **不建立参数映射**。即使在主模板中找到了一个 typedef `using value_type = T`,这里的 `T` 是被查找模板的参数,不是调用方的参数。在 `vector::value_type` 这样的场景下,需要知道 `T = MyType`,但 `HeuristicResolver` 不做参数推导,无法建立这种映射。 -**忽略默认模板参数**:`std::vector` 实际上是 `std::vector>`——第二个参数是默认的。clangd 不展开默认参数,导致依赖于默认参数的名称无法解析。 +- **不处理默认模板参数**。`std::vector` 实际上是 `std::vector>`。第二个参数是默认的,但 `HeuristicResolver` 不展开默认参数,导致依赖于分配器类型的名称无法解析。 -## 设计方案 +clangd 的 issue 中有大量与此相关的用户报告。[#1671](https://github.com/clangd/clangd/issues/1671) 反映了通过 typedef 引用的类型无法解析:用户定义了 `MetaWaldo::Type = Waldo`,对 `Type` 类型的变量调用 `find()` 时,go-to-definition 无法找到目标。[#307](https://github.com/clangd/clangd/issues/307) 反映了依赖基类的成员查找不工作:从 `Base` 继承的成员,在派生类模板中使用时无法跳转。这些问题的共同根源是 `HeuristicResolver` 的解析深度不足——它只做一层浅层查找,不追踪 typedef 链,不匹配偏特化,不穿透基类。 -### 核心思想 +clice 重新设计了依赖名称的解析机制。clice 的 TemplateResolver 实现了一个完整的伪实例化(pseudo-instantiation)引擎,能够在没有具体类型参数的情况下,通过模板参数推导和 typedef 展开,将依赖类型解析为可用的形式。 -clice 实现了一个**伪实例化器(PseudoInstantiator)**——它在没有具体类型参数的情况下,通过启发式方法解析依赖名称。核心思路是:不需要知道 `T` 是 `int` 还是 `string`,只需要追踪 `T` 在模板 typedef 链中的传播路径,将最终结果用 `T` 表达出来。 +## 设计 -例如:`std::vector>::reference` 经过伪实例化后被简化为 `std::vector&`——这是一个具体到足以提供补全的类型,同时保留了模板参数 `T` 的符号含义。 +### 伪实例化 -### 两阶段变换 +TemplateResolver 的核心思想是伪实例化:不需要知道模板参数 `T` 的具体类型,只需要追踪 `T` 在 typedef 链中的传播路径,将最终结果用 `T` 表达出来。 -解析器基于 Clang 的 TreeTransform 基础设施构建,分为两个阶段: +以 `std::vector>::reference` 为例。真正的模板实例化需要将 `T` 替换为具体类型(如 `int`),然后逐步展开所有 typedef。伪实例化的做法是:将 `T` 本身作为"已知的"参数,追踪它在标准库 typedef 链中的传播,最终得到 `std::vector&`。这个结果已经足够语言服务器提供补全和类型信息——用户看到的是一个有意义的类型,而不是 ``。 -**第一阶段:PseudoInstantiator(启发式解析)** +TemplateResolver 对外提供三类操作: -这是主引擎,负责解析各种依赖名称: +- **类型解析**:接受一个依赖类型,返回解析后的类型。例如将 `typename vector::value_type` 解析为 `T`。 +- **名称查找**:在嵌套名称限定符(nested name specifier)中查找声明。例如在 `vector::` 中查找 `iterator`,返回对应的声明。 +- **类型重糖化(resugar)**:将 Clang 内部的规范化模板参数类型(用深度+索引表示)映射回带有参数名称的类型,使结果对用户友好。 -- **DependentNameType**(如 `typename Container::value_type`):在模板的成员中查找 `value_type`,匹配偏特化,展开 typedef,将结果用调用方的参数重新表达 -- **DependentTemplateSpecializationType**(如 `Alloc::rebind`):解析依赖的模板特化,会优先尝试标准库特殊路径 -- **TemplateTypeParmType**(如 `T`):通过实例化栈查找参数的绑定值或默认参数 -- **DecltypeType**(如 `decltype(var)`):对简单变量引用解析其声明类型 +### 两阶段变换 -**第二阶段:SubstituteOnly(循环打破)** +解析依赖名称时,查找和替换两个操作会相互交织。查找一个 typedef 成员后,需要展开它的底层类型;展开后的类型中可能包含新的依赖名称,需要继续查找。但这种递归关系中隐藏着循环的可能:一个 typedef 的底层类型可能通过若干中间步骤又引用到自身。 -当第一阶段在展开 typedef 时,可能再次遇到需要启发式查找的依赖名称,形成循环:typedef A 的底层类型引用了依赖名称 B,查找 B 发现它的类型又涉及 typedef A。 +TemplateResolver 用两阶段设计来打破这种循环: -SubstituteOnly 打破这种循环——它只做参数替换和 typedef 展开,不执行启发式查找。当第一阶段需要展开 typedef 时,委托给 SubstituteOnly 完成,确保不会触发递归的启发式查找。 +**第一阶段**负责启发式查找。它处理各种依赖名称节点——`DependentNameType`(如 `typename Container::value_type`)、`DependentTemplateSpecializationType`(如 `Alloc::rebind`)、`TemplateTypeParmType`(模板参数本身),对它们执行成员查找、偏特化匹配、参数推导等操作。 -### 实例化栈 +**第二阶段**只做参数替换和 typedef 展开,不执行任何启发式查找。当第一阶段需要展开一个 typedef 的底层类型时,它委托给第二阶段完成。第二阶段遇到依赖名称时不会尝试查找——只是替换其中已知的模板参数,然后原样返回。 -模板参数在嵌套模板中可能处于不同的深度层级。实例化栈(InstantiationStack)维护一个参数映射栈,每一帧记录一层模板的参数绑定。 +这样做的效果是:typedef 展开永远不会触发新的启发式查找。启发式查找可能产生需要展开的 typedef,但展开过程不会反过来产生新的查找请求。循环链条被切断。 -当遇到 `TemplateTypeParmType`(一个用深度和索引标识的模板参数)时,解析器从栈顶(最内层)到栈底(最外层)线性搜索,匹配对应深度的参数绑定。如果找到,替换为绑定的类型;如果未找到,尝试使用参数的默认值。当栈为空时(独立的依赖类型),解析器会沿着外围模板声明向上走,推入它们的注入模板参数作为上下文。 +### 实例化栈 -这使得解析器能够处理多层嵌套模板: +模板参数用"深度+索引"来标识——深度对应模板的嵌套层级,索引是同一层中参数的位置。在多层嵌套模板中,不同层级的参数需要同时追踪: ```cpp -template +template struct Outer { - template + template struct Inner { - typename std::pair::first_type member; - // X 在深度 0,Y 在深度 1 - // 两者都需要在栈中追踪才能正确解析 + using type = std::pair; + // X: 深度 0, 索引 0 + // Y: 深度 1, 索引 0 }; }; ``` -### 依赖名称的解析流程 +TemplateResolver 维护一个实例化栈来管理参数绑定。每当解析器进入一层模板的上下文(通过参数推导确定了该层的参数绑定),就向栈中推入一帧。查找参数时,从栈顶向栈底搜索,匹配深度相同的帧。 + +当栈为空时(没有外围模板上下文),解析器会沿声明的外围模板上下文向上遍历,将每一层的注入模板参数(injected template arguments)推入栈中,构建出初始上下文。 + +### 缓存与降级 + +同一个编译单元中,相同的依赖类型节点可能出现在多个位置。TemplateResolver 在编译单元的生命周期内维护一个解析结果缓存,以 AST 节点指针为键。同一个节点在同一个编译单元中不会被重复解析。 + +解析过程中的任何步骤都可能失败——查找不到成员、循环检测触发、递归深度超限。TemplateResolver 在所有失败路径上返回原始的依赖类型,不抛异常也不报错。这意味着解析失败的最坏情况等同于不做解析——LSP 功能退化到显示 ``,但不会崩溃或返回错误的结果。 -以 `typename A::type` 为例: +## 实现 -1. **缓存检查**:如果该节点已解析过,直接返回缓存结果 -2. **循环检测**:如果该节点正在解析中,中止以防止无限递归 -3. **限定符变换**:变换 `A` 部分——替换参数、匹配偏特化 -4. **成员查找**:在变换后的类型中查找 `type`。依次尝试每个偏特化及其基类,然后是主模板及其基类 -5. **参数推导**:使用 Clang 的模板参数推导机制,建立形参到实参的映射 -6. **替换**:通过 SubstituteOnly 展开找到的成员类型中的 typedef,用调用方的参数替换 -7. **递归**:如果结果仍包含依赖名称,递归解析 +### TreeTransform 基础设施 -整个过程有递归深度限制(16 层),防止极端情况下的无限递归。 +Clang 提供了 TreeTransform 机制:通过继承 `TreeTransform` 基类并覆盖特定类型节点的变换方法,可以递归地变换整个类型树。TemplateResolver 的两个阶段各对应一个 `TreeTransform` 子类——第一阶段的 PseudoInstantiator 覆盖了 `DependentNameType`、`DependentTemplateSpecializationType`、`TemplateTypeParmType`、`TypedefType`、`DecltypeType` 等节点的变换方法;第二阶段的 SubstituteOnly 只覆盖 `TemplateTypeParmType`(参数替换)和 `TypedefType`(typedef 展开),不覆盖 `DependentNameType`(不做查找)。 -解析器还维护了两层循环检测:一层防止对同一个 DependentNameType 节点的重入解析,另一层防止在同一个 ClassTemplateDecl 上的递归查找(处理 CRTP 等自引用模式)。 +> 这里的关键约束是:SubstituteOnly 不覆盖 `TransformDependentNameType`。基类的默认行为是替换限定符中的参数后原样重建 `DependentNameType`,不做任何查找。这正是打破循环所需要的行为。 + +### DependentNameType 的解析流程 + +以 `typename vector::value_type` 的解析为例,流程如下: + +1. 检查缓存,如果命中直接返回 +2. 检查是否正在解析同一个节点(循环检测),如果是则返回原始类型 +3. 变换限定符部分 `vector`——替换其中的模板参数,必要时递归解析 +4. 在变换后的类型中查找 `value_type`: + - 提取出类模板声明 + - 依次尝试每个偏特化:推导模板参数,如果匹配成功,在偏特化及其基类中查找成员 + - 如果偏特化都没有匹配,尝试主模板及其基类 +5. 找到成员声明后,提取其底层类型(如 typedef 的底层类型或 record 类型) +6. 通过第二阶段展开底层类型中的 typedef,用当前栈中的参数绑定进行替换 +7. 弹出查找过程中推入的栈帧——替换已经完成,后续处理应该在调用方的上下文中进行 +8. 如果替换结果仍包含依赖名称,递归调用第一阶段继续解析 +9. 将最终结果写入缓存并返回 + +整个过程有 16 层的递归深度限制。 ### 偏特化匹配 -偏特化匹配是伪实例化器相对于 clangd 的关键优势。标准库大量使用偏特化,例如: +当需要在一个类模板中查找成员时,解析器利用 Clang 的模板参数推导机制来匹配偏特化。具体来说,它将可见的模板参数传给 Clang 的 `DeduceTemplateArguments`,尝试将它们与每个偏特化的模式进行匹配。如果推导成功,就在该偏特化的声明上下文中查找目标成员。 + +以一个简化的例子说明: ```cpp -// 主模板——没有定义 value_type -template struct allocator_traits; +template +struct alloc_traits; // 主模板:没有定义 value_type -// 偏特化——定义了 value_type -template struct allocator_traits> { - using value_type = T; +template +struct alloc_traits> { + using value_type = T; // 只在偏特化中定义 }; ``` -当解析 `allocator_traits>::value_type` 时,clangd 在主模板中查找 `value_type` 失败。伪实例化器会尝试将实际参数与每个偏特化的模式进行匹配,找到正确的偏特化后在其中查找成员。 +解析 `alloc_traits>::value_type` 时: + +- 在主模板中查找 `value_type`——不存在 +- 尝试偏特化 `alloc_traits>`,推导出 `T = int`——匹配成功 +- 在偏特化中找到 `value_type = T`,替换 `T` 为 `int`,返回 `int` + +对于符号化的参数(如 `alloc_traits>`,其中 `U` 是调用方的模板参数),同样的推导机制可以建立 `T = U` 的映射,最终返回 `U`。 -### 标准库特殊处理 +### 默认参数处理 -标准库的分配器重绑定链是一个特别深的 typedef 链: +在推导模板参数时,如果提供的参数数量少于模板声明的参数数量,解析器会尝试填充默认参数。默认参数本身可能依赖于已提供的参数(如 `vector` 的默认分配器 `allocator` 依赖于 `T`),因此默认参数的展开也需要通过第二阶段的替换来完成。 + +### 基类查找 + +如果在类模板自身的成员中没有找到目标名称,解析器会继续在其基类中查找。每个基类类型先通过第二阶段替换模板参数,然后递归进入查找流程。这使得通过继承获得的成员也可以被解析——例如 CRTP 模式中,派生类通过基类访问成员类型。 + +解析器对同一个(类模板声明, 目标名称)组合进行循环检测,防止自引用继承链(如 CRTP 中 `callback_traits : callback_traits`)导致无限递归。 + +### 标准库特殊路径 + +标准库的分配器重绑定链是一个特殊的深层 typedef 链。以 `vector` 为例,查找其 `reference` 成员需要经过: ```text -vector → __alloc_traits> → allocator_traits> - → allocator_traits::rebind_alloc +vector> + → __alloc_traits> + → allocator_traits>::rebind_alloc + → allocator::rebind::other (C++17 前) + 或直接替换第一个模板参数 (C++20 后) ``` -这个链条可能超过深度限制。解析器对 `allocator_traits::rebind_alloc` 模式进行特殊短路处理:检测到该模式后,直接尝试 `Alloc::rebind::other`(标准分配器协议),或回退到替换第一个模板参数。 +这个链条的每一步都涉及 `DependentTemplateSpecializationType` 的解析,深度容易超过限制。解析器对 `std::allocator_traits::rebind_alloc` 这一特定模式进行短路处理:检测到该模式后,先尝试标准的分配器重绑定协议(`Alloc::rebind::other`),如果失败(如 C++20 移除了 `allocator::rebind`),则回退到直接替换分配器的第一个模板参数。 + +### ResugarOnly + +Clang 内部用规范化的 `TemplateTypeParmType` 表示模板参数——它只包含深度和索引(如 "深度 0 的第 0 个参数"),不包含参数名称。这在 Clang 内部是充分的,但用户需要看到 `T`,而不是一个匿名参数。 + +ResugarOnly 是一个独立的 TreeTransform,它沿着声明的外围模板上下文向上遍历,收集每一层的模板参数列表,然后将规范化的参数类型映射回对应的命名声明。这个变换在 hover 等需要展示类型信息的场景中使用。 + +## FAQ + +- **为什么不直接使用 Clang 的模板实例化机制?** + + Clang 的 `TemplateInstantiator` 需要完整的实参列表才能工作——它要求将每个模板参数替换为具体类型。在模板定义阶段,这些信息不存在。即使人为选择一个占位类型(如 `int`),也会导致错误的结果:`vector::value_type` 是 `int`,但用户需要看到的是 `T`。伪实例化的设计选择是保留符号参数,让结果中的 `T` 保持其原始含义。此外,编辑器中的代码经常是不完整的或有语法错误的,完整的模板实例化在这些场景下会直接失败。 + +- **TemplateResolver 和 Clang 的 HeuristicResolver 是什么关系?** + + 两者解决的是同一类问题,但深度不同。`HeuristicResolver` 做浅层查找——在主模板中直接查找成员名称,不处理偏特化、不做参数推导、不展开 typedef 链。它的优点是简单且不容易出错。TemplateResolver 做深层的伪实例化——追踪参数传播、匹配偏特化、展开 typedef 链,能覆盖更多场景,但复杂度也高得多。目前 clice 在部分功能中(如 hover 的声明目标查找)仍使用 `HeuristicResolver`,因为那部分代码来源于 clangd 的 find_target 逻辑。TemplateResolver 的集成正在逐步推进。 + +- **为什么不用单阶段设计?** + + 考虑这样的场景:类型 `A::type` 是一个 typedef,底层类型是 `B::value_type`,而 `B::value_type` 又是通过另一个 typedef 间接引用了 `A::type`。单阶段设计中,解析 `A::type` 会展开其 typedef,遇到 `B::value_type`,触发查找,查找结果又需要展开,再次遇到 `A::type`——形成循环。两阶段设计中,typedef 展开由第二阶段完成,第二阶段不做查找,因此 `B::value_type` 在展开过程中不会被解析,循环不会发生。代价是第二阶段可能返回不够精确的结果(某些依赖名称没有被解析),但这比无限递归好得多。 + +- **为什么要对标准库模式做特殊处理?** + + 标准库的分配器重绑定链是实际代码中最常遇到的深层 typedef 链。通用的递归解析在处理这个链条时容易超过深度限制,因为每一步都涉及 `DependentTemplateSpecializationType` 的解析和偏特化匹配。特殊处理直接短路了这个常见模式,代价是代码中多了一个 ad-hoc 的分支。未来的方向是改进通用解析能力或引入一个可扩展的模式匹配机制来替代硬编码。 -### 类型重糖化 +## 已知局限 -Clang 内部用规范化的 `TemplateTypeParmType`(深度+索引)表示模板参数。用户看到的应该是参数名(如 `T`),而不是 `parameter 0 of depth 0`。ResugarOnly 变换将规范化的参数类型映射回其原始声明,用于对用户展示友好的类型信息。 +- **多元素包展开**:目前只处理单元素参数包(常见的包转发场景),不支持多元素包的展开。 -### 优雅降级 + ```cpp + template + struct Tuple { + using first = typename std::tuple_element<0, std::tuple>::type; + // Ts... = {int, float, double} 时无法正确展开 + }; + ``` -解析过程中的任何步骤失败(查找失败、循环检测触发、深度超限),解析器返回原始的依赖类型,而不是报错。这意味着 LSP 功能不会因为解析失败而完全中断——只是退化到"无法解析"的状态,与不做伪实例化的效果相同。 +- **非类型模板参数的默认值**:只处理类型模板参数(`typename T`)的默认值,不处理非类型模板参数(`int N`)和模板模板参数的默认值。 -## 设计决策与权衡 + ```cpp + template + struct Array { + using type = T; // 当 N 有默认值时,参数检查可能失败 + }; + ``` -**为什么是伪实例化而不是真正的模板实例化?** 真正的实例化需要具体的类型参数(如 `T = int`),但在模板定义阶段没有这些信息。而且语言服务器中的代码经常是不完整的,无法进行完整的实例化。伪实例化通过保留符号参数(`T` 本身),避免了对具体类型的依赖。 +- **依赖成员表达式**:`expr.template member()` 形式的依赖成员访问尚未实现。 -**为什么需要两阶段设计?** 单阶段设计中,typedef 展开会触发新的启发式查找,而查找结果可能又需要 typedef 展开——形成循环。两阶段设计通过职责分离打破循环:启发式查找只在第一阶段执行,typedef 展开委托给不做查找的第二阶段。 + ```cpp + template + void foo(T& obj) { + obj.template bar(); // 无法解析 bar 的声明 + } + ``` -**为什么要特殊处理标准库模式?** 标准库的分配器重绑定是实际代码中最常见的深层 typedef 链。通用的递归解析在这里会超过深度限制。特殊处理虽然不优雅,但覆盖了最高频的用户场景。未来计划用通用机制替代这些特殊处理。 +- **复杂的 decltype 表达式**:只处理 `decltype(var)` 形式的简单变量引用,不支持成员访问、函数调用等更复杂的表达式。 -**为什么选择优雅降级而不是报错?** 模板解析本质上是启发式的——不可能覆盖所有 C++ 模板的边界情况。优雅降级确保解析失败不会比"不做解析"更差,同时在能解析的情况下提供显著的体验提升。 +- **依赖上下文中的运算符**:运算符在 Clang 的 AST 中没有 `IdentifierInfo`,无法通过名称查找来解析。 -## 已知限制 + ```cpp + template + auto add(T a, T b) { + return a + b; // 如果 T 重载了 operator+,无法解析到该声明 + } + ``` -- **包展开(pack expansion)**:目前只处理单元素包(如包转发),不支持多元素包(如 `Us... = {int, float}` 的展开)。 -- **非类型模板参数的默认值**:只支持类型模板参数(TemplateTypeParmDecl)的默认值替换,不支持非类型参数(如 `template`)和模板模板参数的默认值。 -- **依赖成员表达式**:`x.template foo()` 形式的依赖成员访问尚未实现。 -- **复杂的 decltype**:只处理 `decltype(var)` 的简单情况,不支持成员访问、函数调用等复杂表达式。 -- **运算符查找**:依赖上下文中的运算符(如 `a + b`,其中 `a` 类型依赖于模板参数)无法解析,因为运算符缺少标识符信息。 +- **表达式级别的解析**:目前 TemplateResolver 主要处理类型级别的依赖名称(`DependentNameType` 等)。表达式级别的解析(如 `CXXUnresolvedConstructExpr`、`UnresolvedLookupExpr`)已声明接口但尚未实现。此外,TemplateResolver 尚未完全集成到所有 LSP 功能中——部分功能(如 inlay hints 的依赖调用、语义标记的依赖名称着色)仍在等待接入。 diff --git a/docs/zh/index.md b/docs/zh/index.md index 1d074b809..47157b75f 100644 --- a/docs/zh/index.md +++ b/docs/zh/index.md @@ -20,16 +20,16 @@ hero: alt: clice features: - - icon: T - title: 更好的模板处理 - details: 使用伪实例化器处理依赖的模板名,对于复杂模板也可以有代码补全 - - icon: H - title: 头文件上下文 - details: 支持头文件在不同源文件的上下文之间进行状态切换,当然也完全支持非自包含文件 - - icon: M - title: 模块 - details: 良好的 C++20 模块支持,从代码补全到高亮到跳转全都适配 - - icon: I - title: 内存占用更低,速度更快 - details: 优秀的异步任务调度,支持编译任务取消,缓存必要的信息,避免无意义的 CPU 浪费 + - icon: 📝 + title: 编译上下文 + details: 第一个将编译上下文作为正式概念的语言服务器。用户可以查询和切换编译上下文,支持非自包含头文件和多配置项目 + - icon: 📦 + title: C++20 模块 + details: 基于引用计数的实时模块编译 DAG,支持取消和依赖级联。代码补全、语义高亮、跳转定义全面适配模块语法 + - icon: 🔍 + title: 模板解析 + details: 通过伪实例化技术解析依赖名称,在模板定义内部也能提供准确的代码补全和跳转 + - icon: ⚡ + title: 多进程架构 + details: Master + Worker 进程模型,隔离 Clang 的崩溃和内存泄漏。支持优先级调度、实时内存监控和进程自动恢复 --- diff --git a/docs/zh/sidebar.yaml b/docs/zh/sidebar.yaml index 66b5288f5..7410a64cf 100644 --- a/docs/zh/sidebar.yaml +++ b/docs/zh/sidebar.yaml @@ -28,17 +28,17 @@ design: items: - overview - compilation-context - - command - - index-design + - command-resolve + - symbol-index - multi-process - - module - - incremental + - module-graph + - incremental-parse - template-resolver - dependency-scanning dev: label: 开发 - collapsed: true + collapsed: false items: - build - contribution