Repository navigation
feat(bench): benchmark harness with per-TU profiler and perf logs #605
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
10 commits
Select commit
Hold shift + click to select a range
76e17d9
feat(bench): benchmark harness, per-TU profiler, perf instrumentation
16bit-ykiko fac1cd0
fix(bench): review fixes - verify_pch bound, worker-log breakdown, pe…
16bit-ykiko bc1b178
docs(bench): single-file clangd --check comparison recipe
16bit-ykiko a8167de
feat(index): index_detail perf topic - build/serialize stage splits
16bit-ykiko 96fcd97
fix(bench): resolve review threads - probe position, perf windows, ch…
16bit-ykiko 6c2df40
refactor(bench): self-review fixes - shared helpers, dead code, perf …
16bit-ykiko 88b38ef
fix(bench): second review round - stage mirroring, exact perf window,…
16bit-ykiko 16f984e
fix(bench): prefix-safe AST-load sources, scope/rel series discrimina…
16bit-ykiko c2bdaab
fix(bench): verify against full chain content, pin CDB, decouple time…
16bit-ykiko 3c45fdc
fix(bench): pin logging_dir overlay, size workers by available cpus
16bit-ykiko File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,148 @@ | ||
| # Benchmarks | ||
|
|
||
| Performance work on clice runs on three layers. Pick the layer that answers | ||
| your question: | ||
|
|
||
| 1. **Instrumentation** — _where does the time go in a real run?_ Every | ||
| server run emits `[perf:<topic>] key=value` log lines (see | ||
| `src/support/logging.h`) covering startup phases, per-file compiles, | ||
| PCH/PCM/index builds, cache hits, request latencies and index queries. | ||
| `tools/bench/perf_report.ts` aggregates any log into per-series | ||
| percentiles and can export a Chrome trace. This layer needs no special | ||
| build and works on logs users attach to issue reports. | ||
|
|
||
| 2. **Scenario harness** — _what does the user experience end to end?_ | ||
| `tools/bench/bench.ts` drives a server over LSP through fixed scenarios | ||
| (cold start, warm start, edit loop, warm feature requests) on a real | ||
| workspace and reports client-observed percentiles. It is server-agnostic: | ||
| point it at clangd with `--server clangd` to A/B the same scenario. | ||
|
|
||
| 3. **Component benchmarks** — _which design alternative is faster?_ | ||
| Standalone binaries in this directory, built with | ||
| `-DCLICE_ENABLE_BENCHMARK=ON`, each answering one decision: | ||
| - `scan_benchmark` — dependency-graph scan over a real CDB. | ||
| - `pipeline_benchmark` — per-TU stage profile (preprocess with/without | ||
| TokenBuffer, parse, index build/serialize, preamble PCH build incl. | ||
| preamble indexing, reparse over PCH incl. interactive indexing), one | ||
| result per file. | ||
| - `pch_chain_benchmark` — monolithic vs chained PCH strategy (ported | ||
| from PR #405). | ||
|
|
||
| ## Building | ||
|
|
||
| ```bash | ||
| cmake -B build/RelWithDebInfo -G Ninja -DCMAKE_BUILD_TYPE=RelWithDebInfo \ | ||
| -DCMAKE_TOOLCHAIN_FILE=cmake/toolchain.cmake -DCLICE_ENABLE_BENCHMARK=ON | ||
| ninja -C build/RelWithDebInfo scan_benchmark pipeline_benchmark pch_chain_benchmark | ||
| ``` | ||
|
|
||
| Always benchmark `RelWithDebInfo`; Debug numbers are meaningless. | ||
|
|
||
| ## Workloads | ||
|
|
||
| Benchmarks against real projects must be pinned to be comparable. | ||
| `workloads.json` records project + ref + configure command; | ||
| `fetch_workload.py <name>` materializes one under `benchmarks/workloads/` | ||
| (shallow clone + CMake configure — no build needed, only the | ||
| compile_commands.json): | ||
|
|
||
| ```bash | ||
| python benchmarks/fetch_workload.py llvm | ||
| ``` | ||
|
|
||
| clice's own CDB (`build/RelWithDebInfo/compile_commands.json`) doubles as | ||
| an always-available medium workload. | ||
|
|
||
| ## Typical sessions | ||
|
|
||
| Stage profile of the 100 largest TUs plus a Chrome trace of one: | ||
|
|
||
| ```bash | ||
| ./build/RelWithDebInfo/bin/pipeline_benchmark --limit 100 --json /tmp/pipeline.json \ | ||
| benchmarks/workloads/llvm/build/compile_commands.json | ||
| ./build/RelWithDebInfo/bin/pipeline_benchmark --filter SemaExpr.cpp --runs 3 \ | ||
| --time-trace /tmp/traces benchmarks/workloads/llvm/build/compile_commands.json | ||
| ``` | ||
|
|
||
| `--time-trace` writes clang's own `-ftime-trace` profile of the parse | ||
| stage per file — open it in [Perfetto](https://ui.perfetto.dev) to see the | ||
| frontend-internal breakdown (preprocessing, parsing, Sema, PCH | ||
| deserialization) that wall-clock stage timing cannot separate. | ||
|
|
||
| `--log-level info` additionally surfaces the `[perf:index_detail]` lines | ||
| from inside the index stages: semantics-table build vs projection vs | ||
| finishing within `TUIndex::build`, and the path-rekeying copy vs the | ||
| flatbuffers pack within `serialize`. The same lines appear in worker logs | ||
| of a real session, so production runs decompose identically. | ||
|
|
||
| E2E scenarios, clice vs clangd: | ||
|
|
||
| ```bash | ||
| node tools/bench/bench.ts --workspace benchmarks/workloads/llvm \ | ||
| --file clang/lib/Sema/SemaExpr.cpp --json /tmp/clice.json | ||
| node tools/bench/bench.ts --workspace benchmarks/workloads/llvm \ | ||
| --file clang/lib/Sema/SemaExpr.cpp --server clangd --json /tmp/clangd.json | ||
| ``` | ||
|
|
||
| Breakdown of a real (non-benchmark) session from its logs: | ||
|
|
||
| ```bash | ||
| node tools/bench/perf_report.ts <logging_dir>/<session>/*.log --trace /tmp/trace.json | ||
| ``` | ||
|
|
||
| Single-file stage comparison against clangd — pair `pipeline_benchmark` | ||
| (one file selected via `--filter`) with `clangd --check`, which prints its | ||
| preamble build and AST build times for the same TU without a server or | ||
| background indexing in the way: | ||
|
|
||
| ```bash | ||
| ./build/RelWithDebInfo/bin/pipeline_benchmark --filter SemaExpr.cpp --runs 5 \ | ||
| benchmarks/workloads/llvm/build/compile_commands.json | ||
| clangd --check=benchmarks/workloads/llvm/clang/lib/Sema/SemaExpr.cpp \ | ||
| --compile-commands-dir=benchmarks/workloads/llvm/build 2>&1 | grep -E "preamble|AST" | ||
| ``` | ||
|
|
||
| Read them side by side as: clangd "Built preamble in N s" vs our | ||
| `pch_build`, clangd "Building AST" gap vs our `parse_pch`. Everything our | ||
| `pch_build` spends beyond clangd's preamble number is the work clice adds | ||
| to the critical path (TokenBuffer collection, preamble indexing). | ||
|
|
||
| ## Method rules | ||
|
|
||
| - **Fix the machine, compare on the machine.** Absolute numbers are not | ||
| comparable across hosts; run both sides of any A/B on the same machine in | ||
| the same session. | ||
| - **Cold vs warm is a protocol, not an accident.** The harness's | ||
| `cold_start` wipes the cache dirs; everything else is warm. For component | ||
| benchmarks the first run warms the OS file cache — use `--runs` and look | ||
| at percentiles, not single samples. | ||
| - **Idle machine.** No concurrent builds. On WSL2 specifically, do not run | ||
| ninja alongside a benchmark: page-cache-sensitive numbers wobble because | ||
| WSL2 reclaims mmap'd cache aggressively. | ||
| - **One variable at a time.** The pipeline stages and the A/B knobs | ||
| (`collect_tokens`, PCH on/off) exist so a comparison changes exactly one | ||
| thing. | ||
|
|
||
| ## Reference numbers | ||
|
|
||
| Indicative magnitudes from past measured runs — **not** authoritative | ||
| baselines (single machine, dated). Re-measure locally before drawing | ||
| conclusions. | ||
|
|
||
| **Interactive path** (WSL2, RelWithDebInfo, 2026-07; probe TU ≈ 150k | ||
| semantic nodes over `<iostream>/<vector>/<string>`): | ||
|
|
||
| | shape | time | | ||
| | ---------------------------------------------- | -------------------------------- | | ||
| | parse, no PCH | ~330 ms | | ||
| | parse over warm preamble PCH (didChange path) | ~50 ms | | ||
| | TUIndex build, interactive (`interested_only`) | ~0.1–1.4 ms | | ||
| | warm hover end-to-end (tiny project) | ~0.9 ms (feature itself ~0.2 ms) | | ||
|
|
||
| The didChange experience is parse-dominated; features and index are | ||
| sub-millisecond next to it. | ||
|
|
||
| **Monolithic vs chained PCH** (PR #405, LLVM 21 era, 70 stdlib headers): | ||
| full chain build ~2× the monolithic build, but appending one header is | ||
| ~35× faster than the monolithic full rebuild, and AST-load overhead of the | ||
| chain stays within +6% even under full deserialization. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,116 @@ | ||
| #!/usr/bin/env python3 | ||
| """Materialize a pinned benchmark workload from benchmarks/workloads.json. | ||
|
|
||
| Usage: | ||
| python benchmarks/fetch_workload.py <name> | ||
|
|
||
| Clones the workload's repository at its pinned ref into | ||
| benchmarks/workloads/<name> (shallow) and runs the configure command to | ||
| generate compile_commands.json. Both steps are skipped when already done, | ||
| so re-running is cheap. | ||
| """ | ||
|
|
||
| import json | ||
| import subprocess | ||
| import sys | ||
| from pathlib import Path | ||
|
|
||
| BENCH_DIR = Path(__file__).resolve().parent | ||
|
|
||
|
|
||
| def run(args: list[str], cwd: Path) -> None: | ||
| print(f"+ {' '.join(args)}") | ||
| subprocess.run(args, cwd=cwd, check=True) | ||
|
|
||
|
|
||
| def head_commit(root: Path) -> str | None: | ||
| result = subprocess.run( | ||
| ["git", "rev-parse", "--verify", "HEAD"], | ||
| cwd=root, | ||
| capture_output=True, | ||
| text=True, | ||
| ) | ||
| return result.stdout.strip() if result.returncode == 0 else None | ||
|
|
||
|
|
||
| def tracked_files_dirty(root: Path) -> bool: | ||
| """Untracked files are invisible here on purpose: the configure step | ||
| always leaves build output (including the CDB) untracked in the | ||
| checkout, so only tracked-file modifications are detectable drift.""" | ||
| result = subprocess.run( | ||
| ["git", "status", "--porcelain", "--untracked-files=no"], | ||
| cwd=root, | ||
| capture_output=True, | ||
| text=True, | ||
| ) | ||
| return result.returncode != 0 or result.stdout.strip() != "" | ||
|
|
||
|
|
||
| def checkout(root: Path, workload: dict) -> bool: | ||
| """Materialize the pinned ref; returns True when the tree changed. | ||
|
|
||
| The marker file records which ref+commit a completed checkout | ||
| produced: a fetch that died halfway, a ref bumped in workloads.json, | ||
| or a checkout moved by hand all fail the comparison and are redone. | ||
| A matching marker with locally modified tracked files is hard-reset — | ||
| edits must not leak into a run the script reports as pinned. | ||
| """ | ||
| marker = root / ".git" / "workload-ref" | ||
| if (root / ".git").exists(): | ||
| head = head_commit(root) | ||
| if head is not None and marker.exists(): | ||
| if marker.read_text() == f"{workload['ref']} {head}": | ||
| if not tracked_files_dirty(root): | ||
| print(f"{root} already at {workload['ref']}") | ||
| return False | ||
| print(f"{root} has local modifications; resetting") | ||
| run(["git", "reset", "--hard", "--quiet", "HEAD"], cwd=root) | ||
| return True | ||
|
|
||
| root.mkdir(parents=True, exist_ok=True) | ||
| run(["git", "init", "--quiet"], cwd=root) | ||
| run( | ||
| ["git", "fetch", "--depth", "1", workload["git"], workload["ref"]], | ||
| cwd=root, | ||
| ) | ||
| run(["git", "checkout", "--quiet", "FETCH_HEAD"], cwd=root) | ||
| marker.write_text(f"{workload['ref']} {head_commit(root)}") | ||
| return True | ||
|
|
||
|
|
||
| def main() -> int: | ||
| workloads = json.loads((BENCH_DIR / "workloads.json").read_text())["workloads"] | ||
|
|
||
| if len(sys.argv) != 2 or sys.argv[1] not in workloads: | ||
| names = ", ".join(sorted(workloads)) | ||
| print(f"usage: fetch_workload.py <name> (available: {names})") | ||
| return 1 | ||
|
|
||
| name = sys.argv[1] | ||
| workload = workloads[name] | ||
| root = BENCH_DIR / "workloads" / name | ||
|
|
||
| fresh = checkout(root, workload) | ||
|
|
||
| cdb = root / workload["cdb"] | ||
| if fresh or not cdb.exists(): | ||
| run(workload["configure"], cwd=root) | ||
| else: | ||
| print(f"{cdb} already generated") | ||
|
|
||
| print(f"\nworkload ready: {root}") | ||
| print(f"compile_commands.json: {cdb}") | ||
| print("suggested runs:") | ||
| print(f" ./build/RelWithDebInfo/bin/scan_benchmark {cdb}") | ||
| print(f" ./build/RelWithDebInfo/bin/pipeline_benchmark --limit 100 {cdb}") | ||
| scenario = workload.get("scenario") | ||
| if scenario: | ||
| print( | ||
| f" node tools/bench/bench.ts --workspace {root}" | ||
| f" --file {root / scenario['file']}" | ||
| ) | ||
| return 0 | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| sys.exit(main()) |
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.