Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,9 @@
temp/
.cache/

# Materialized benchmark workloads (fetch_workload.py)
benchmarks/workloads/

.llvm*/
.clice/
compile_commands.json
Expand Down
19 changes: 12 additions & 7 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -234,13 +234,18 @@ if(CLICE_ENABLE_TEST)
endif()

if(CLICE_ENABLE_BENCHMARK)
add_executable(scan_benchmark
"${PROJECT_SOURCE_DIR}/benchmarks/scan_benchmark.cpp"
)
target_include_directories(scan_benchmark PRIVATE
"${PROJECT_SOURCE_DIR}/src"
)
target_link_libraries(scan_benchmark PRIVATE clice::core kota::deco)
foreach(benchmark scan_benchmark pipeline_benchmark pch_chain_benchmark)
add_executable(${benchmark}
"${PROJECT_SOURCE_DIR}/benchmarks/${benchmark}.cpp"
)
target_include_directories(${benchmark} PRIVATE
"${PROJECT_SOURCE_DIR}/src"
)
target_link_libraries(${benchmark} PRIVATE clice::core kota::deco)
Comment thread
16bit-ykiko marked this conversation as resolved.
# The benchmarks locate lib/clang next to their binary; an explicit
# `ninja <benchmark>` must not race or skip the resource copy.
add_dependencies(${benchmark} copy_clang_resource)
endforeach()
endif()

include("${PROJECT_SOURCE_DIR}/cmake/release.cmake")
148 changes: 148 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,148 @@
# Benchmarks

Performance work on clice runs on three layers. Pick the layer that answers
your question:

1. **Instrumentation** — _where does the time go in a real run?_ Every
server run emits `[perf:<topic>] key=value` log lines (see
`src/support/logging.h`) covering startup phases, per-file compiles,
PCH/PCM/index builds, cache hits, request latencies and index queries.
`tools/bench/perf_report.ts` aggregates any log into per-series
percentiles and can export a Chrome trace. This layer needs no special
build and works on logs users attach to issue reports.

2. **Scenario harness** — _what does the user experience end to end?_
`tools/bench/bench.ts` drives a server over LSP through fixed scenarios
(cold start, warm start, edit loop, warm feature requests) on a real
workspace and reports client-observed percentiles. It is server-agnostic:
point it at clangd with `--server clangd` to A/B the same scenario.

3. **Component benchmarks** — _which design alternative is faster?_
Standalone binaries in this directory, built with
`-DCLICE_ENABLE_BENCHMARK=ON`, each answering one decision:
- `scan_benchmark` — dependency-graph scan over a real CDB.
- `pipeline_benchmark` — per-TU stage profile (preprocess with/without
TokenBuffer, parse, index build/serialize, preamble PCH build incl.
preamble indexing, reparse over PCH incl. interactive indexing), one
result per file.
- `pch_chain_benchmark` — monolithic vs chained PCH strategy (ported
from PR #405).

## Building

```bash
cmake -B build/RelWithDebInfo -G Ninja -DCMAKE_BUILD_TYPE=RelWithDebInfo \
-DCMAKE_TOOLCHAIN_FILE=cmake/toolchain.cmake -DCLICE_ENABLE_BENCHMARK=ON
ninja -C build/RelWithDebInfo scan_benchmark pipeline_benchmark pch_chain_benchmark
```

Always benchmark `RelWithDebInfo`; Debug numbers are meaningless.

## Workloads

Benchmarks against real projects must be pinned to be comparable.
`workloads.json` records project + ref + configure command;
`fetch_workload.py <name>` materializes one under `benchmarks/workloads/`
(shallow clone + CMake configure — no build needed, only the
compile_commands.json):

```bash
python benchmarks/fetch_workload.py llvm
```

clice's own CDB (`build/RelWithDebInfo/compile_commands.json`) doubles as
an always-available medium workload.

## Typical sessions

Stage profile of the 100 largest TUs plus a Chrome trace of one:

```bash
./build/RelWithDebInfo/bin/pipeline_benchmark --limit 100 --json /tmp/pipeline.json \
benchmarks/workloads/llvm/build/compile_commands.json
./build/RelWithDebInfo/bin/pipeline_benchmark --filter SemaExpr.cpp --runs 3 \
--time-trace /tmp/traces benchmarks/workloads/llvm/build/compile_commands.json
```

`--time-trace` writes clang's own `-ftime-trace` profile of the parse
stage per file — open it in [Perfetto](https://ui.perfetto.dev) to see the
frontend-internal breakdown (preprocessing, parsing, Sema, PCH
deserialization) that wall-clock stage timing cannot separate.

`--log-level info` additionally surfaces the `[perf:index_detail]` lines
from inside the index stages: semantics-table build vs projection vs
finishing within `TUIndex::build`, and the path-rekeying copy vs the
flatbuffers pack within `serialize`. The same lines appear in worker logs
of a real session, so production runs decompose identically.

E2E scenarios, clice vs clangd:

```bash
node tools/bench/bench.ts --workspace benchmarks/workloads/llvm \
--file clang/lib/Sema/SemaExpr.cpp --json /tmp/clice.json
node tools/bench/bench.ts --workspace benchmarks/workloads/llvm \
--file clang/lib/Sema/SemaExpr.cpp --server clangd --json /tmp/clangd.json
```

Breakdown of a real (non-benchmark) session from its logs:

```bash
node tools/bench/perf_report.ts <logging_dir>/<session>/*.log --trace /tmp/trace.json
```

Single-file stage comparison against clangd — pair `pipeline_benchmark`
(one file selected via `--filter`) with `clangd --check`, which prints its
preamble build and AST build times for the same TU without a server or
background indexing in the way:

```bash
./build/RelWithDebInfo/bin/pipeline_benchmark --filter SemaExpr.cpp --runs 5 \
benchmarks/workloads/llvm/build/compile_commands.json
clangd --check=benchmarks/workloads/llvm/clang/lib/Sema/SemaExpr.cpp \
--compile-commands-dir=benchmarks/workloads/llvm/build 2>&1 | grep -E "preamble|AST"
```

Read them side by side as: clangd "Built preamble in N s" vs our
`pch_build`, clangd "Building AST" gap vs our `parse_pch`. Everything our
`pch_build` spends beyond clangd's preamble number is the work clice adds
to the critical path (TokenBuffer collection, preamble indexing).

## Method rules

- **Fix the machine, compare on the machine.** Absolute numbers are not
comparable across hosts; run both sides of any A/B on the same machine in
the same session.
- **Cold vs warm is a protocol, not an accident.** The harness's
`cold_start` wipes the cache dirs; everything else is warm. For component
benchmarks the first run warms the OS file cache — use `--runs` and look
at percentiles, not single samples.
- **Idle machine.** No concurrent builds. On WSL2 specifically, do not run
ninja alongside a benchmark: page-cache-sensitive numbers wobble because
WSL2 reclaims mmap'd cache aggressively.
- **One variable at a time.** The pipeline stages and the A/B knobs
(`collect_tokens`, PCH on/off) exist so a comparison changes exactly one
thing.

## Reference numbers

Indicative magnitudes from past measured runs — **not** authoritative
baselines (single machine, dated). Re-measure locally before drawing
conclusions.

**Interactive path** (WSL2, RelWithDebInfo, 2026-07; probe TU ≈ 150k
semantic nodes over `<iostream>/<vector>/<string>`):

| shape | time |
| ---------------------------------------------- | -------------------------------- |
| parse, no PCH | ~330 ms |
| parse over warm preamble PCH (didChange path) | ~50 ms |
| TUIndex build, interactive (`interested_only`) | ~0.1–1.4 ms |
| warm hover end-to-end (tiny project) | ~0.9 ms (feature itself ~0.2 ms) |

The didChange experience is parse-dominated; features and index are
sub-millisecond next to it.

**Monolithic vs chained PCH** (PR #405, LLVM 21 era, 70 stdlib headers):
full chain build ~2× the monolithic build, but appending one header is
~35× faster than the monolithic full rebuild, and AST-load overhead of the
chain stays within +6% even under full deserialization.
116 changes: 116 additions & 0 deletions benchmarks/fetch_workload.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,116 @@
#!/usr/bin/env python3
"""Materialize a pinned benchmark workload from benchmarks/workloads.json.

Usage:
python benchmarks/fetch_workload.py <name>

Clones the workload's repository at its pinned ref into
benchmarks/workloads/<name> (shallow) and runs the configure command to
generate compile_commands.json. Both steps are skipped when already done,
so re-running is cheap.
"""

import json
import subprocess
import sys
from pathlib import Path

BENCH_DIR = Path(__file__).resolve().parent


def run(args: list[str], cwd: Path) -> None:
print(f"+ {' '.join(args)}")
subprocess.run(args, cwd=cwd, check=True)


def head_commit(root: Path) -> str | None:
result = subprocess.run(
["git", "rev-parse", "--verify", "HEAD"],
cwd=root,
capture_output=True,
text=True,
)
return result.stdout.strip() if result.returncode == 0 else None


def tracked_files_dirty(root: Path) -> bool:
"""Untracked files are invisible here on purpose: the configure step
always leaves build output (including the CDB) untracked in the
checkout, so only tracked-file modifications are detectable drift."""
result = subprocess.run(
["git", "status", "--porcelain", "--untracked-files=no"],
cwd=root,
capture_output=True,
text=True,
)
return result.returncode != 0 or result.stdout.strip() != ""


def checkout(root: Path, workload: dict) -> bool:
"""Materialize the pinned ref; returns True when the tree changed.

The marker file records which ref+commit a completed checkout
produced: a fetch that died halfway, a ref bumped in workloads.json,
or a checkout moved by hand all fail the comparison and are redone.
A matching marker with locally modified tracked files is hard-reset —
edits must not leak into a run the script reports as pinned.
"""
marker = root / ".git" / "workload-ref"
if (root / ".git").exists():
head = head_commit(root)
if head is not None and marker.exists():
if marker.read_text() == f"{workload['ref']} {head}":
if not tracked_files_dirty(root):
print(f"{root} already at {workload['ref']}")
return False
print(f"{root} has local modifications; resetting")
run(["git", "reset", "--hard", "--quiet", "HEAD"], cwd=root)
return True

root.mkdir(parents=True, exist_ok=True)
run(["git", "init", "--quiet"], cwd=root)
run(
["git", "fetch", "--depth", "1", workload["git"], workload["ref"]],
cwd=root,
)
run(["git", "checkout", "--quiet", "FETCH_HEAD"], cwd=root)
marker.write_text(f"{workload['ref']} {head_commit(root)}")
return True


def main() -> int:
workloads = json.loads((BENCH_DIR / "workloads.json").read_text())["workloads"]

if len(sys.argv) != 2 or sys.argv[1] not in workloads:
names = ", ".join(sorted(workloads))
print(f"usage: fetch_workload.py <name> (available: {names})")
return 1

name = sys.argv[1]
workload = workloads[name]
root = BENCH_DIR / "workloads" / name

fresh = checkout(root, workload)

cdb = root / workload["cdb"]
if fresh or not cdb.exists():
run(workload["configure"], cwd=root)
else:
print(f"{cdb} already generated")

print(f"\nworkload ready: {root}")
print(f"compile_commands.json: {cdb}")
print("suggested runs:")
print(f" ./build/RelWithDebInfo/bin/scan_benchmark {cdb}")
print(f" ./build/RelWithDebInfo/bin/pipeline_benchmark --limit 100 {cdb}")
scenario = workload.get("scenario")
if scenario:
print(
f" node tools/bench/bench.ts --workspace {root}"
f" --file {root / scenario['file']}"
)
return 0


if __name__ == "__main__":
sys.exit(main())
Loading