Skip to content

feat(server): concurrent background indexing with priority control - #432

Merged
16bit-ykiko merged 14 commits into
mainfrom
feat/concurrent-background-indexing
Apr 23, 2026
Merged

16bit-ykiko merged 14 commits into
mainfrom
feat/concurrent-background-indexing

Conversation

@16bit-ykiko

@16bit-ykiko 16bit-ykiko commented Apr 21, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Rewrite serial background indexing to concurrent dispatch (up to stateless_worker_count / 2 parallel tasks)
  • Add depth-counted pause/resume mechanism: completion and signature-help handlers pause new index dispatches to prioritize user requests
  • Report indexing progress via LSP $/progress notifications (percentage + file count)
  • Lower thread scheduling priority (nice +10) for index tasks in stateless workers via RAII ScopedNice guard

Test plan

  • pixi run format — no changes
  • pixi run unit-test Debug — 551 passed, 9 skipped (pre-existing)
  • pixi run smoke-test Debug — 2/2 passed
  • pixi run integration-test Debug — 121 passed, 3 failed (all pre-existing on main: header_context x2, staleness x1)
  • Manual test: open a large project (e.g. LLVM), verify progress bar appears and completion remains responsive during indexing

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Pause/resume controls for background indexing
    • Concurrent, adaptive background indexing with configurable concurrency
    • LSP progress reporting (create/begin/report/end) and updated completion metrics
  • Behavior Change

    • Code completion and signature help temporarily pause indexing for responsiveness
    • Background indexing runs with reduced scheduling priority on non-Windows and logs "files dispatched" at finish
  • Tests

    • Test client fixture defaults init options and sets workspace cache dir

@coderabbitai

coderabbitai Bot commented Apr 21, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds pause/resume controls and concurrency tuning to background indexing, extracts per-file indexing and resource-monitor coroutines, reports LSP $/progress, lowers niceness for index worker processes on non‑Windows, and wires the Indexer into MasterServer to pause around completion/signature-help requests.

Changes

Cohort / File(s) Summary
Indexer core
src/server/indexer.h, src/server/indexer.cpp
Public API: set_peer(), set_max_concurrency(), pause_indexing(), resume_indexing(), ScopedPause, scoped_pause(). New coroutine tasks: index_one() and monitor_resources(); run_background_indexing() refactored into concurrent dispatcher/collector; added fields for pause_depth, resume_event, inflight, finished, monitor_generation, progress reporting (LSP $/progress).
MasterServer integration
src/server/master_server.cpp
Indexer wired into server (set_peer, set_max_concurrency). Completion and SignatureHelp handlers now acquire indexer.scoped_pause() to suspend background indexing during those requests; handler registration adjusted.
Stateless worker niceness
src/server/stateless_worker.cpp
Added platform-guarded ScopedNice (non‑Windows) to lower scheduling priority during BuildKind::Index handling; applied it around handle_index(params).
Tests fixture
tests/conftest.py
client fixture normalizes init_options (use {} when absent or copy marker dict), ensures init_options["project"]["cache_dir"] points to workspace / ".clice". Improved server shutdown logging: check server.returncode, read stderr if present, and expanded stderr filter to include Sanitizer.

Sequence Diagram(s)

sequenceDiagram
    participant LSP as LSP Client
    participant MS as MasterServer
    participant IDX as Indexer
    participant WRK as StatelessWorker
    participant DB as Disk/Build System

    LSP->>MS: Completion / SignatureHelp request
    activate MS
    MS->>IDX: scoped_pause() (pause_indexing)
    activate IDX
    IDX->>IDX: increment pause_depth, wait gate
    deactivate IDX

    MS->>WRK: forward build / handle_completion
    activate WRK
    WRK->>WRK: ScopedNice (non‑Windows) lowers priority
    WRK->>DB: perform build/index
    DB-->>WRK: build result (TUIndex / error)
    WRK-->>MS: return result
    deactivate WRK

    MS->>IDX: resume_indexing()
    activate IDX
    IDX->>IDX: decrement pause_depth, signal resume_event if zero
    IDX-->>MS: resumed
    deactivate IDX
    MS-->>LSP: return result
    deactivate MS

    par Background indexing (concurrent)
        IDX->>IDX: monitor_resources() adjusts concurrency
        loop dispatch loop (up to max_concurrent)
            IDX->>IDX: dispatch index_one(server_path_id)
            activate WRK
            WRK->>DB: index file
            DB-->>WRK: TUIndex/outcome
            WRK-->>IDX: completion_event
            deactivate WRK
        end
    end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Poem

🐰 I pause my paws when builders hum,

I hop and dispatch, then softly come,
I lower steps so nighttime sings,
I count the ticks of progress rings,
And nibble carrots made of things.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 30.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main changes: introducing concurrent background indexing and priority control mechanisms via pause/resume and scheduling priority adjustments.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/concurrent-background-indexing

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (2)
src/server/indexer.h (1)

82-83: Include <algorithm> for std::max.

Indexer::set_max_concurrency() uses std::max, but indexer.h does not include <algorithm>. Relying on transitive includes makes this header fragile.

Proposed include fix
 `#include` <cstdint>
+#include <algorithm>
 `#include` <functional>
 `#include` <memory>
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/server/indexer.h` around lines 82 - 83, Add a direct include of the
<algorithm> header in indexer.h so Indexer::set_max_concurrency can safely use
std::max; update indexer.h to `#include` <algorithm> near the other standard
includes and ensure the declaration for void set_max_concurrency(std::size_t n)
and the member max_concurrent remain unchanged.
src/server/indexer.cpp (1)

3-5: Include <format> where std::format is used.

This file now calls std::format in progress reporting but does not explicitly include <format>.

Proposed include fix
+#include <format>
 `#include` <string>
 `#include` <variant>
 `#include` <vector>
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/server/indexer.cpp` around lines 3 - 5, Add a direct include for <format>
at the top of this translation unit so uses of std::format in the progress
reporting code compile; specifically insert `#include` <format> alongside the
existing includes (which currently show <string>, <variant>, <vector>) so calls
to std::format in the progress reporting / reporting function (where std::format
is invoked) resolve correctly. Ensure the file is compiled with C++20 or later
if not already.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/server/indexer.cpp`:
- Around line 703-713: The progress total currently uses total =
index_queue.size() which counts already-processed items; change the progress
denominator to the current batch/remaining count by computing remaining =
index_queue.size() - index_queue_pos (use std::size_t to match types) and pass
that remaining value into progress->begin(...) and any later progress updates
(e.g., the updates around the progress->report or progress->end calls at the
later block referenced). Update references to total in the progress messages and
percentage calculations to use remaining so progress moves from 0 to 100% for
each run.
- Around line 721-759: The loop currently uses a binary kota::event
(completion_event) and increments completed once per wait, which misses multiple
task completions; fix by adding an atomic counter (e.g.,
std::atomic<std::uint32_t> finished_count) that each child task increments after
co_await self->index_one(id) (inside the lambda launched by loop.schedule), then
still call completion_event.set(); in the main loop replace the single
++completed after co_await completion_event.wait() with reading and zeroing the
atomic (delta = finished_count.exchange(0)) and doing completed += delta (so you
track the exact number of completed tasks); update references to inflight and
index_one as needed but keep the event for waking the loop.

In `@src/server/master_server.cpp`:
- Around line 684-688: The pause/resume around co_await is not exception-safe;
create a small RAII helper (e.g., class ScopedIndexingPause) that calls
indexer.pause_indexing() in its constructor and indexer.resume_indexing() in its
destructor, then replace the explicit pause_indexing()/resume_indexing() pairs
around compiler.handle_completion(...) (and the similar pair at the other
handler) by instantiating ScopedIndexingPause before the co_await so
resume_indexing() is guaranteed even on exceptions/unwinding.

In `@src/server/stateless_worker.cpp`:
- Around line 22-36: The ScopedNice RAII currently enabled on all non-Windows
platforms should be restricted to Linux only and must check for failures from
getpriority/setpriority: change the guard to `#if` defined(__linux__) and in
ScopedNice's constructor call getpriority(PRIO_PROCESS, 0) while checking for -1
and errno to detect errors before using the value (store a safe sentinel if
getpriority fails), and only call setpriority if the syscall succeeds (check its
return value and errno); similarly in the destructor, check the return of
setpriority and avoid restoring if the saved value is invalid—use the existing
ScopedNice, constructor, destructor, getpriority, and setpriority identifiers
when making these changes.

---

Nitpick comments:
In `@src/server/indexer.cpp`:
- Around line 3-5: Add a direct include for <format> at the top of this
translation unit so uses of std::format in the progress reporting code compile;
specifically insert `#include` <format> alongside the existing includes (which
currently show <string>, <variant>, <vector>) so calls to std::format in the
progress reporting / reporting function (where std::format is invoked) resolve
correctly. Ensure the file is compiled with C++20 or later if not already.

In `@src/server/indexer.h`:
- Around line 82-83: Add a direct include of the <algorithm> header in indexer.h
so Indexer::set_max_concurrency can safely use std::max; update indexer.h to
`#include` <algorithm> near the other standard includes and ensure the declaration
for void set_max_concurrency(std::size_t n) and the member max_concurrent remain
unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 4d1e41fe-b4e0-4231-b436-d2eab3604b33

📥 Commits

Reviewing files that changed from the base of the PR and between 17e6801 and 53bfc84.

📒 Files selected for processing (4)
  • src/server/indexer.cpp
  • src/server/indexer.h
  • src/server/master_server.cpp
  • src/server/stateless_worker.cpp

Comment thread src/server/indexer.cpp Outdated
Comment thread src/server/indexer.cpp Outdated
Comment thread src/server/master_server.cpp Outdated
Comment thread src/server/stateless_worker.cpp Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/server/indexer.cpp`:
- Around line 690-718: monitor_resources can outlive a run because
run_background_indexing flips indexing_active to false but doesn't await the
monitor coroutine, allowing an old monitor suspended in kota::sleep to resume
after a new run sets indexing_active true and thus two monitors mutate
max_concurrent/baseline_concurrent concurrently; fix by ensuring per-run
cancellation/awaiting: either make run_background_indexing signal and await
monitor_resources completion (use a generation token, kota::event or
promise/future per run) so monitor_resources exits when its run is stopped, or
co_await the monitor task before returning from run_background_indexing;
alternatively shorten the kota::sleep and check a run-scoped cancellation token
each loop iteration to ensure the old monitor stops before a new one starts
(refer to monitor_resources, run_background_indexing, indexing_active,
max_concurrent, baseline_concurrent, and kota::sleep).
- Around line 775-781: The lambda passed to loop.schedule currently co_awaits
Indexer::index_one and only decrements inflight and calls completion_event.set()
on the happy path, which leaks a concurrency slot if index_one throws; modify
the scheduled lambda (the kota::task<> created in loop.schedule) to guarantee
the decrement of the Indexer::inflight counter and a completion_event.set() on
every exit path — e.g. wrap the co_await self->index_one(id) in a try/finally
pattern or install a small RAII/scope-guard local (created before the co_await)
that in its destructor decrements inflight and calls done.set(), so regardless
of exceptions the dispatcher won’t deadlock.

In `@tests/conftest.py`:
- Around line 113-117: The code currently mutates the marker's original dict via
init_options = init_options_marker.args[0]; instead, create a shallow copy
(e.g., init_options = dict(init_options_marker.args[0]) if init_options_marker
else {}) before modifying it so the original marker data is not mutated across
test runs; then ensure the copied init_options has a "project" dict and set
default "cache_dir" to str(workspace / ".clice") on that copy (leave
marker.args[0] untouched).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: c17a4015-71c9-4844-bb3d-e186c74b4663

📥 Commits

Reviewing files that changed from the base of the PR and between 53bfc84 and a40e8b9.

📒 Files selected for processing (3)
  • src/server/indexer.cpp
  • src/server/indexer.h
  • tests/conftest.py

Comment thread src/server/indexer.cpp Outdated
Comment thread src/server/indexer.cpp Outdated
Comment thread tests/conftest.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
src/server/indexer.cpp (1)

779-784: ⚠️ Potential issue | 🟠 Major

Exception from index_one still leaks a concurrency slot.

The dispatch lambda unconditionally runs --inflight; ++finished; done.set(); after co_await self->index_one(id). If index_one (or anything it awaits — pool.send_stateless, merge, allocations in fill_compile_args/fill_pcm_deps) throws, the coroutine unwinds past those lines. inflight stays pinned above zero and no further done.set() is delivered, so the outer loop hangs on completion_event.wait() and save() never runs.

Wrap the body so the decrement and signal fire on every exit path:

🛡️ Exception-safe dispatch
             loop.schedule([](Indexer* self, std::uint32_t id, kota::event& done) -> kota::task<> {
-                co_await self->index_one(id);
-                --self->inflight;
-                ++self->finished;
-                done.set();
+                try {
+                    co_await self->index_one(id);
+                } catch(const std::exception& e) {
+                    LOG_WARN("Background index task threw: {}", e.what());
+                } catch(...) {
+                    LOG_WARN("Background index task threw unknown exception");
+                }
+                --self->inflight;
+                ++self->finished;
+                done.set();
             }(this, server_path_id, completion_event));
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/server/indexer.cpp` around lines 779 - 784, The lambda passed to
loop.schedule currently does co_await self->index_one(id) and then
unconditionally does --self->inflight; ++self->finished; done.set(); but if
index_one or any awaited operation throws the decrement/signal are skipped; fix
by making the lambda exception-safe: wrap the co_await in a try/catch (or use a
small scope-guard) inside the dispatched coroutine so that --self->inflight,
++self->finished and done.set() are executed on every exit path (both success
and on any exception), and rethrow the exception after performing those cleanup
actions so behavior and error propagation remain intact; target the anonymous
lambda in loop.schedule and the cleanup of inflight/finished/done there.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/server/indexer.cpp`:
- Around line 735-800: The progress denominator `batch` is captured once before
dispatch and can go stale as `index_queue` grows, causing percentages >100;
capture the starting queue position (e.g. store initial_index_queue_pos =
index_queue_pos at the top), then when reporting compute the live total as
current_batch = index_queue.size() - initial_index_queue_pos (use that instead
of the original `batch`) and clamp the computed pct to a max of 100 before
calling progress->report(...) (or as a safety also clamp `completed` to
current_batch). Update references around `batch`, `index_queue`,
`index_queue_pos`, `completed`, and the call to progress->report(...)
accordingly.

---

Duplicate comments:
In `@src/server/indexer.cpp`:
- Around line 779-784: The lambda passed to loop.schedule currently does
co_await self->index_one(id) and then unconditionally does --self->inflight;
++self->finished; done.set(); but if index_one or any awaited operation throws
the decrement/signal are skipped; fix by making the lambda exception-safe: wrap
the co_await in a try/catch (or use a small scope-guard) inside the dispatched
coroutine so that --self->inflight, ++self->finished and done.set() are executed
on every exit path (both success and on any exception), and rethrow the
exception after performing those cleanup actions so behavior and error
propagation remain intact; target the anonymous lambda in loop.schedule and the
cleanup of inflight/finished/done there.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 15609e4d-7d32-4c9c-a096-fa7bb7facf1f

📥 Commits

Reviewing files that changed from the base of the PR and between a40e8b9 and 72c46b7.

📒 Files selected for processing (4)
  • src/server/indexer.cpp
  • src/server/indexer.h
  • src/server/master_server.cpp
  • tests/conftest.py

Comment thread src/server/indexer.cpp

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/server/indexer.h (1)

82-99: Optional: consider making ScopedPause movable.

ScopedPause deletes copy and does not declare a move, so scoped_pause() only works at the call site thanks to C++17 mandatory prvalue copy elision (e.g. auto g = indexer.scoped_pause();). Transferring ownership across scopes (storing in an optional, returning further up, moving into a member) won’t compile. If that’s intentional (guard is meant to stay put), no change needed; otherwise add a move constructor that null-checks on destruction.

Proposed move-enabled variant (only if needed)
     struct [[nodiscard]] ScopedPause {
-        Indexer& indexer;
+        Indexer* indexer;

-        explicit ScopedPause(Indexer& idx) : indexer(idx) {
-            indexer.pause_indexing();
+        explicit ScopedPause(Indexer& idx) : indexer(&idx) {
+            indexer->pause_indexing();
         }

         ~ScopedPause() {
-            indexer.resume_indexing();
+            if(indexer) indexer->resume_indexing();
         }

         ScopedPause(const ScopedPause&) = delete;
         ScopedPause& operator=(const ScopedPause&) = delete;
+        ScopedPause(ScopedPause&& o) noexcept : indexer(o.indexer) { o.indexer = nullptr; }
+        ScopedPause& operator=(ScopedPause&&) = delete;
     };
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/server/indexer.h` around lines 82 - 99, ScopedPause currently deletes
copy and has no move, preventing transfers; change its internals to support move
semantics by storing an Indexer* (e.g. Indexer* indexer) instead of a reference,
keep the explicit constructor to set the pointer and call
indexer->pause_indexing(), implement a move constructor that steals the pointer
(setting the source.pointer to nullptr) and a move assignment (or delete move
assignment if you prefer), keep copy operations deleted, and modify the
destructor to only call resume_indexing() when the pointer is non-null; ensure
scoped_pause() still returns a ScopedPause by value.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@src/server/indexer.h`:
- Around line 82-99: ScopedPause currently deletes copy and has no move,
preventing transfers; change its internals to support move semantics by storing
an Indexer* (e.g. Indexer* indexer) instead of a reference, keep the explicit
constructor to set the pointer and call indexer->pause_indexing(), implement a
move constructor that steals the pointer (setting the source.pointer to nullptr)
and a move assignment (or delete move assignment if you prefer), keep copy
operations deleted, and modify the destructor to only call resume_indexing()
when the pointer is non-null; ensure scoped_pause() still returns a ScopedPause
by value.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 13754cf8-1ce5-4611-8f05-816234334c56

📥 Commits

Reviewing files that changed from the base of the PR and between fb6c2d4 and 5da721e.

📒 Files selected for processing (2)
  • src/server/indexer.cpp
  • src/server/indexer.h
✅ Files skipped from review due to trivial changes (1)
  • src/server/indexer.cpp

@16bit-ykiko
16bit-ykiko force-pushed the feat/concurrent-background-indexing branch 2 times, most recently from 11c4fe3 to 59c11ad Compare April 22, 2026 08:24
16bit-ykiko added a commit that referenced this pull request Apr 23, 2026
…d missing PCH cache dir (#435)

## Summary

Three pre-existing bugs cause worker processes to crash with SEGV or
SIGABRT. On the main branch these crashes are silent (workers die,
requests fail fast with "transport closed", tests still pass because
null responses are accepted). However when combined with #432's worker
respawn mechanism, the crash-respawn-crash cycle on low-core CI machines
causes request timeouts and smoke test hangs.

### Fixes

- **compilation.cpp**: `ProxyAction::CreateASTConsumer` now checks for
null before passing to `MultiplexConsumer`. When the wrapped action's
`CreateASTConsumer` fails (e.g. missing system headers during PCH
generation), this previously caused a null pointer dereference, SEGV,
ASAN kills the stateless worker.
- **compilation_unit.cpp**: `file_path()` returns empty `StringRef` on
invalid `FileID` instead of asserting. The assert fired when
`IncludeGraph::from()` called `file_path(interested_file())` on an AST
compiled with synthesized default commands (no compile_commands.json,
clang++ -std=c++20 fallback, no system headers, invalid main file ID),
SIGABRT, stateful worker crash.
- **compiler.cpp**: `ensure_pch` now creates the PCH cache directory
before sending the build request. Previously, when `load_workspace()`
exited early (no compile_commands.json), the cache subdirectories were
never created, causing every PCH write to fail with "No such file or
directory".
- **master_server.cpp/h**: `load_workspace()` changed from
`kota::task<>` to plain `void` -- it contains only synchronous
filesystem operations and no co_await, so the coroutine wrapper was
unnecessary. Called directly instead of via `loop.schedule()`.

## Test plan

- [x] Verified zero SEGV/SIGABRT/assertion crashes in worker stderr
after fix
- [x] rapid_edit.jsonl smoke test passes 3/3 runs consistently (34s
each)
- [x] Behavior matches main branch (both return 134 responses, 0
pending)
- [x] Debug build with ASAN (detect_leaks=0) -- clean run, no sanitizer
reports

<!-- codesmith:footer -->
---
<a
href="https://app.blacksmith.sh/clice-io/codesmith/clice/pr/435"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://pr-comments-assets.blacksmith.sh/codesmith/view-in-codesmith-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://pr-comments-assets.blacksmith.sh/codesmith/view-in-codesmith-light.svg"><img
alt="View in Codesmith"
src="https://pr-comments-assets.blacksmith.sh/codesmith/view-in-codesmith-dark.svg"></picture></a>
<sup>Codesmith can help with this PR — just tag <code>@codesmith</code>
or enable autofix.</sup>

- [ ] Autofix CI and bot reviews
<!-- /codesmith:footer -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved error handling for AST consumer creation with null checks and
a clear failure path.
* Safer file-path access that returns empty for invalid identifiers
instead of asserting.
* PCH cache handling now validates cache configuration, attempts
directory creation, logs warnings, and aborts PCH builds on failure.

* **Refactor**
* Workspace loading changed from asynchronous to synchronous execution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
16bit-ykiko and others added 14 commits April 23, 2026 10:40
Rewrite background indexing from serial to concurrent dispatch, with
pause/resume for user request priority and LSP progress reporting.

- Dispatch up to max_concurrent index tasks simultaneously via kota
  event loop scheduling, with inflight counting and event signaling
- Add pause/resume mechanism (depth-counted) so completion and
  signature-help requests temporarily halt new index dispatches
- Report indexing progress via LSP $/progress notifications
- Lower thread priority (nice +10) for index tasks in stateless
  workers so user-facing requests are not starved

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Integration tests could fail when stale PCH files from a previous
Clang version remained in the XDG cache directory. Force cache_dir
into the workspace via initializationOptions so the existing .clice/
cleanup in the workspace fixture prevents cross-run cache pollution.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add a monitor_resources() coroutine that runs alongside background
indexing and adjusts max_concurrent based on system memory availability.
Reduces concurrency when available memory drops below 15% (respecting
cgroup limits), and restores toward the baseline when memory recovers
above 30%.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Progress total: use current batch size instead of lifetime queue size
- Completed counter: use per-task `finished` counter to avoid binary
  event coalescence undercounting
- Pause/resume: add RAII ScopedPause guard for exception safety across
  co_await boundaries
- Monitor lifetime: use generation counter so a stale monitor_resources
  coroutine exits when its owning run ends
- conftest.py: shallow-copy marker init_options to avoid cross-test
  mutation

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Aids CI debugging by printing the server process exit code and any
AddressSanitizer output in test teardown.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Move completion_event to Indexer member to fix heap-use-after-free when
server shuts down with inflight index tasks (the coroutine frame was
destroyed while lambdas still held a reference to the local event).

Enable background indexing for module interface units: compile their PCMs
via CompileGraph before dispatching to the stateless worker.  Partition
the queue so modules are indexed first, ensuring PCMs are available for
non-module files that import them.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Stateless worker count now defaults to parallelism()/2 instead of
a hardcoded 3, scaling with the machine's CPU cores.  Indexer
max_concurrent uses the full stateless count since interactive requests
go through stateful workers.

Worker processes are now monitored: on unexpected exit the crash signal
and stderr output (stack traces, sanitizer reports) are logged, and the
worker is automatically respawned up to 5 times.  send_stateless skips
dead workers in the round-robin.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When respawn_worker() replaced a WorkerProcess, the old BincodePeer was
destroyed while its run()/write_loop() coroutines were still scheduled
in the event loop. On the next tick they'd resume into freed memory,
crashing the master server. Fix by calling close() on the old peer and
moving it to a retired_peers list so it stays alive until its coroutines
finish naturally.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
BindingDecl::getType() can return a null QualType in error-recovery or
dependent contexts. decl_of() dereferenced it unconditionally via
type->getAs<>(), crashing the worker with SIGSEGV. Add an early null
check to protect all callers.

This was the sole cause of all 60 worker crashes observed during LLVM
background indexing (4695 files).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace POSIX-only getpriority/setpriority with kota::sys::priority()
and kota::sys::set_priority(), removing the #ifndef _WIN32 guard.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
send_stateful() and notify_stateful() lacked alive checks, causing the
master server to hang when a stateful worker died and exceeded max
restarts. Also add alive-awareness to pick_least_loaded() so new
documents are never assigned to dead workers.

Additionally, enforce a 5-minute wall-clock timeout per test in both
smoke tests (replay.py --wall-timeout) and integration tests
(pytest-timeout), preventing CI from hanging indefinitely.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Prevents CI jobs from hanging for hours if a test hangs despite
the per-test timeouts in replay.py and pytest-timeout.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add concurrency group to cancel in-progress CI on non-main branches
when new commits are pushed. Split the monolithic "Run tests" step into
separate Unit/Integration/Smoke test steps with individual timeouts for
better visibility into which test phase hangs or fails.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Change load_workspace() to synchronous call (not scheduled as async task)
- Increase smoke test CI timeout to 15 minutes
- Force line-buffered stdout in replay.py for CI output visibility

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@16bit-ykiko
16bit-ykiko force-pushed the feat/concurrent-background-indexing branch from ec15896 to ba42027 Compare April 23, 2026 03:45
@16bit-ykiko
16bit-ykiko merged commit 939ab6d into main Apr 23, 2026
35 of 37 checks passed
@16bit-ykiko
16bit-ykiko deleted the feat/concurrent-background-indexing branch April 23, 2026 05:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant