Skip to content

fix(server): structured shutdown with when_any cancellation - #444

Merged
16bit-ykiko merged 10 commits into
mainfrom
fix/server-shutdown-asan
Jun 8, 2026
Merged

16bit-ykiko merged 10 commits into
mainfrom
fix/server-shutdown-asan

Conversation

@16bit-ykiko

@16bit-ykiko 16bit-ykiko commented May 31, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Structured shutdown via when_any: All server modes (pipe, socket, daemon) now use when_any to race the main transport loop against shutdown_event.wait(). When shutdown is signaled, when_any cancels all sibling tasks — including pending I/O — then runs shutdown_and_cleanup() sequentially. This replaces the old pattern of fire-and-forget loop.schedule() + loop.stop().
  • Bump kotatsu to 2a8c147 (deferred sync resume): kotatsu's sync primitives (event, mutex, semaphore, cv) now defer waiter resumes to the event loop instead of resuming inline. This eliminates the reentrancy that previously required a sentinel workaround (child == self) and made cancel unable to propagate through a task whose coroutine frame was still on the call stack.
  • Sanitizer-clean exit enforcement in tests: conftest.py now checks every server exit for non-zero return code and sanitizer output (AddressSanitizer, LeakSanitizer, etc.) via assert_server_exited_cleanly(). Any sanitizer finding or unclean exit fails the test.
  • Fix indexer monitor task lifetime: run_background_indexing previously spawned the resource monitor into bg_tasks (a class member). Now uses a local task_group that is joined before returning, ensuring the monitor is fully stopped before the indexer reports completion.
  • Fix worker_pool dangling reference: monitor_worker held a reference to workers[index] across co_await proc.wait(), but the vector could reallocate during the wait. The reference is now taken after the await.

Design decisions

  • resilient_file_watcher wrapper (daemon mode): file_watcher_task() co_returns on creation failure. Without wrapping, this would become the when_any winner and trigger daemon shutdown. The wrapper suspends on shutdown_event.wait() after watcher exit, so only the real shutdown signal terminates the daemon.
  • schedule_shutdown guards on Exited, not ShuttingDown: The LSP shutdown request sets lifecycle to ShuttingDown before the exit notification calls schedule_shutdown(). If the guard checked ShuttingDown, the event would never be signaled. event.set() is idempotent, so repeated calls are safe.
  • Accept loop wrapped in task_group.spawn: The accept loop is spawned as a child of a task_group rather than running directly in accept_connections. This ensures when_any cancellation propagates through the task_group to both the accept loop and all active connection tasks.

Summary by CodeRabbit

  • Bug Fixes

    • More reliable shutdown with staged cleanup and persisted index/cache
    • Resilient file-watching that reliably reloads workspace changes
  • Improvements

    • Reworked connection handling with per-connection tasks and a single lazy LSP client
    • Coordinated server modes and acceptor behavior; optional daemon file-watching
    • More predictable worker pool startup and monitoring
  • Tests

    • Stronger test assertions enforcing clean server exit and stderr checks

@coderabbitai

coderabbitai Bot commented May 31, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: c9ca1c6d-cd96-4b2b-b321-c9da96971d40

📥 Commits

Reviewing files that changed from the base of the PR and between cf7d774 and 4fc239f.

📒 Files selected for processing (1)
  • tests/pytest.ini
✅ Files skipped from review due to trivial changes (1)
  • tests/pytest.ini

📝 Walkthrough

Walkthrough

Refactors MasterServer lifecycle: separates shutdown signaling from async cleanup, moves file watching into a coroutine, uses task groups for connection handling, coordinates modes with kota::when_any, preallocates worker pools, and adds sanitizer-aware test assertions for clean server exit.

Changes

Server Shutdown and Async Lifecycle Management

Layer / File(s) Summary
Shutdown signaling and cleanup coordination
src/server/service/master_server.h, src/server/service/master_server.cpp
Header declares file_watcher_task() and shutdown_and_cleanup() coroutines. schedule_shutdown() signals the shutdown event; shutdown_and_cleanup() persists index/cache, concurrently stops indexer and compiler, stops the worker pool, and transitions to Exited state.
File-watching as async task
src/server/service/master_server.cpp
file_watcher_task() coroutine replaces scheduler-based watcher, creates fs event watcher on workspace root, reloads workspace on compile_commands.json changes, and marks supported source/header/module-like file paths as saved.
Connection acceptance with task groups and lazy LSP client
src/server/service/master_server.cpp
accept_connections() uses kota::task_group for accept loop and per-connection spawned tasks, lazily constructs shared LSPClient on first LSP registration, and cleans up connections on disconnect.
Pipe, socket, and daemon mode orchestration
src/server/service/master_server.cpp
Pipe mode conditionally starts agent acceptor (when opts.port > 0), coordinates components with kota::when_any before cleanup. Socket mode wraps accept and shutdown with when_any. Daemon mode introduces daemon_accept() task group and resilient_file_watcher() helper, races them against shutdown-event, and conditionally enables file watching when workspace is provided.
Worker pool preallocation and monitor ordering
src/server/worker/worker_pool.cpp
WorkerPool reserves capacity for stateful/stateless workers before spawning. monitor_worker constructs the worker name earlier and delays binding workers[index] until after awaiting process exit.
Test shutdown validation infrastructure
tests/conftest.py
Introduces SANITIZER_MARKERS constant and _server_stderr_excerpt() helper to filter sanitizer/runtime error lines. New assert_server_exited_cleanly() coroutine awaits server exit, validates exit code, detects errors, and prints filtered stderr. _shutdown_client() delegates clean-exit assertion to helper in finally block.
Integration test updates
tests/integration/agentic/test_agentic.py
test_rpc_shutdown replaces manual polling with assert_server_exited_cleanly() and cleans up background tasks. test_shutdown_during_indexing wraps initialization with try/except, replaces polling with assertion, and uses a finally block to cancel client tasks unconditionally.
Kotatsu dependency & pytest config
cmake/package.cmake, tests/pytest.ini
Updates pinned kotatsu GIT_TAG to a newer commit hash and adds a pytest filterwarnings entry to ignore pygls DeprecationWarning.

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant MasterServer
  participant Acceptor
  participant LSPPeer
  participant Indexer
  participant Compiler
  participant WorkerPool

  Client->>MasterServer: connect / request
  MasterServer->>Acceptor: accept_connections() (task_group)
  Acceptor->>LSPPeer: spawn per-connection task
  Client->>LSPPeer: LSP registration / requests
  MasterServer->>Indexer: coordinate stop (on shutdown)
  MasterServer->>Compiler: coordinate stop (on shutdown)
  MasterServer->>WorkerPool: stop workers (on shutdown)
  MasterServer->>MasterServer: shutdown_and_cleanup() persists index/cache and sets Exited
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

"I nibble at the shutdown bell tonight,
coroutines hum beneath the moonlight,
watchers wake and workers pre-warm,
task groups dance to keep things warm,
tests clap softly when exits are right."

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 16.67% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately and concisely summarizes the main change: refactoring server shutdown to use structured coroutine-based approach with when_any cancellation, replacing the previous loop-based shutdown pattern.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/server-shutdown-asan

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@16bit-ykiko
16bit-ykiko marked this pull request as ready for review June 1, 2026 01:31

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
tests/integration/agentic/test_agentic.py (1)

533-597: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Ensure the server is still torn down on early test failures.

If c.initialize(...) fails on Lines 570-575 while the subprocess is still alive, or anything raises before Line 592, the finally block only cancels client tasks and leaves the spawned server running. That can leak a background server into later tests and mask the original failure.

💡 Suggested fix

Add _shutdown_client to the import on Line 533, then fall back to it unless the explicit clean-exit assertion already ran:

-    from tests.conftest import _find_free_port, assert_server_exited_cleanly
+    from tests.conftest import (
+        _find_free_port,
+        _shutdown_client,
+        assert_server_exited_cleanly,
+    )
     c = CliceClient()
     await c.start_io(*cmd)

+    clean_exit_asserted = False
     try:
         init_options = {
             "project": {
                 "cache_dir": str(workspace / ".clice"),
                 "idle_timeout_ms": 0,
@@
         rpc.sock.close()

         await assert_server_exited_cleanly(c._server, timeout=15.0)
+        clean_exit_asserted = True
     finally:
-        c._stop_event.set()
-        for task in c._async_tasks:
-            task.cancel()
-        await asyncio.sleep(0.1)
+        try:
+            if c._server.returncode is None and not clean_exit_asserted:
+                await _shutdown_client(c)
+        finally:
+            c._stop_event.set()
+            for task in c._async_tasks:
+                task.cancel()
+            await asyncio.sleep(0.1)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/integration/agentic/test_agentic.py` around lines 533 - 597, The test
can leave the spawned server running if initialize() or other code raises before
the clean-exit assertion; import the helper _shutdown_client and in the finally
block (before cancelling tasks) check c._server and whether c._server.returncode
is None, then attempt to call await assert_server_exited_cleanly(c._server,
timeout=15.0) and if that fails or the server is still running call await
_shutdown_client(c) to forcibly stop the server (reference CliceClient,
c._server, initialize, assert_server_exited_cleanly, and _shutdown_client to
locate changes).
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@tests/integration/agentic/test_agentic.py`:
- Around line 533-597: The test can leave the spawned server running if
initialize() or other code raises before the clean-exit assertion; import the
helper _shutdown_client and in the finally block (before cancelling tasks) check
c._server and whether c._server.returncode is None, then attempt to call await
assert_server_exited_cleanly(c._server, timeout=15.0) and if that fails or the
server is still running call await _shutdown_client(c) to forcibly stop the
server (reference CliceClient, c._server, initialize,
assert_server_exited_cleanly, and _shutdown_client to locate changes).

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 3bcf29e7-2d75-4016-a417-7276787aaf62

📥 Commits

Reviewing files that changed from the base of the PR and between cc5b25d and dfeda4d.

📒 Files selected for processing (6)
  • src/server/service/agent_client.cpp
  • src/server/service/lsp_client.cpp
  • src/server/service/master_server.cpp
  • src/server/worker/worker_pool.cpp
  • tests/conftest.py
  • tests/integration/agentic/test_agentic.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/server/service/master_server.cpp (1)

432-435: ⚡ Quick win

Redundant peer closure handlers may cause double-close.

Lines 432-435 schedule a coroutine that closes lsp_peer on shutdown. The new code at lines 460-464 defines close_peer_on_shutdown which does the exact same thing, and it's invoked at lines 469 and 473 within when_all. Both will execute on shutdown, calling peer.close() twice.

Additionally, context from lsp_client.cpp shows LSPClient already calls peer.close() on exit notification, so there could be three close() calls.

Consider removing the pre-existing handler at lines 432-435 since the new when_all orchestration already handles shutdown-triggered closure.

Proposed fix
         kota::ipc::JsonPeer lsp_peer(loop, std::move(final_transport));
         LSPClient lsp_client(server, lsp_peer);
-        loop.schedule([](MasterServer& server, kota::ipc::JsonPeer& peer) -> kota::task<> {
-            co_await server.get_shutdown_event().wait();
-            peer.close();
-        }(server, lsp_peer));

         kota::tcp::acceptor agent_acceptor;

Also applies to: 460-464

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/server/service/master_server.cpp` around lines 432 - 435, Remove the
redundant shutdown closure: delete the earlier loop.schedule(...) lambda that
calls peer.close() (the anonymous coroutine that captures MasterServer& and
kota::ipc::JsonPeer& and awaits server.get_shutdown_event().wait()), and rely on
the new close_peer_on_shutdown helper used inside when_all; ensure only
close_peer_on_shutdown (and LSPClient's exit path) perform peer.close() to avoid
double-close.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@src/server/service/master_server.cpp`:
- Around line 432-435: Remove the redundant shutdown closure: delete the earlier
loop.schedule(...) lambda that calls peer.close() (the anonymous coroutine that
captures MasterServer& and kota::ipc::JsonPeer& and awaits
server.get_shutdown_event().wait()), and rely on the new close_peer_on_shutdown
helper used inside when_all; ensure only close_peer_on_shutdown (and LSPClient's
exit path) perform peer.close() to avoid double-close.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: d8ef350d-5d3e-40c8-80cd-5e9b379de160

📥 Commits

Reviewing files that changed from the base of the PR and between dfeda4d and 42a3d20.

📒 Files selected for processing (1)
  • src/server/service/master_server.cpp

16bit-ykiko and others added 5 commits June 5, 2026 21:46
Replace explicit peer.close()/acceptor.stop() shutdown handlers with
when_any-based cancellation propagation, fix indexer monitor_resources
task tracking with dedicated task_group, and bump kotatsu.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Update kotatsu to ea7d99b which fixes cancel propagation for
reentrantly-cancelled tasks (#158, #160).  Use when_any-based
structured shutdown so all tasks are properly cancelled before
cleanup, avoiding ASAN use-after-free on exit.

Add close_peer_on_shutdown workaround for pipe mode: the sentinel
prevents cancel from reaching pending I/O when shutdown fires
inline from peer.run()'s exit handler.
kotatsu 5e059a2 defers sync primitive resumes to the event loop idle
tick, eliminating inline reentrancy.  cancel() now propagates cleanly
through when_any to the transport io_op, so the close_peer_on_shutdown
workaround is no longer needed.
- Fix lifecycle state: schedule_shutdown sets ShuttingDown (not Exited),
  shutdown_and_cleanup sets Exited after cleanup completes. Guard only
  checks Exited to avoid blocking the exit notification when the LSP
  shutdown request already set ShuttingDown.
- Restore connection cleanup: erase Connection from list when peer
  disconnects to prevent unbounded accumulation.
- Wrap file_watcher_task in resilient_file_watcher so watcher creation
  failure doesn't become a when_any winner that shuts down the daemon.
- Remove unused `this` capture in agent_client shutdown handler.
Picks up stabilized cancellation handling, grant abandonment for
mutex/semaphore, and immediate drain of deferred resumes after the
outermost coroutine resume returns.
@16bit-ykiko
16bit-ykiko force-pushed the fix/server-shutdown-asan branch from 65a5a39 to 549f208 Compare June 8, 2026 14:32
@16bit-ykiko 16bit-ykiko changed the title fix(server): make shutdown sanitizer-clean fix(server): structured shutdown with when_any cancellation Jun 8, 2026
@16bit-ykiko

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 8, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/server/service/master_server.cpp`:
- Around line 339-345: The current run_connection coroutine erases the
Connection but never clears the LSP registration state, causing lsp_registered
to remain true after the owning client disconnects; before calling
connections.erase(pos) in run_connection (and the similar cleanup block at
353-389) check the Connection at *pos for its LSP ownership (e.g., a boolean
like lsp_registered or a non-null LSP client member) and if set, clear the
global/manager flag and reset/unregister the LSP client slot (e.g., null out the
LSPClient pointer or call the Connection/LSP manager unregister method) so
subsequent connections can acquire the LSP slot.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 3a942816-9992-47fb-8926-99dcb731ab92

📥 Commits

Reviewing files that changed from the base of the PR and between ca616ee and 549f208.

📒 Files selected for processing (2)
  • cmake/package.cmake
  • src/server/service/master_server.cpp

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Inline review comments failed to post. This is likely due to GitHub's internal server error or limits when posting large numbers of comments. If you are seeing this consistently it is likely a permissions issue. Please check "Moderation" -> "Code review limits" under your organization settings.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/server/service/master_server.cpp`:
- Around line 339-345: The current run_connection coroutine erases the
Connection but never clears the LSP registration state, causing lsp_registered
to remain true after the owning client disconnects; before calling
connections.erase(pos) in run_connection (and the similar cleanup block at
353-389) check the Connection at *pos for its LSP ownership (e.g., a boolean
like lsp_registered or a non-null LSP client member) and if set, clear the
global/manager flag and reset/unregister the LSP client slot (e.g., null out the
LSPClient pointer or call the Connection/LSP manager unregister method) so
subsequent connections can acquire the LSP slot.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 3a942816-9992-47fb-8926-99dcb731ab92

📥 Commits

Reviewing files that changed from the base of the PR and between ca616ee and 549f208.

📒 Files selected for processing (2)
  • cmake/package.cmake
  • src/server/service/master_server.cpp
🛑 Comments failed to post (1)
src/server/service/master_server.cpp (1)

339-345: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Reset the LSP slot when the owning connection disconnects.

lsp_registered flips to true for the first accepted socket client and is never cleared when that Connection is erased. After that client disconnects, every later connection is built without an LSPClient, so editor reconnects lose all LSP handlers until the server restarts.

Suggested fix
 static kota::task<> run_connection(kota::ipc::JsonPeer* peer,
                                    std::list<Connection>& connections,
-                                   std::list<Connection>::iterator pos) {
+                                   std::list<Connection>::iterator pos,
+                                   bool& lsp_registered) {
     co_await peer->run();
     LOG_INFO("Client disconnected");
+    if(pos->lsp_client) {
+        lsp_registered = false;
+    }
     connections.erase(pos);
 }
 ...
-            group.spawn(run_connection(peer_ptr, connections, it));
+            group.spawn(run_connection(peer_ptr, connections, it, lsp_registered));

Also applies to: 353-389

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/server/service/master_server.cpp` around lines 339 - 345, The current
run_connection coroutine erases the Connection but never clears the LSP
registration state, causing lsp_registered to remain true after the owning
client disconnects; before calling connections.erase(pos) in run_connection (and
the similar cleanup block at 353-389) check the Connection at *pos for its LSP
ownership (e.g., a boolean like lsp_registered or a non-null LSP client member)
and if set, clear the global/manager flag and reset/unregister the LSP client
slot (e.g., null out the LSPClient pointer or call the Connection/LSP manager
unregister method) so subsequent connections can acquire the LSP slot.

Revert indexer monitor task to bg_tasks.spawn() to prevent UAF when
background indexing is cancelled, and improve test_rpc_shutdown to
check for sanitizer errors via assert_server_exited_cleanly().
@16bit-ykiko

Copy link
Copy Markdown
Member Author

@codex

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/integration/agentic/test_agentic.py`:
- Around line 448-451: The cleanup (c._stop_event.set() and cancelling
c._async_tasks) must be executed regardless of assertion failures from
assert_server_exited_cleanly; wrap the await
assert_server_exited_cleanly(c._server) call in a try/finally and move
c._stop_event.set() and the loop that cancels each task in c._async_tasks into
the finally block (mirror the pattern used in test_shutdown_during_indexing) so
background tasks are always stopped even if the assertion raises.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: af953f98-94e2-4c5f-a5b7-d220623e99c7

📥 Commits

Reviewing files that changed from the base of the PR and between 549f208 and 55eed02.

📒 Files selected for processing (1)
  • tests/integration/agentic/test_agentic.py

Comment thread tests/integration/agentic/test_agentic.py
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Chef's kiss.

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

The 3s default was too tight for CI environments, especially Debug
builds where indexer/compiler cleanup takes longer. Bumped to 10s.
@16bit-ykiko
16bit-ykiko merged commit 4c1ca76 into main Jun 8, 2026
20 checks passed
@16bit-ykiko
16bit-ykiko deleted the fix/server-shutdown-asan branch June 8, 2026 16:40
16bit-ykiko added a commit that referenced this pull request Jun 8, 2026
## Summary

- **Structured shutdown via `when_any`**: All server modes (pipe,
socket, daemon) now use `when_any` to race the main transport loop
against `shutdown_event.wait()`. When shutdown is signaled, `when_any`
cancels all sibling tasks — including pending I/O — then runs
`shutdown_and_cleanup()` sequentially. This replaces the old pattern of
fire-and-forget `loop.schedule()` + `loop.stop()`.
- **Bump kotatsu to 2a8c147 (deferred sync resume)**: kotatsu's sync
primitives (`event`, `mutex`, `semaphore`, `cv`) now defer waiter
resumes to the event loop instead of resuming inline. This eliminates
the reentrancy that previously required a sentinel workaround (`child ==
self`) and made cancel unable to propagate through a task whose
coroutine frame was still on the call stack.
- **Sanitizer-clean exit enforcement in tests**: `conftest.py` now
checks every server exit for non-zero return code and sanitizer output
(`AddressSanitizer`, `LeakSanitizer`, etc.) via
`assert_server_exited_cleanly()`. Any sanitizer finding or unclean exit
fails the test.
- **Fix indexer monitor task lifetime**: `run_background_indexing`
previously spawned the resource monitor into `bg_tasks` (a class
member). Now uses a local `task_group` that is joined before returning,
ensuring the monitor is fully stopped before the indexer reports
completion.
- **Fix worker_pool dangling reference**: `monitor_worker` held a
reference to `workers[index]` across `co_await proc.wait()`, but the
vector could reallocate during the wait. The reference is now taken
after the await.

## Design decisions

- **`resilient_file_watcher` wrapper (daemon mode)**:
`file_watcher_task()` co_returns on creation failure. Without wrapping,
this would become the `when_any` winner and trigger daemon shutdown. The
wrapper suspends on `shutdown_event.wait()` after watcher exit, so only
the real shutdown signal terminates the daemon.
- **`schedule_shutdown` guards on `Exited`, not `ShuttingDown`**: The
LSP `shutdown` request sets lifecycle to `ShuttingDown` before the
`exit` notification calls `schedule_shutdown()`. If the guard checked
`ShuttingDown`, the event would never be signaled. `event.set()` is
idempotent, so repeated calls are safe.
- **Accept loop wrapped in `task_group.spawn`**: The accept loop is
spawned as a child of a `task_group` rather than running directly in
`accept_connections`. This ensures `when_any` cancellation propagates
through the task_group to both the accept loop and all active connection
tasks.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * More reliable shutdown with staged cleanup and persisted index/cache
  * Resilient file-watching that reliably reloads workspace changes

* **Improvements**
* Reworked connection handling with per-connection tasks and a single
lazy LSP client
* Coordinated server modes and acceptor behavior; optional daemon
file-watching
  * More predictable worker pool startup and monitoring

* **Tests**
* Stronger test assertions enforcing clean server exit and stderr checks
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant