Repository navigation
feat(tui_gateway): async-side shared heavy-read bound for session.list (seam a) - #215
Conversation
Share a loop-local async SessionDB heavy-read gate between dashboard REST reads and WebSocket session-list RPCs. The gate reads dashboard.heavy_read_max_concurrency lazily, queues with a bounded wait, logs queue_wait, exposes shed stats, and returns retryable backend-busy errors on saturation.\n\nVerified:\n- scripts/run_tests.sh tests/test_web_server_sessiondb_eventloop.py -v (14 passed)\n- scripts/run_tests.sh -j 16 tests/tui_gateway tests/test_tui_gateway*.py tests/test_web_server*.py tests/hermes_cli/test_web_server*.py (1133 passed)\n- git diff --check\n- /Users/alexgierczyk/.hermes/hermes-agent/venv/bin/python -m py_compile hermes_cli/session_db_heavy_gate.py hermes_cli/web_server.py tui_gateway/server.py tui_gateway/ws.py\n\nNot pushed.
| # drained. Replacing while permits are checked out would let old and | ||
| # new semaphores admit work simultaneously and temporarily exceed | ||
| # both bounds. | ||
| if getattr(existing_semaphore, "_value", 0) < existing_limit: | ||
| return existing_semaphore | ||
| semaphore = asyncio.Semaphore(limit) |
There was a problem hiding this comment.
Queued Waiters Cross Gate Swap
When the config value changes while tasks are already waiting on the old semaphore, _value can still look fully drained because queued acquire() calls have not decremented it yet. This replaces the loop entry with a new semaphore, so new requests can enter through the new gate while old waiters later enter through the old gate, temporarily exceeding the configured heavy-read bound.
| def session_db_heavy_read_stats() -> dict[str, float | int]: | ||
| with _STATS_LOCK: | ||
| stats = dict(_STATS) | ||
| stats["max_concurrency"] = _configured_max_concurrency() | ||
| stats["queue_timeout_seconds"] = _QUEUE_WAIT_TIMEOUT_S |
There was a problem hiding this comment.
session_db_heavy_read_stats() now runs inside the async status endpoint, but it calls _configured_max_concurrency(), which synchronously loads dashboard config. Frequent status polling or a slow config filesystem can block the event loop and delay unrelated REST or websocket work.
| try: | ||
| async with session_db_heavy_read_slot( | ||
| surface="ws", | ||
| operation=str(req_method), | ||
| ): | ||
| return await asyncio.to_thread(server.handle_request_bound, req, transport) |
There was a problem hiding this comment.
For heavy methods that are also long handlers, this path bypasses server.dispatch() and runs the full handler in asyncio.to_thread() instead of the gateway RPC pool. With a small or shared default executor, two slow session.list or projects.tree scans can occupy the default executor and delay otherwise cheap websocket offloads.
| """ | ||
| t = transport or _stdio_transport | ||
| token = bind_transport(t) | ||
| try: |
There was a problem hiding this comment.
Handler Errors Become Dispatch Crashes
handle_request_bound() calls handle_request() without the exception mapping that dispatch() applies for long-handler workers. If a gated websocket handler raises, clients now receive a generic -32603 internal error from handle_ws instead of the previous handler error response, and telemetry counts it as a dispatch crash.
* Revert "fix(state): force recency backfill v3 (#216)" This reverts commit cd05a73. * Revert "feat(tui_gateway): bound heavy session reads (#215)" This reverts commit 66d98d3. * Revert "feat(state): denormalize session.list recency (effective_last_active + two-stage query) (#213)" This reverts commit 5d67b3f.
… 3 premises re-checked (t_2a1bd9cd) #115 DROP->KEEP (swiftui-skills skill is live), #215/#441 DROP->UNRESOLVED (tui_gateway/ws.py still consumes the gate; dashboard live), nopr:8a8b81638c UPSTREAM->UNRESOLVED (leak needs the fork-only auto-attach detector). Branches: 5 revert branches built + targeted pytest green; upstream-617 and upstream-466 hand-ported onto upstream/main with tests green.
Phase 2 of docs/desktop/2026-07-05-session-list-loop-starvation-PRD.md (D-4/D-5/INV-4/INV-6). Shared per-loop async SessionDB heavy-read gate (default 2, config dashboard.heavy_read_max_concurrency, lazy/loop-bound NOT import-frozen). REST + WS list_sessions_rich (and siblings) acquire the SAME semaphore async-side BEFORE asyncio.to_thread dispatch (seam a — not inside the worker pool, so cheap ws ops never starve). Saturation: bounded wait + queue_wait log + shed stats + retryable backend_busy (503/JSON-RPC). Dashboard status exposes session_db_heavy_reads stats. 349 gateway tests green (Apollo re-ran per-file runner). Reviewed by Apollo. Live K=8 incident-regime certify = Apollo post-merge.