Skip to content

fix(gateway): bound turn concurrency + boot-resume fan-out (P1 2026-09-21 starvation) - #827

Merged
Kyzcreig merged 4 commits into
mainfrom
fix/gateway-turn-concurrency-cap
Sep 22, 2026
Merged

Kyzcreig merged 4 commits into
mainfrom
fix/gateway-turn-concurrency-cap

Conversation

@Kyzcreig

@Kyzcreig Kyzcreig commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Incident

On 2026-09-21 from 02:52–03:07 PT, the default gateway became effectively deaf for 8–10 minutes per message after three interrupted restarts. A live py-spy capture on pid 99910 measured:

  • 29 simultaneous run_conversation frames in one interpreter
  • 265 process threads, including 73 executor workers
  • 110% CPU and host load 27/32
  • 200–340 second turns, 10-second pre-tool hook timeouts, and ~90-second platform-send lag

Each boot scheduled about 11 restart resumes at once, while Kanban injections, subagent work, and live user turns also entered the unbounded gateway executor.

Summary

  • Add a gateway-wide turn admission controller. New installs default to 8 concurrent turns; a missing raw setting or 0 preserves legacy unbounded behavior.
  • Reserve user capacity by restricting internal gateway turns to max(1, cap - 2) slots.
  • Acquire admission before transcript leases and retain it until a timed-out/cancelled executor worker really exits.
  • Send one interim queue acknowledgment after 15 seconds and emit structured wait/acquire gauges.
  • Bound startup resume bodies to 3 by default while synchronously claiming every resume slot, preserving FIFO order and the existing startup-restore drain deadline.
  • Document the new gateway.max_concurrent_turns and gateway.startup_resume_concurrency settings.
 GatewayRunner
   _handle_message_with_agent
+    TurnAdmission.slot
     SessionTurnLeaseRegistry.acquire
       _run_agent
+        reuse task-owned admission
         run_conversation executor
+        retain permit until worker exit

   _schedule_resume_pending_sessions
     claim every session sentinel
-    create_task(resume) for every entry
+    StartupResumePool.submit (FIFO, concurrency 3)

max_concurrent_sessions remains the separate new-session rejection gate. Nested delegate_task agents remain governed by delegation limits and do not traverse this gateway admission path.

Fail-before proof

Against clean fork/main at d13c5f25ca, the new real-executor regression was copied unchanged into a temporary worktree and run with both parameter variants:

FAILED test_run_agent_bounds_real_executor[2]
AssertionError: executor has no turn admission gate
PASSED test_run_agent_bounds_real_executor[None]

The bounded case therefore fails on the pre-change production path, while the explicit unbounded control preserves legacy behavior.

Verification

HERMES_TEST_FILE_RETRIES=0 nice -n 19 scripts/run_tests.sh -j 2 --file-timeout 900 \
  tests/gateway/test_restart_resume_pending.py \
  tests/gateway/test_turn_concurrency.py \
  tests/gateway/test_startup_restore_gate.py \
  tests/gateway/test_startup_resume_protection_e2e.py \
  tests/gateway/test_startup_resume_interrupt_demotion.py \
  tests/gateway/test_turn_lease.py \
  tests/gateway/test_stop_zombie_turn_lease.py \
  tests/gateway/test_no_atomic_write_reachable_from_loop.py \
  tests/gateway/test_fast_command.py \
  tests/gateway/test_proxy_mode.py

184 passed, 0 failed in 47.1s

Round-3 hardening (7c9072f) — three CI regressions this PR introduced, now closed

Each was reproduced RED on the previous head 966ed208 with the current tests, then GREEN here:

  1. Atomic-write ratchet (test_no_atomic_write_reachable_from_loop.py). The admission split
    renamed two pre-existing reachable roots. The baseline entries were renamed in place
    (_handle_message_with_agent → _handle_message_with_agent_admitted,
    _run_agent_inner → _run_agent_admitted). No entries added; allowlist not widened.
    RED on 966ed208: 1 failed.
  2. Invalid cap normalization (test_proxy_mode.py). _get_turn_admission() fed a non-int
    config value straight into asyncio.Semaphore (TypeError on a MagicMock cap). Both config
    boundaries are now normalized at runtime: only a positive int is bounded; None/0/bool/
    any non-int → unbounded for turns, → 3 for startup_resume_concurrency; a present-but-invalid
    value logs once at WARNING. Covered by
    test_runtime_non_integer_turn_cap_is_unbounded_and_warns_once and
    test_runtime_non_integer_startup_resume_cap_defaults_to_three.
    RED on 966ed208: TypeError: '<' not supported between instances of 'object' and 'int'.
  3. Direct-caller handler contract (test_fast_command.py, 8 failures). The new outer
    _handle_message_with_agent generation guard treated a direct caller with no session state as
    stale and returned None before the existing fail-closed routing ran. The guard now only applies
    when generation state actually exists (_peek_session_state(...) is not None), so queued turns
    invalidated by /stop while waiting for admission are still dropped
    (test_queued_handler_drops_generation_invalidated_while_waiting) while direct/internal callers
    keep the legacy contract (test_direct_handler_without_generation_state_keeps_legacy_contract).
    RED on 966ed208: 8 failed.

None of the three pre-existing test files were edited except the ratchet baseline rename, which the
ratchet's own docstring prescribes.

Also passed:

  • ruff check on all changed Python files
  • python3 -m py_compile on all changed Python files
  • git diff --check

The new tests cover bounded and unbounded real executor behavior, internal/user partitioning, admission-before-lease ordering, delayed one-time user acknowledgment, cancellation retention, FIFO boot-resume throttling for 4 and 7 entries, restore-drain timeout preservation, config coercion/precedence, and ten repeated admission bursts returning permits and ownership state to baseline.

Rollback and deployment

  • Set gateway.max_concurrent_turns: 0 to restore unbounded turn execution.
  • Set gateway.startup_resume_concurrency to a sufficiently large positive value to approximate the prior eager boot-resume fan-out.

This PR does not edit the live ~/.hermes/config.yaml and does not restart the gateway. The fix is not live until Apollo merges it, configures the limits, and reloads the gateway through safe-restart.py. Existing installs must set gateway.max_concurrent_turns; an absent key remains unbounded for legacy compatibility.

Remaining risk

The cap controls gateway-originated turns only; nested delegated agents intentionally remain outside this layer. The default of 8 is a new-install recommendation, while existing raw configs that omit the key remain unbounded until configured.

@Kyzcreig Kyzcreig changed the title fix(gateway): bound turn and boot-resume concurrency fix(gateway): bound turn concurrency + boot-resume fan-out (P1 2026-09-21 starvation) Sep 21, 2026
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Apollo review of PR #827 (head 46e0f57) — NOT accepting yet; review run #4 is in flight, this is a comment.

What holds (verified by me, not from the summary): admission is claimed in _handle_message_with_agent BEFORE SessionTurnLeaseRegistry.acquire (so a slot-waiter can never be a stale_lease_holder) ✅; _run_agent_inner re-enters the same task-owned permit for recursive/queued turns ✅; retain_worker(_executor_task) keeps the permit until a timed-out executor thread actually exits (the 224-thread leak shape) ✅; internal turns capped at max(1, cap-2) so user turns keep a reserved slice ✅; ack reuses the #258 helper with _interim_send metadata ✅; max_concurrent_sessions semantics untouched ✅; fail-before is a real RED (executor has no turn admission gate) with a passing unbounded control ✅.

BLOCKING — CI slice 14/16 is red because of this PR, and it is the exact seam the brief said not to regress.
tests/gateway/test_restart_resume_pending.py — 3 failures: test_reconnect_reschedule_is_platform_scoped (adapter.handle_message awaited 0 times), test_startup_restore_waits_for_resume_before_draining_inbound (seen == [], expected ['resume-start']), test_startup_restore_gate_releases_when_a_resume_turn_runs_long.
Mechanism: StartupResumePool.submit does call_soon(self._pump), so the resume body is not even create_task'd until the NEXT loop tick, and then needs another tick to run. Old code create_task'd immediately. The startup-restore inbound gate and these tests observe the resume between those ticks → gate sees zero started resumes → drains/answers inbound before any resume has begun. That is a real behaviour change in production too (a user message arriving in that window is no longer queued behind the resume), not a test artefact.
Fix: the comment "defer the pump until the synchronous scheduler has claimed ALL slots" is solving a problem that does not exist — the _AGENT_PENDING_SENTINEL claim + _persist_active_agents() for every entry happens synchronously BEFORE submit is called (run.py ~15960). So pump synchronously in submit (create the first concurrency tasks immediately, queue the rest), or have _schedule_resume_pending_sessions call pool.pump() once after its loop. Either way, the first N resumes must be on the loop when _schedule_resume_pending_sessions returns, exactly as before. Re-run tests/gateway/test_restart_resume_pending.py and the 66 you already ran; the 3 must go green without editing them.

Wiring-path finding (mine to close, not yours, but state it in the PR body): max_concurrent_turns defaults to None = unbounded when the raw config omits the key. Ace's production ~/.hermes/config.yaml omits it. So as merged, this fix is INERT on the gateway that had the incident until I add gateway.max_concurrent_turns: 8 before the safe-restart. I'll do that at deploy; add one sentence to the Deployment section saying "existing installs must set the key; absent = unbounded (legacy)". startup_resume_concurrency correctly defaults to 3 regardless — good.

Also confirm in the handback: the 66-test verification line should list test_restart_resume_pending.py in the set once it is green.

Everything else can land as-is. Ping when the head moves.

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Apollo review r2 of PR #827 head 966ed20 — the resume-timing fix is CORRECT and I mutation-proved it myself (reverted only the pump line with the tests kept: 3 failed/48 passed; restored: 78/78 across test_restart_resume_pending + test_turn_concurrency). Good.

But CI on this head has THREE new reds, all caused by this PR, all mechanical. Review run #6 is in flight so this is a comment; fix all three in ONE rework:

  1. slice 13 — tests/gateway/test_no_atomic_write_reachable_from_loop.py::test_reachable_atomic_writes_do_not_grow. This is a ratchet whose baseline is keyed on FUNCTION NAMES (gateway/run.py _handle_message_with_agent -> atomic_json_write, gateway/run.py _run_agent_inner -> atomic_json_write, lines ~112/119). Your split/rename (_handle_message_with_agent_admitted, _run_agent_admitted) makes the same two pre-existing paths appear as two NEW entries. No new atomic writes were added, so the correct fix is to RENAME the two baseline entries to the new function names (the test's own docstring says an entry that disappears must be removed from the baseline). Do not add entries; do not widen the allowlist.

  2. slice 4 — tests/gateway/test_proxy_mode.py::TestRunAgentProxyDispatch::test_run_agent_delegates_to_proxy: TypeError: '<' not supported between instances of 'MagicMock' and 'int'. That fixture sets runner.config = MagicMock(), so getattr(config, "max_concurrent_turns") returns a MagicMock and TurnAdmission(cap) feeds it to asyncio.Semaphore. This is a REAL hardening gap, not just a test problem: _get_turn_admission must coerce the cap — accept only a positive int, anything else (None, 0, non-int, Mock) → unbounded, and log once at WARNING if the value was present but invalid. Same for startup_resume_concurrency in _schedule_resume_pending_sessions (currently getattr(...) or 3 — a MagicMock is truthy and would become the concurrency). Add a test: config attribute is a non-int object → admission is unbounded, pool concurrency is 3.

  3. slice 16 — tests/gateway/test_fast_command.py ×8 (e.g. test_malformed_persisted_identity_roundtrip_fails_closed_before_enrichment[*]). These call runner._handle_message_with_agent(...) directly with a minimal runner. Almost certainly the same MagicMock-cap crash or a missing attribute the new wrapper reads (_adapter_for_source / _thread_metadata_for_source inside the ack lambda, or event.internal). Fix at the source (the wrapper must tolerate a runner without those, and the ack must be lazily resolved only when a wait actually happens), not by editing those tests.

Then: re-run the full tests/gateway/ slice set via CI (push and let it run — do NOT run the whole dir locally on this host), confirm 0 FAILURE across all 16 slices, and update the PR Verification block. Report {head_sha, ci_run_url, slices_green}.

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 21, 2026
@Kyzcreig
Kyzcreig removed this pull request from the merge queue due to a manual request Sep 21, 2026
@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 21, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 21, 2026
Add gateway turn admission with reserved user capacity and retain permits
until executor workers actually exit. Bound startup resume execution while
preserving synchronous session claims and the restore drain timeout.

Verified:
- 66 focused gateway tests passed via scripts/run_tests.sh
- ruff passed on all changed Python files
- git diff --check and py_compile passed
- clean fork/main bounded-executor regression failed while the unbounded
  control passed
Restore the pre-existing create_task timing for the first bounded startup
resume workers so inbound restore handling observes recovery as started.

Verified:
- 117 focused gateway tests passed
- ruff check gateway/turn_admission.py
- python3 -m py_compile gateway/turn_admission.py
- git diff --check
Normalize runtime concurrency values before constructing semaphores, preserve direct handler callers without generation state, and update the atomic-write ratchet for renamed admitted handlers.

Verified:
- 184 focused gateway tests passed across 10 files
- ruff check passed on changed Python files
- py_compile passed on changed Python files
- git diff --check passed
@Kyzcreig
Kyzcreig force-pushed the fix/gateway-turn-concurrency-cap branch from 7c9072f to 17888bd Compare September 22, 2026 00:13
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

PARTIAL — ensemble escalated: healthy floor: 1 of 3 members completed

This review did not reach a trusted verdict, so it is not a gate pass and the findings below may be incomplete. They are posted so they can be read rather than lost in a terminal record.

Reviewed with 1 of 3 model families — anthropic, openai unavailable.

Confidence: 1/5

Findings

  • P1 gateway/turn_admission.py:107 — Broken cancel contract
  • P1 gateway/turn_admission.py:58 — Slot leak on cancel

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $33.62 · duration: 37m 46s · rounds: 2 · files examined: 7

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Apollo r3 on head 17888bd (rebased): FleetReview's two P1s (slot leak on cancel in TurnAdmission.slot; bare-Future cancel contract in StartupResumePool.submit) are CONFIRMED real by reading the code; the red slice-9/14 test_charge_only_after_dispatch[cancel] is the second one. Round-4 rework dispatched (see kanban). Not landing this head.

…me cancel owns its task

Two FleetReview P1s on 17888bd.

P1-A slot leak on cancel. TurnAdmission.slot() is an async generator, so a
raise BEFORE its first yield skips __aexit__ entirely and the pre-yield
region is the ONLY release path. The old code awaited the notice task after
total_acquired=True; that await sat behind ack() -> adapter.send following a
15 s wait, a wide window in which a CancelledError permanently burned one of
`cap` slots. At zero, every turn blocks in acquire forever — the starvation
this gate exists to prevent. Pre-yield now releases internal/total (and
unwinds in_flight/_owners) on ANY BaseException and re-raises; the notice is
cancelled and reaped detached via add_done_callback, never awaited inside the
critical section.

P1-B broken cancel contract. StartupResumePool.submit() handed back a bare
Future, which marks itself done the instant it is cancelled even while the
admitted resume task keeps running. gateway/run.py's shutdown path reads
exactly that done/cancelled state to tell "never started -> cancel + re-mark
the session" from "in progress -> leave it alone", so a running resume could
be re-marked and restored a second time on top of the original turn.
Admitted entries now get _AdmittedResumeHandle, whose cancel() delegates to
the inner task and which stays pending until the task's done callback
resolves it. Queued entries keep the plain Future contract.

Verified (venv python3.11, PYTHONPATH=worktree, HERMES_HOME=/tmp):
- 8-file focused set: 176 passed, 0 failed.
- tests/gateway/test_resume_cap_hardening.py::test_charge_only_after_dispatch
  [cancel] (#801's red test on this head) green, file unedited.
- Mutation-proof P1-A: delete the pre-yield except-BaseException release ->
  test_pre_yield_failure_releases_acquired_permits[False,True] RED
  (total._value 2 != 3, the burned slot).
- Mutation-proof P1-B: submit() back to a bare Future ->
  test_admitted_resume_handle_cancel_delegates_to_its_task RED AND
  test_charge_only_after_dispatch[cancel] RED (assert 2 == 0).
- 4 round-3 guard mutants re-run, all still RED: cap coercion, startup
  coercion, legacy contract, stop-guard.

Class sweep (AST over every .py): exactly 2 async context managers acquire a
permit before yield. The other, hermes_cli/session_db_heavy_gate.py
session_db_heavy_read_slot, has zero awaits between acquire and yield, so it
has no cancellation-delivery point and is not an instance of this class.

ruff check, py_compile, git diff --check all clean.
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Round 4 — both FleetReview P1s closed. Head 043eae7e26. CI 16/16 green.

CI run: https://github.com/ANG-Ventures/hermes-agent/actions/runs/35687032169 — All required checks pass ✅, slices 1..16 all success. Slices 9 and 14 (the red test_charge_only_after_dispatch[cancel]) are green, and that file was not edited.

P1-A — slot leak on cancel (gateway/turn_admission.py TurnAdmission.slot)

slot() is an async generator: a raise before its first yield skips __aexit__ entirely, so the pre-yield region is the only release path. The old finally: notice.cancel(); await asyncio.gather(notice, ...) ran after total_acquired = True, and that await sits behind ack() → adapter.send after a 15 s wait — a wide window in which a CancelledError permanently burned one of cap slots. At zero, every turn blocks in acquire forever: the exact starvation this PR exists to fix, made permanent.

Fix, as specified:

  • the whole pre-yield region is wrapped in except BaseException: which unwinds in_flight/_owners, releases total and internal for whatever was actually acquired, and re-raises;
  • the notice task is never awaited inside the critical section — it is cancelled and reaped detached via add_done_callback(_reap_notice).

P1-B — broken cancel contract (StartupResumePool.submit)

A bare Future marks itself done the instant it is cancelled while the admitted resume task keeps running; gateway/run.py's _cancel_pending_boot_resumes_for_shutdown reads exactly that done/cancelled state to tell "never started → cancel + re-mark the session" from "in progress → leave it alone", so a running resume could be re-marked and restored a second time on top of the original turn.

Fix: admitted entries get _AdmittedResumeHandle(asyncio.Future) whose cancel() delegates to the inner task (returning the task's cancelled state) and which stays pending until the task's own done callback resolves it. future._task is bound before any await point so a cancel racing admission reaches the task rather than completing the handle early. Queued (unadmitted) entries keep the plain Future contract — covered by test_queued_resume_handle_keeps_plain_cancel_contract.

Verification (local: ~/.hermes/hermes-agent/.venv/bin/python 3.11.15, pytest 9.1.1, PYTHONPATH=<worktree>, throwaway HERMES_HOME)

8-file focused set — test_turn_concurrency, test_restart_resume_pending, test_startup_restore_gate, test_startup_resume_protection_e2e, test_turn_lease, test_resume_cap_hardening, test_fast_command, test_proxy_mode:

176 passed, 2 warnings in 25.29s

Mutation proofs — 6 mutants, all RED:

# Mutant Result
P1-A delete the pre-yield except BaseException: release test_pre_yield_failure_releases_acquired_permits[False] + [True] RED — assert 2 == 3 on total._value, i.e. the burned slot itself
P1-B submit() back to a bare create_future() test_admitted_resume_handle_cancel_delegates_to_its_task RED and test_resume_cap_hardening.py::test_charge_only_after_dispatch[cancel] RED (assert 2 == 0) — confirming the fix is what turns slice 9/14 green
r3-1 cap coercion removed RED (test_runtime_non_integer_turn_cap_is_unbounded_and_warns_once / test_proxy_mode)
r3-2 startup coercion → or 3 RED (test_runtime_non_integer_startup_resume_cap_defaults_to_three)
r3-3 legacy contract guard removed RED (test_direct_handler_without_generation_state_keeps_legacy_contract / test_fast_command)
r3-4 stop-guard removed RED (test_queued_handler_drops_generation_invalidated_while_waiting)

Every mutant was reverted and the file byte-compared back to the original before the next one.

Class sweep (P1-A's shape, not just its site)

AST pass over every .py in the tree for async context managers that acquire a permit before their first yield: exactly 2 sites. gateway/turn_admission.py::slot (acquire@68, yield@100, an await at 70 between them — the bug) and hermes_cli/session_db_heavy_gate.py::session_db_heavy_read_slot (acquire@136, yield@167, zero awaits between them). Cancellation is only delivered at an await, so the second site has no delivery point and is not an instance of this class. No other file needs the fix.

Also clean: ruff check, py_compile, git diff --check on both changed files. Only gateway/turn_admission.py and tests/gateway/test_turn_concurrency.py are touched this round; no pre-existing test file was edited.

Note for deploy (unchanged from r2): gateway.max_concurrent_turns must be SET in ~/.hermes/config.yaml — absent = unbounded = this fix is inert.

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: t_6745cd39 · gate: BYPASS: no FleetReview run exists for 043eae7 and the router is degraded (t_8d3c4eeb); argus artifact review stands in · why: argus r1 APPROVED (artifact lens): both FleetReview P1s (slot leak on cancel; bare-Future resume cancel contract) REPRODUCED on pre-fix head 17888bd and closed on 043eae7; 9 mutants RED incl #801's [cancel] test; CI 16/16 + 'All required checks pass'. FleetReview has no record and nothing queued for this head (router at ~50% escalated, t_8d3c4eeb).

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: t_6745cd39 · gate: BYPASS: no FleetReview run exists for 043eae7; router degraded (t_8d3c4eeb); argus artifact review stands in · why: argus r1 APPROVED (artifact lens): both FleetReview P1s reproduced pre-fix + closed on 043eae7; 9 mutants RED; CI 16/16. No FR run exists for this head; router degraded (t_8d3c4eeb).

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 22, 2026
Merged via the queue into main with commit 2334f26 Sep 22, 2026
54 checks passed
@Kyzcreig
Kyzcreig deleted the fix/gateway-turn-concurrency-cap branch September 22, 2026 05:27
Kyzcreig pushed a commit that referenced this pull request Sep 22, 2026
… raises

session_db_heavy_read_slot is an @asynccontextmanager that acquires its
semaphore permit before the first yield. A generator that raises before
its first yield never runs __aexit__, so the try/finally around the yield
never executes and the pre-yield window is the ONLY release path. Any
raise in that window burned a permit permanently; cap repeats killed the
gate and every dashboard/TUI session-list read shed
SessionDBHeavyReadBusy until the process restarted.

Second member of the class PR #827 fixed at gateway/turn_admission.py.
#827's sweep cleared this site with an await-only discriminator ("a
cancel is only delivered at an await, and this window has none") — but
#827's own regression test injects a SYNCHRONOUS raise, so the class is
"the pre-yield region raises for ANY reason".

Reproduced on the live file before fixing (cap=2, _record_stats injected
to raise): RAISE arm final=1 entered_body=False; STARVATION arm final=0
healthy_caller_served=False. After the fix: final=2 / final=2 / served.

Also lands the enforcement mechanism, not just the patch:
scripts/check_preyield_permit_release.py asks the right question — "does
any statement between the acquire and the first yield have a raise path
(Call/Attribute/Subscript/Await), and is it covered by a try that
releases?" — so a third such context manager cannot land unguarded. Its
DOES NOT COVER section states the boundaries (acquire-failure path,
release machinery, alias releases) rather than hiding them.

Verified:
- 12/12 tests/scripts/test_preyield_permit_release.py pass
- mutants, each restored byte-identical after:
  A release deleted from the new guard          -> RED (3 tests)
  B _RAISE_CAPABLE narrowed to (ast.Await,)     -> RED (3 tests)
  C except Exception accepted as sufficient     -> RED (1 test)
  D decorator discovery broken                  -> RED (7 tests)
- guard run against gateway/turn_admission.py at 043eae7 (post-#827)
  -> PASS; at its pre-fix parent -> FAIL at the known leak line
- ruff clean on all three files
- tests/test_web_server_sessiondb_eventloop.py + tests/scripts/ :
  89 passed, 2 failed — both failures reproduce on the unmodified base
  (ramscratch HERMES_HOME sandbox-guard + manifest nodeid collection),
  unrelated to this change
Kyzcreig pushed a commit that referenced this pull request Sep 22, 2026
… raises

session_db_heavy_read_slot is an @asynccontextmanager that acquires its
semaphore permit before the first yield. A generator that raises before
its first yield never runs __aexit__, so the try/finally around the yield
never executes and the pre-yield window is the ONLY release path. Any
raise in that window burned a permit permanently; cap repeats killed the
gate and every dashboard/TUI session-list read shed
SessionDBHeavyReadBusy until the process restarted.

Second member of the class PR #827 fixed at gateway/turn_admission.py.
#827's sweep cleared this site with an await-only discriminator ("a
cancel is only delivered at an await, and this window has none") — but
#827's own regression test injects a SYNCHRONOUS raise, so the class is
"the pre-yield region raises for ANY reason".

Reproduced on the live file before fixing (cap=2, _record_stats injected
to raise): RAISE arm final=1 entered_body=False; STARVATION arm final=0
healthy_caller_served=False. After the fix: final=2 / final=2 / served.

Also lands the enforcement mechanism, not just the patch:
scripts/check_preyield_permit_release.py asks the right question — "does
any statement between the acquire and the first yield have a raise path
(Call/Attribute/Subscript/Await), and is it covered by a try that
releases?" — so a third such context manager cannot land unguarded. Its
DOES NOT COVER section states the boundaries (acquire-failure path,
release machinery, alias releases) rather than hiding them.

Verified:
- 12/12 tests/scripts/test_preyield_permit_release.py pass
- mutants, each restored byte-identical after:
  A release deleted from the new guard          -> RED (3 tests)
  B _RAISE_CAPABLE narrowed to (ast.Await,)     -> RED (3 tests)
  C except Exception accepted as sufficient     -> RED (1 test)
  D decorator discovery broken                  -> RED (7 tests)
- guard run against gateway/turn_admission.py at 043eae7 (post-#827)
  -> PASS; at its pre-fix parent -> FAIL at the known leak line
- ruff clean on all three files
- tests/test_web_server_sessiondb_eventloop.py + tests/scripts/ :
  89 passed, 2 failed — both failures reproduce on the unmodified base
  (ramscratch HERMES_HOME sandbox-guard + manifest nodeid collection),
  unrelated to this change
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant