test: real multi-process E2E suites for state.db integrity, compaction, liveness, history replay and entrypoint parity - #120171
Merged
Merged
Conversation
teknium1
force-pushed
the
tests/core-e2e
branch
from
September 23, 2026 13:09
6a32ff3 to
2cf408e
Compare
This was referenced Sep 23, 2026
teknium1
force-pushed
the
tests/core-e2e
branch
from
September 23, 2026 16:50
2cf408e to
f601e22
Compare
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…ced handler errors, no-retry CI) Independent review of #120171 found checks that could not fail. Each is now proven red by a mutation that the old version reported as XFAIL or pass. - chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used raises=AssertionError and RpcError subclasses it, so a gateway crash counted as the expected failure. Every invariant is now asserted normally; only the known leftovers (surviving tool tree and the tool_call it leaves without a result, both fixed by #120306) raise ToolOutlivedGateway, the only exception the xfail accepts. The DB check used to sit behind the orphan assert and never ran; running it exposed the dangling tool_call half of the same bug. - history/test_prefix_stability: surface_switch's strict xfail tripped at the first prefix break, before usage and integrity. Messages/system prompt, usage and integrity are asserted first; the tools-array drift is checked last and raises ToolsArrayDrift, the only exception the xfail accepts. - history/test_transcript_ledger: scripted steer/interrupt callables run on the fake provider's handler thread, where an assert only dropped the connection. Script records those failures and the test re-raises them after every turn; steer must land and the interrupted turn must report interrupted=True within 30 s. - fakes/fake_llm_provider: Hang drops the connection at its deadline instead of leaving a kept-alive client waiting past it. - parity: the API server port was picked, released, then bound by the child. Readiness now requires our child's pid from authenticated /health/detailed and retries on a fresh port when the child reports it in use. The fixture guard refused any HERMES_HOME under ~/.hermes, failing all parity tests whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root or a real profile. - chaos/_gateway_harness: the gateway stays in pytest's process group, so the runner's kill of a timed-out file reaches it. - sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable constructor failed: messages_fts"; the delete arm's busy tolerance keys on the result code. A failed episode's roles are stopped so the shared chamber and rig no longer fail every later episode. - chaos, compaction, parity homes: updates.check=false (history already had it). The passive update check made a GitHub round-trip from every test surface, and on a blobless clone whose objects lag upstream its `git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that the 5 s timeout orphans; the orphan scans then failed on git processes. - chaos/test_agent_turn_liveness: a PROBE failure now carries the provider call counts and the agent's stale-kill log, so a cross-turn breaker trip can be told apart from a slow probe. - tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never retried into green; own uv cache entry (cache-suffix: e2e).
૮ >ﻌ< ა ci reviewran on 5535d03 — test(e2e): micro-compaction display check is a plain test no ❌ Job failuresJS & TS checks / JS & TS checks · View jobJob JS & TS checks / JS & TS checks failed. ℹ️ InfoCI-sensitive file review · View jobPR touches sensitive files, but the Sensitive files changed: debug infoCI timingsCI timings · View report · View jobWall time 5m47s vs 9m41s (-40.3%). 9 job(s) slower, 10 faster,
|
teknium1
force-pushed
the
tests/core-e2e
branch
from
September 23, 2026 17:52
e658a83 to
5535d03
Compare
Class C1: state.db corruption, WAL generations unlinked under a live holder, lost/duplicated rows, repair destroying data, fd leaks. Every role is a real OS process on one WAL state.db through the production SessionDB: gateway- and TUI-like writers with intent/ack journals, a dashboard reader living across episodes, short-lived openers (SessionDB, bare sqlite3, real `hermes sessions list/stats`), FTS rebuild/optimize and repair_state_db_schema. Nine seeded episodes inject kill -9 mid-write, SIGTERM graceful close, POSIX lock cancellation by a stray in-process open/close, chmod flips, concurrent and killed FTS rebuilds, repair against a live and an offline store, FTS corruption and a whole-fleet SIGKILL, then assert the same invariants: integrity_check ok and still WAL, no child holding a (deleted) db/-wal/-shm (/proc fd monitor), every acked append stored exactly once, counts only grow, repair never lowers them, FTS == canonical rows and search finds acked rows once, no role errors, bounded fds for the long-lived reader and churner. Red-proof: disabling the OFD WAL lock guard (75e155a, whose revert no longer applies cleanly) makes the lock_cancellation episode kill the gateway writer with SIGBUS / fail sibling openers, 3/3 runs.
Class C1, "state.db corrupting from compaction": real AIAgent processes run real turns (user -> terminal tool call -> answer) through the loopback fake provider on one shared WAL state.db while a plain writer, a reader and an open/close churner work the same file. A gateway-like agent micro-compacts every turn; a TUI-like agent runs /compress here 2; faults are clean, kill -9 mid-turn and kill -9 timed into the micro-compaction summary/commit window; a fresh process then resumes the session. Invariants: every acked user turn live exactly once, its tool result and answer recoverable (live or compacted history) and never live twice, the /compress-kept exchanges live exactly once, canonical rows never shrink (external sampler + reader), integrity ok, FTS == canonical, resumed request carries every acked user turn exactly once, plain-writer appends exactly once. Red-proof: reverting 91df541 (micro-compaction tool rows) turns every absorbed tool result/answer into active=0,compacted=0 debris; reverting 526d135 (/compress here tail) leaves the kept exchanges with no live row.
… the fake provider Backward compatible: FakeLLMServer(prompt_tokens_fn=...) reports usage.prompt_tokens computed from each request body (token-driven logic such as compaction triggers sees a realistic, growing count instead of a constant 100), and Text(finish_reason=...) lets a scripted reply end with e.g. 'length' (truncated summaries). Defaults are unchanged.
Class C4 (context compression correctness/liveness). A real AIAgent + SessionDB on disk drives seeded random sessions (tool-pair heavy incl. parallel calls, huge single turns, image parts, long answers, single tool results larger than the whole threshold, a cron-shaped single job turn with 14 tool rounds, a provider whose real window is smaller than the configured one) against the recording fake provider. After every turn and for every request it asserts: bounded wall time and summarizer calls, all scripted steps consumed or an explicit error, no user message lost from state.db (live or compacted=1), superseded rows always have a live/ archived copy, tool_call/tool_result pairs matched in live rows and in every request, system prompt stable, the in-flight ask (cron job prompt) after every summary (#100818), and persisted == sent: state.db at request time is exactly the request's history prefix, live rows equal the in-memory history, and a fresh agent resumed from state.db continues identically. Exposed the stale current_turn_user_idx bug fixed in the previous commit.
Same real chain and per-turn invariants, with the summarizer answering empty / refusal / truncated (finish_reason=length) / slower than the compression timeout / HTTP 500. A bad summary must never archive history (P0 #94448 empty-summary deletes the middle; refusal accepted as a summary), never reach the model or state.db, must be surfaced to the user, and must not loop (877-summarizer-call class). A 500 may commit only the designed deterministic fallback, never with abort_on_summary_failure.
… == sent (C4) Drives manual compaction through the TUI/Desktop choke point (tui_gateway _compress_session_history -> compress_now) and rolling micro-compaction through the real turn finalizer, then asserts no user message lost, live rows == in-memory history, kept exchanges stay live verbatim rows, failed summaries change nothing, and the next turn from memory AND from a fresh agent resumed off state.db both send exactly the persisted history. Covers the 526d135 (/compress here N tail lost from state.db) and 91df541 (micro-compaction marks summarized rows rewind-only) data-loss class.
… (C4) A real child process runs /compress on a seeded session's state.db; the parent SIGKILLs it while the summary request is in flight and inside the archive_and_compact write transaction (parked by a child-side hook after archive + inserts, before COMMIT). Asserts integrity_check ok, rows byte-identical to pre-compaction, message_count consistent, then a fresh agent reclaims the dead holder's lease, compacts and continues with persisted == sent.
… step) Class C3 (conversation-history persistence & replay integrity). A real AIAgent + SessionDB runs a scenario matrix (plain turns, parallel tool batches, /steer mid-turn, interrupt, stream drop + 500 retry, manual /compress, /compress here N, auto compaction, opt-in micro-compaction) against the recording fake provider. After EVERY step it asserts: exactly-once at the storage layer and in the model view, next request replays the persisted view, nothing the model saw ever leaves the durable/display history, each typed input shown once, rows monotonic, usage == provider-billed; then a FRESH 'hermes chat --resume' process must replay a byte-identical request prefix. Red-proofs: reverting 526d135 (compress-here tail) and 91df541 (micro-compaction rewind flags) and an injected double row insert all go red. A strict xfail documents a live micro-compaction bug (merged user row duplicates both inputs in the resumed display). The fake provider now records the usage it billed per request (record['usage']).
Class C17 (prompt-cache prefix byte-stability + usage/cost accounting). One durable session is driven for 10+ turns through FRESH processes of the real entrypoints (tui_gateway stdio JSON-RPC server, 'hermes chat -q --resume' oneshot), flipping cwd between hops, with a parallel tool batch and a manual session.compress. Asserts every main request is a byte-identical extension of the previous one (tools array, system prompt, earlier messages) with exactly one sanctioned break at the compaction, every process boundary replays the persisted model view, storage rows are exactly-once and monotonic, and sessions / session_model_usage equal what the fake provider billed. tui_gateway_restarts went red on base (the gateway's manual /compress persisted a prompt with the process HOME as cwd; fixed in the previous commit). surface_switch is a strict xfail documenting live TUI<->oneshot tools-array divergence (tool_search catalog rebuilt per surface; pinned tools lose dynamic schema overrides; -q --resume prunes skill_manage).
One fixture HERMES_HOME (shell + plugin pre_llm_call hooks, AGENTS.md, a skill, a memory entry, a stdio MCP server that spawns a grandchild, the recording fake provider, a disabled toolset) and one scripted turn per entrypoint: hermes -z, hermes chat -q, tui_gateway stdio, hermes serve over the Desktop WS, GatewayRunner with a fake adapter, api_server, ACP stdio, cron run-now. The fake model calls the MCP canary tool, then answers. Same invariants on every surface, first turn after a cold start: context file / skill index / memory in the recorded system prompt, both hooks fired and injected, MCP tool offered and a REAL call returned the canary (the inverted stdio-liveness burst failed exactly here everywhere), documented toolset's feature tools present and the disabled toolset absent, answer delivered to the client, zero MCP server/grandchild survivors after the surface's normal shutdown. Catches the 'works in the CLI, missing in Desktop/gateway/ACP/cron' class; found #95577 (oneshot row). ACP's disabled_toolsets gap (#74582) is a strict known-red cell.
Spawns hermes serve exactly as the Desktop does and tui_gateway.entry as the Ink TUI does. serve: the first stdout line must be the READY sentinel, resolved to the live port by the Desktop's REAL parser (apps/desktop/electron/backend-ready.ts run under Node, no regex copy), and the port answers an authenticated /api/status. tui_gateway: the first stdout frame is a contract-valid gateway.ready and stdout stays pure JSON-RPC. Found setup.ready leaking onto serve stdout before READY.
… (C15) tui_gateway host + the real stdio MCP fixture: - fast death: the server crashes mid-call leaving a helper that inherited its stdio (npx/uvx wrapper shape). The call must surface as an outcome-uncertain tool error exactly once (never replayed onto the respawned server), the host survives, the turn completes within a bound 12x under the configured call timeout, and shutdown reaps the helper. - host crash: SIGKILL the host after a turn; the death supervisor must still reap the MCP server and its grandchild. Also pins mcp_discovery_timeout in the fixture home: interactive surfaces wait only ~1.5 s for discovery by design, which a loaded box exceeds, so the first-turn-complete invariant needs the documented knob raised.
…veness) Class C8 (agent-turn liveness): hangs, runaway retries, lost tool results, wedged agents, orphan processes. A real AIAgent + SessionDB runs in a child process with config.yaml liveness knobs (agent.api_max_retries, providers.custom request/stale timeouts, agent.max_turns) against the scripted loopback provider. 22 fault modes (hang, reasoning-model stall, 500/429/400/garbage forever, drop mid-stream, hang then recover, slow trickle, malformed/unknown tool calls, 8-way parallel, hung/stdin/huge tools, background swarms, endless tool loop, interrupts) are each followed by a probe turn on the same agent and checked against the same invariants: bounded termination, bounded provider calls, reusability with history kept, every tool_call answered (in state.db and on the wire), no surviving tagged process, integrity_check ok.
python -m gateway.run with a fake platform adapter: each fault mode must end the turn within its deadline (or promptly on /stop / SIGTERM), keep provider calls bounded, release the turn lease so the session answers the next message (#104303 class), keep the event loop responsive (heartbeat), answer every tool_call (#93251 class) and leave no orphan process.
…server python -m tui_gateway.entry over stdio: every fault mode ends with a terminal message.complete within its deadline (interrupt within seconds), the dispatcher keeps answering during a wedged turn, the busy flag is released so the same session streams the next prompt, tool results are complete, and exit on EOF/SIGTERM mid-turn leaves no orphans (the foreground-tool orphan on exit is a known base bug, strict xfail).
…s per-file budget
…ythons without os.pidfd_open CI's pinned uv only knows CPython 3.11.14, whose bundled SQLite has the WAL-reset bug, so Hermes ran state.db in DELETE mode and every torture-chamber episode skipped (green over zero coverage). The chaos cleanup also called os.pidfd_open, which that build lacks, turning two strict xfails into errors.
…DELETE journal mode too Both suites skipped outright on a WAL-reset-vulnerable SQLite, so the DELETE-mode population (every user whose Python bundles SQLite 3.7.0-3.51.2, e.g. uv CPython 3.11.14's 3.50.4, and the network/FUSE homes Hermes also falls back to DELETE on) had no multi-process integrity coverage. - Parametrize both module fixtures over journal mode [wal, delete]. The delete arm goes through the production decision path: each child (_roles.py, including the `hermes` CLI run as __main__) pins the version hermes_state_wal.is_sqlite_wal_reset_vulnerable() reports to 3.50.4, and apply_wal_with_fallback picks DELETE itself; the harness never issues a journal-mode pragma. The wal arm skips only where Hermes would not run WAL; the delete arm runs everywhere. - Every invariant that holds in both modes stays on in both: integrity_check, acknowledged writes exactly once, monotonic counts, repair never lowers rows, FTS parity, fd bounds, persisted == replayed. New in both: the store is still in the arm's mode and every SessionDB role reported the matching _wal_active (vacuity guard for the seam). WAL-only: the deleted -wal/-shm fd scan (the deleted main-file scan runs in both). - DELETE mode blocks readers on writes and has no writer fairness: in that arm a SQLITE_BUSY refusal of a read/open/FTS pass is waited out and counted (reported in failure context), and plain writers are paced by 20 ms; in the wal arm busy still fails the role. - chmod_flip runs 3-5 flips instead of 6-10 (same fault, process start-up dominated the time). Red-proof (delete arm red, wal arm green): repair without live-writer checks/exclusion -> "acked ... stored 0x"; a DELETE-mode commit misreported busy so the retry re-runs the insert -> "acked ... stored 2x" (torture and compaction); default config enabling WAL on a vulnerable SQLite -> "journal_mode is 'wal' (arm is delete)" + "_wal_active=True in the delete arm".
…ced handler errors, no-retry CI) Independent review of #120171 found checks that could not fail. Each is now proven red by a mutation that the old version reported as XFAIL or pass. - chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used raises=AssertionError and RpcError subclasses it, so a gateway crash counted as the expected failure. Every invariant is now asserted normally; only the known leftovers (surviving tool tree and the tool_call it leaves without a result, both fixed by #120306) raise ToolOutlivedGateway, the only exception the xfail accepts. The DB check used to sit behind the orphan assert and never ran; running it exposed the dangling tool_call half of the same bug. - history/test_prefix_stability: surface_switch's strict xfail tripped at the first prefix break, before usage and integrity. Messages/system prompt, usage and integrity are asserted first; the tools-array drift is checked last and raises ToolsArrayDrift, the only exception the xfail accepts. - history/test_transcript_ledger: scripted steer/interrupt callables run on the fake provider's handler thread, where an assert only dropped the connection. Script records those failures and the test re-raises them after every turn; steer must land and the interrupted turn must report interrupted=True within 30 s. - fakes/fake_llm_provider: Hang drops the connection at its deadline instead of leaving a kept-alive client waiting past it. - parity: the API server port was picked, released, then bound by the child. Readiness now requires our child's pid from authenticated /health/detailed and retries on a fresh port when the child reports it in use. The fixture guard refused any HERMES_HOME under ~/.hermes, failing all parity tests whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root or a real profile. - chaos/_gateway_harness: the gateway stays in pytest's process group, so the runner's kill of a timed-out file reaches it. - sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable constructor failed: messages_fts"; the delete arm's busy tolerance keys on the result code. A failed episode's roles are stopped so the shared chamber and rig no longer fail every later episode. - chaos, compaction, parity homes: updates.check=false (history already had it). The passive update check made a GitHub round-trip from every test surface, and on a blobless clone whose objects lag upstream its `git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that the 5 s timeout orphans; the orphan scans then failed on git processes. - chaos/test_agent_turn_liveness: a PROBE failure now carries the provider call counts and the agent's stale-kill log, so a cross-turn breaker trip can be told apart from a slow probe. - tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never retried into green; own uv cache entry (cache-suffix: e2e).
…20316 landed Drop the strict xfail on the resumed-display test and re-enable the input check for the micro_compaction ledger scenario. The ledger no-loss oracle now exempts model-only rows (the merged user turn #120316 flags model_only): they stand in for originals that stay displayed and were each tracked when sent. A/B: removing the model_only flag in agent/micro_compaction.py turns both micro cases red; with it, the history suites are green.
teknium1
force-pushed
the
tests/core-e2e
branch
from
September 23, 2026 18:10
5535d03 to
276f56e
Compare
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…ced handler errors, no-retry CI) Independent review of #120171 found checks that could not fail. Each is now proven red by a mutation that the old version reported as XFAIL or pass. - chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used raises=AssertionError and RpcError subclasses it, so a gateway crash counted as the expected failure. Every invariant is now asserted normally; only the known leftovers (surviving tool tree and the tool_call it leaves without a result, both fixed by #120306) raise ToolOutlivedGateway, the only exception the xfail accepts. The DB check used to sit behind the orphan assert and never ran; running it exposed the dangling tool_call half of the same bug. - history/test_prefix_stability: surface_switch's strict xfail tripped at the first prefix break, before usage and integrity. Messages/system prompt, usage and integrity are asserted first; the tools-array drift is checked last and raises ToolsArrayDrift, the only exception the xfail accepts. - history/test_transcript_ledger: scripted steer/interrupt callables run on the fake provider's handler thread, where an assert only dropped the connection. Script records those failures and the test re-raises them after every turn; steer must land and the interrupted turn must report interrupted=True within 30 s. - fakes/fake_llm_provider: Hang drops the connection at its deadline instead of leaving a kept-alive client waiting past it. - parity: the API server port was picked, released, then bound by the child. Readiness now requires our child's pid from authenticated /health/detailed and retries on a fresh port when the child reports it in use. The fixture guard refused any HERMES_HOME under ~/.hermes, failing all parity tests whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root or a real profile. - chaos/_gateway_harness: the gateway stays in pytest's process group, so the runner's kill of a timed-out file reaches it. - sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable constructor failed: messages_fts"; the delete arm's busy tolerance keys on the result code. A failed episode's roles are stopped so the shared chamber and rig no longer fail every later episode. - chaos, compaction, parity homes: updates.check=false (history already had it). The passive update check made a GitHub round-trip from every test surface, and on a blobless clone whose objects lag upstream its `git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that the 5 s timeout orphans; the orphan scans then failed on git processes. - chaos/test_agent_turn_liveness: a PROBE failure now carries the provider call counts and the agent's stale-kill log, so a cross-turn breaker trip can be told apart from a slow probe. - tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never retried into green; own uv cache entry (cache-suffix: e2e).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Real multi-process E2E suites for the failure classes users keep hitting: state.db corruption, compaction losing or duplicating history, hung turns, history that doesn't replay, and features that work in one entrypoint but not another. They found 6 production bugs, each landed as its own PR (#120175, #120178, #120161, #120158, #120152, #120176). This branch is rebased on those fixes and carries tests and CI only.
Follow-up to #120071 (the low-value test purge). The Desktop transcript-integrity suite (the duplicate-message class) and the remaining lanes (tenancy, delivery, upgrade, live-provider, terminal) will be added to this PR as they finish; see Pending below.
What's in it
Every suite drives real Hermes code in real OS processes:
AIAgent,SessionDBon a real WALstate.db,hermes serve, the messaging gateway, the TUI backend (tui_gateway),hermes chat --resume, cron, ACP and a real stdio MCP server. Only the model is faked.tests/fakes/fake_llm_provider.pyis a scripted, recording loopback server that speaks OpenAI chat-completions and can inject faults (stall mid-stream, 5xx, 429 with Retry-After, truncated or refused summaries, raw bodies). Each suite checks invariants after every step, not a single happy-path outcome.tests/e2e/core/…)sqlite/test_torture_chamber.py,test_compaction_contention.pyintegrity_checkok and still WAL; every acknowledged write stored exactly once; row counts never go down; repair never loses rows; FTS matches the real rows; no process holds a deleted db/-wal/-shm file; open file counts stay flat. Faults: kill -9, SIGTERM, lock cancellation, chmod flips, killed FTS rebuilds, repair on a live then an offline DB, killing every process at once.compaction/(4 files, 34 tests)/compress,/compress here N, micro-compaction, kill -9 mid-summary and mid-commit. Every user message stays live or archived; tool calls and results stay paired in the DB and in every request; the DB matches the request exactly; resume continues identically; bounded summarizer calls and turn time.history/test_transcript_ledger.py,test_prefix_stability.py--resumeprocess sends a byte-identical prefix. Tools array, system prompt and history stay byte-stable across TUI, oneshot and working-directory hops, with one allowed break at compaction. Usage matches what the provider billed.chaos/(agent, gateway, tui_gateway)parity/chat -q, tui_gateway, serve over WS, gateway, api_server, ACP, cron run-now). The READY handshake is parsed by Desktop's realbackend-ready.tsunder Node. MCP server death, restart and replay transitions.Production bugs found (each merged as its own PR; this branch carries no production code)
current_turn_user_idxstale, so the next request dropped this turn's earlier tool calls and results while state.db kept them.hermes -zignored the launch directory's AGENTS.md and ran commands in$HOME(One-shot CLI (hermes chat -q) resolves project-context discovery against$HOMEinstead of the launch directory #95577).hermes serveboot wrote asetup.readyJSON-RPC line to stdout ahead of READY./compressin the TUI saved the system prompt with HOME as the working directory.hermes gateway stopcould exit 1, and a restart-on-failure supervisor brought the gateway back.bash -licshell ignores SIGTERM orphaned the children it spawned during the grace window.Validation
Against the merged fixes: rebased on main after all six landed;
scripts/run_tests.sh tests/e2e/coregives 14 files, 116 passed, 0 failed, 4 strict xfails (the bugs below).CI e2e job: runs on CPython 3.11.15 / SQLite 3.53.1, and fails if SQLite is WAL-vulnerable. On the pinned 3.11.14 the state.db suites had silently skipped.
Red-proven: 50 bug reinsertions, each turning at least one suite red. Some are historical reverts:
526d135a96a,91df54184d3,80154cf3cfe(providers.<id>.stale_timeout_seconds cannot shorten a hung stream — reasoning-model stale-timeout floor silently overrides the explicit config #115024), the WAL-unlink lock guard, and each of this PR's own fixes. Others are injected: retry budget ignored,interrupt()doing nothing, truncated summary accepted, split write transaction, repair without its live-writer checks, pooled connections leaking, READY sent to stderr, and more. The per-suite red tables are in each lane report.Stability: every file passed 10 of 10 consecutive runs on a shared box at load average 40–170.
CI: the
e2ejob now runs onubuntu-latest-32-corewith Node 26 and a 900 s per-file budget, one subprocess per file, in parallel.Known bugs pinned by strict xfail (they flip red once fixed)
tool_searchcatalog rebuilt per surface,skill_manageoverride lost), so every switch is a full cache miss.disabled_toolsets([Bug]: ACP _make_agent drops config agent.disabled_toolsets — tools can't be disabled for ACP/Buzz agents #74582).Pending (will land on this branch)
Infographic