Skip to content

test: real multi-process E2E suites for state.db integrity, compaction, liveness, history replay and entrypoint parity - #120171

Merged
teknium1 merged 28 commits into
mainfrom
tests/core-e2e
Sep 23, 2026
Merged

teknium1 merged 28 commits into
mainfrom
tests/core-e2e

Conversation

@teknium1

@teknium1 teknium1 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Real multi-process E2E suites for the failure classes users keep hitting: state.db corruption, compaction losing or duplicating history, hung turns, history that doesn't replay, and features that work in one entrypoint but not another. They found 6 production bugs, each landed as its own PR (#120175, #120178, #120161, #120158, #120152, #120176). This branch is rebased on those fixes and carries tests and CI only.

Follow-up to #120071 (the low-value test purge). The Desktop transcript-integrity suite (the duplicate-message class) and the remaining lanes (tenancy, delivery, upgrade, live-provider, terminal) will be added to this PR as they finish; see Pending below.

What's in it

Every suite drives real Hermes code in real OS processes: AIAgent, SessionDB on a real WAL state.db, hermes serve, the messaging gateway, the TUI backend (tui_gateway), hermes chat --resume, cron, ACP and a real stdio MCP server. Only the model is faked. tests/fakes/fake_llm_provider.py is a scripted, recording loopback server that speaks OpenAI chat-completions and can inject faults (stall mid-stream, 5xx, 429 with Retry-After, truncated or refused summaries, raw bodies). Each suite checks invariants after every step, not a single happy-path outcome.

Suite (tests/e2e/core/…) Class What must hold after every step
sqlite/test_torture_chamber.py, test_compaction_contention.py C1 state.db integrity integrity_check ok and still WAL; every acknowledged write stored exactly once; row counts never go down; repair never loses rows; FTS matches the real rows; no process holds a deleted db/-wal/-shm file; open file counts stay flat. Faults: kill -9, SIGTERM, lock cancellation, chmod flips, killed FTS rebuilds, repair on a live then an offline DB, killing every process at once.
compaction/ (4 files, 34 tests) C4 compaction Seeded random sessions (parallel tool calls, huge turns, images, oversize tool results), summarizer faults (empty, refusal, truncated, timeout, 5xx), /compress, /compress here N, micro-compaction, kill -9 mid-summary and mid-commit. Every user message stays live or archived; tool calls and results stay paired in the DB and in every request; the DB matches the request exactly; resume continues identically; bounded summarizer calls and turn time.
history/test_transcript_ledger.py, test_prefix_stability.py C3 history, C17 prompt cache Each row stored once, and the next request replays exactly what was saved, through steer, interrupt, stream drop plus 5xx retry and every compaction mode. A fresh --resume process sends a byte-identical prefix. Tools array, system prompt and history stay byte-stable across TUI, oneshot and working-directory hops, with one allowed break at compaction. Usage matches what the provider billed.
chaos/ (agent, gateway, tui_gateway) C8 turn liveness 45 fault scenarios. The turn ends within its deadline (15 s after an interrupt, /stop or SIGTERM); provider calls stay bounded; the same session answers the next message; every tool call has its result; no orphan processes; the event loop stays responsive during a stuck turn.
parity/ C19 entrypoint parity, C5 boot, C15 MCP 8 entrypoints × 16 cells (oneshot, chat -q, tui_gateway, serve over WS, gateway, api_server, ACP, cron run-now). The READY handshake is parsed by Desktop's real backend-ready.ts under Node. MCP server death, restart and replay transitions.

Production bugs found (each merged as its own PR; this branch carries no production code)

Validation

  • Against the merged fixes: rebased on main after all six landed; scripts/run_tests.sh tests/e2e/core gives 14 files, 116 passed, 0 failed, 4 strict xfails (the bugs below).

  • CI e2e job: runs on CPython 3.11.15 / SQLite 3.53.1, and fails if SQLite is WAL-vulnerable. On the pinned 3.11.14 the state.db suites had silently skipped.

  • Red-proven: 50 bug reinsertions, each turning at least one suite red. Some are historical reverts: 526d135a96a, 91df54184d3, 80154cf3cfe (providers.<id>.stale_timeout_seconds cannot shorten a hung stream — reasoning-model stale-timeout floor silently overrides the explicit config #115024), the WAL-unlink lock guard, and each of this PR's own fixes. Others are injected: retry budget ignored, interrupt() doing nothing, truncated summary accepted, split write transaction, repair without its live-writer checks, pooled connections leaking, READY sent to stderr, and more. The per-suite red tables are in each lane report.

  • Stability: every file passed 10 of 10 consecutive runs on a shared box at load average 40–170.

  • CI: the e2e job now runs on ubuntu-latest-32-core with Node 26 and a 900 s per-file budget, one subprocess per file, in parallel.

Known bugs pinned by strict xfail (they flip red once fixed)

Pending (will land on this branch)

  • Desktop Playwright core suite: transcript oracle (every message rendered exactly once, in order, never duplicated even for a frame), switch-back race, onboarding → first chat, boot/kill/quit/relaunch orphan census, and the clarify/approval round trip. Stability is currently 9/10; one boot-census failure is still being diagnosed.
  • Lanes still running: tenancy/profile isolation, cron/delivery, upgrade/install, live-provider wire, terminal/execution.

Infographic

core e2e

@teknium1
teknium1 requested a review from a team September 23, 2026 12:11
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/gateway Gateway runner, session dispatch, delivery comp/cli CLI entry point, hermes_cli/, setup wizard comp/cron Cron scheduler and job management comp/tui Terminal UI (ui-tui/ + tui_gateway/) comp/desktop Electron desktop app (apps/desktop/*) tool/terminal Terminal execution and process management area/compression Context compression and continuation sessions area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation labels Sep 23, 2026
@teknium1 teknium1 changed the title test: real multi-process E2E suites for state.db integrity, compaction, liveness, history replay and entrypoint parity (+6 bugs they found) test: real multi-process E2E suites for state.db integrity, compaction, liveness, history replay and entrypoint parity Sep 23, 2026
teknium1 added a commit that referenced this pull request Sep 23, 2026
…ced handler errors, no-retry CI)

Independent review of #120171 found checks that could not fail. Each is now
proven red by a mutation that the old version reported as XFAIL or pass.

- chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used
  raises=AssertionError and RpcError subclasses it, so a gateway crash counted
  as the expected failure. Every invariant is now asserted normally; only the
  known leftovers (surviving tool tree and the tool_call it leaves without a
  result, both fixed by #120306) raise ToolOutlivedGateway, the only exception
  the xfail accepts. The DB check used to sit behind the orphan assert and
  never ran; running it exposed the dangling tool_call half of the same bug.
- history/test_prefix_stability: surface_switch's strict xfail tripped at the
  first prefix break, before usage and integrity. Messages/system prompt,
  usage and integrity are asserted first; the tools-array drift is checked
  last and raises ToolsArrayDrift, the only exception the xfail accepts.
- history/test_transcript_ledger: scripted steer/interrupt callables run on
  the fake provider's handler thread, where an assert only dropped the
  connection. Script records those failures and the test re-raises them after
  every turn; steer must land and the interrupted turn must report
  interrupted=True within 30 s.
- fakes/fake_llm_provider: Hang drops the connection at its deadline instead
  of leaving a kept-alive client waiting past it.
- parity: the API server port was picked, released, then bound by the child.
  Readiness now requires our child's pid from authenticated /health/detailed
  and retries on a fresh port when the child reports it in use. The fixture
  guard refused any HERMES_HOME under ~/.hermes, failing all parity tests
  whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root
  or a real profile.
- chaos/_gateway_harness: the gateway stays in pytest's process group, so
  the runner's kill of a timed-out file reaches it.
- sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable
  constructor failed: messages_fts"; the delete arm's busy tolerance keys on
  the result code. A failed episode's roles are stopped so the shared chamber
  and rig no longer fail every later episode.
- chaos, compaction, parity homes: updates.check=false (history already had
  it). The passive update check made a GitHub round-trip from every test
  surface, and on a blobless clone whose objects lag upstream its
  `git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that
  the 5 s timeout orphans; the orphan scans then failed on git processes.
- chaos/test_agent_turn_liveness: a PROBE failure now carries the provider
  call counts and the agent's stale-kill log, so a cross-turn breaker trip
  can be told apart from a slow probe.
- tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never
  retried into green; own uv cache entry (cache-suffix: e2e).
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 5535d03 — test(e2e): micro-compaction display check is a plain test no

❌ Job failures

JS & TS checks / JS & TS checks · View job

Job JS & TS checks / JS & TS checks failed.


ℹ️ Info

CI-sensitive file review · View job

PR touches sensitive files, but the ci-reviewed label has been added, approving them.

Sensitive files changed:


debug info

CI timings

CI timings · View report · View job

Wall time 5m47s vs 9m41s (-40.3%). 9 job(s) slower, 10 faster,

  • OS-specific tests / Windows-only tests: -231.0s
  • Check contributors / check-attribution: -137.0s
  • Python lints / ruff enforcement (blocking): -101.0s
  • Python lints / Windows footguns (blocking): +75.0s
  • OS-specific tests / macOS-only tests: +73.0s

Class C1: state.db corruption, WAL generations unlinked under a live holder,
lost/duplicated rows, repair destroying data, fd leaks. Every role is a real
OS process on one WAL state.db through the production SessionDB: gateway- and
TUI-like writers with intent/ack journals, a dashboard reader living across
episodes, short-lived openers (SessionDB, bare sqlite3, real `hermes sessions
list/stats`), FTS rebuild/optimize and repair_state_db_schema. Nine seeded
episodes inject kill -9 mid-write, SIGTERM graceful close, POSIX lock
cancellation by a stray in-process open/close, chmod flips, concurrent and
killed FTS rebuilds, repair against a live and an offline store, FTS
corruption and a whole-fleet SIGKILL, then assert the same invariants:
integrity_check ok and still WAL, no child holding a (deleted) db/-wal/-shm
(/proc fd monitor), every acked append stored exactly once, counts only grow,
repair never lowers them, FTS == canonical rows and search finds acked rows
once, no role errors, bounded fds for the long-lived reader and churner.

Red-proof: disabling the OFD WAL lock guard (75e155a, whose revert no
longer applies cleanly) makes the lock_cancellation episode kill the gateway
writer with SIGBUS / fail sibling openers, 3/3 runs.
Class C1, "state.db corrupting from compaction": real AIAgent processes run
real turns (user -> terminal tool call -> answer) through the loopback fake
provider on one shared WAL state.db while a plain writer, a reader and an
open/close churner work the same file. A gateway-like agent micro-compacts
every turn; a TUI-like agent runs /compress here 2; faults are clean, kill -9
mid-turn and kill -9 timed into the micro-compaction summary/commit window;
a fresh process then resumes the session. Invariants: every acked user turn
live exactly once, its tool result and answer recoverable (live or
compacted history) and never live twice, the /compress-kept exchanges live
exactly once, canonical rows never shrink (external sampler + reader),
integrity ok, FTS == canonical, resumed request carries every acked user
turn exactly once, plain-writer appends exactly once.

Red-proof: reverting 91df541 (micro-compaction tool rows) turns every
absorbed tool result/answer into active=0,compacted=0 debris; reverting
526d135 (/compress here tail) leaves the kept exchanges with no live row.
… the fake provider

Backward compatible: FakeLLMServer(prompt_tokens_fn=...) reports usage.prompt_tokens
computed from each request body (token-driven logic such as compaction triggers sees a
realistic, growing count instead of a constant 100), and Text(finish_reason=...) lets a
scripted reply end with e.g. 'length' (truncated summaries). Defaults are unchanged.
Class C4 (context compression correctness/liveness). A real AIAgent + SessionDB on disk drives
seeded random sessions (tool-pair heavy incl. parallel calls, huge single turns, image parts,
long answers, single tool results larger than the whole threshold, a cron-shaped single job
turn with 14 tool rounds, a provider whose real window is smaller than the configured one)
against the recording fake provider. After every turn and for every request it asserts:
bounded wall time and summarizer calls, all scripted steps consumed or an explicit error,
no user message lost from state.db (live or compacted=1), superseded rows always have a live/
archived copy, tool_call/tool_result pairs matched in live rows and in every request, system
prompt stable, the in-flight ask (cron job prompt) after every summary (#100818), and
persisted == sent: state.db at request time is exactly the request's history prefix, live rows
equal the in-memory history, and a fresh agent resumed from state.db continues identically.
Exposed the stale current_turn_user_idx bug fixed in the previous commit.
Same real chain and per-turn invariants, with the summarizer answering empty / refusal /
truncated (finish_reason=length) / slower than the compression timeout / HTTP 500. A bad
summary must never archive history (P0 #94448 empty-summary deletes the middle; refusal
accepted as a summary), never reach the model or state.db, must be surfaced to the user, and
must not loop (877-summarizer-call class). A 500 may commit only the designed deterministic
fallback, never with abort_on_summary_failure.
… == sent (C4)

Drives manual compaction through the TUI/Desktop choke point (tui_gateway
_compress_session_history -> compress_now) and rolling micro-compaction through the real turn
finalizer, then asserts no user message lost, live rows == in-memory history, kept exchanges
stay live verbatim rows, failed summaries change nothing, and the next turn from memory AND from
a fresh agent resumed off state.db both send exactly the persisted history. Covers the
526d135 (/compress here N tail lost from state.db) and 91df541 (micro-compaction
marks summarized rows rewind-only) data-loss class.
… (C4)

A real child process runs /compress on a seeded session's state.db; the parent SIGKILLs it
while the summary request is in flight and inside the archive_and_compact write transaction
(parked by a child-side hook after archive + inserts, before COMMIT). Asserts integrity_check
ok, rows byte-identical to pre-compaction, message_count consistent, then a fresh agent
reclaims the dead holder's lease, compacts and continues with persisted == sent.
… step)

Class C3 (conversation-history persistence & replay integrity). A real AIAgent + SessionDB
runs a scenario matrix (plain turns, parallel tool batches, /steer mid-turn, interrupt,
stream drop + 500 retry, manual /compress, /compress here N, auto compaction, opt-in
micro-compaction) against the recording fake provider. After EVERY step it asserts:
exactly-once at the storage layer and in the model view, next request replays the
persisted view, nothing the model saw ever leaves the durable/display history, each typed
input shown once, rows monotonic, usage == provider-billed; then a FRESH
'hermes chat --resume' process must replay a byte-identical request prefix.

Red-proofs: reverting 526d135 (compress-here tail) and 91df541 (micro-compaction
rewind flags) and an injected double row insert all go red. A strict xfail documents a live
micro-compaction bug (merged user row duplicates both inputs in the resumed display).

The fake provider now records the usage it billed per request (record['usage']).
Class C17 (prompt-cache prefix byte-stability + usage/cost accounting). One durable session
is driven for 10+ turns through FRESH processes of the real entrypoints (tui_gateway stdio
JSON-RPC server, 'hermes chat -q --resume' oneshot), flipping cwd between hops, with a
parallel tool batch and a manual session.compress. Asserts every main request is a
byte-identical extension of the previous one (tools array, system prompt, earlier messages)
with exactly one sanctioned break at the compaction, every process boundary replays the
persisted model view, storage rows are exactly-once and monotonic, and sessions /
session_model_usage equal what the fake provider billed.

tui_gateway_restarts went red on base (the gateway's manual /compress persisted a prompt
with the process HOME as cwd; fixed in the previous commit). surface_switch is a strict
xfail documenting live TUI<->oneshot tools-array divergence (tool_search catalog rebuilt per
surface; pinned tools lose dynamic schema overrides; -q --resume prunes skill_manage).
One fixture HERMES_HOME (shell + plugin pre_llm_call hooks, AGENTS.md, a
skill, a memory entry, a stdio MCP server that spawns a grandchild, the
recording fake provider, a disabled toolset) and one scripted turn per
entrypoint: hermes -z, hermes chat -q, tui_gateway stdio, hermes serve over
the Desktop WS, GatewayRunner with a fake adapter, api_server, ACP stdio,
cron run-now. The fake model calls the MCP canary tool, then answers.

Same invariants on every surface, first turn after a cold start: context
file / skill index / memory in the recorded system prompt, both hooks fired
and injected, MCP tool offered and a REAL call returned the canary (the
inverted stdio-liveness burst failed exactly here everywhere), documented
toolset's feature tools present and the disabled toolset absent, answer
delivered to the client, zero MCP server/grandchild survivors after the
surface's normal shutdown. Catches the 'works in the CLI, missing in
Desktop/gateway/ACP/cron' class; found #95577 (oneshot row). ACP's
disabled_toolsets gap (#74582) is a strict known-red cell.
Spawns hermes serve exactly as the Desktop does and tui_gateway.entry as
the Ink TUI does. serve: the first stdout line must be the READY sentinel,
resolved to the live port by the Desktop's REAL parser
(apps/desktop/electron/backend-ready.ts run under Node, no regex copy), and
the port answers an authenticated /api/status. tui_gateway: the first
stdout frame is a contract-valid gateway.ready and stdout stays pure
JSON-RPC. Found setup.ready leaking onto serve stdout before READY.
… (C15)

tui_gateway host + the real stdio MCP fixture:
- fast death: the server crashes mid-call leaving a helper that inherited
  its stdio (npx/uvx wrapper shape). The call must surface as an
  outcome-uncertain tool error exactly once (never replayed onto the
  respawned server), the host survives, the turn completes within a bound
  12x under the configured call timeout, and shutdown reaps the helper.
- host crash: SIGKILL the host after a turn; the death supervisor must
  still reap the MCP server and its grandchild.

Also pins mcp_discovery_timeout in the fixture home: interactive surfaces
wait only ~1.5 s for discovery by design, which a loaded box exceeds, so
the first-turn-complete invariant needs the documented knob raised.
…veness)

Class C8 (agent-turn liveness): hangs, runaway retries, lost tool results, wedged
agents, orphan processes. A real AIAgent + SessionDB runs in a child process with
config.yaml liveness knobs (agent.api_max_retries, providers.custom request/stale
timeouts, agent.max_turns) against the scripted loopback provider. 22 fault modes
(hang, reasoning-model stall, 500/429/400/garbage forever, drop mid-stream, hang then
recover, slow trickle, malformed/unknown tool calls, 8-way parallel, hung/stdin/huge
tools, background swarms, endless tool loop, interrupts) are each followed by a probe
turn on the same agent and checked against the same invariants: bounded termination,
bounded provider calls, reusability with history kept, every tool_call answered (in
state.db and on the wire), no surviving tagged process, integrity_check ok.
python -m gateway.run with a fake platform adapter: each fault mode must end the turn
within its deadline (or promptly on /stop / SIGTERM), keep provider calls bounded,
release the turn lease so the session answers the next message (#104303 class), keep
the event loop responsive (heartbeat), answer every tool_call (#93251 class) and leave
no orphan process.
…server

python -m tui_gateway.entry over stdio: every fault mode ends with a terminal
message.complete within its deadline (interrupt within seconds), the dispatcher keeps
answering during a wedged turn, the busy flag is released so the same session streams
the next prompt, tool results are complete, and exit on EOF/SIGTERM mid-turn leaves no
orphans (the foreground-tool orphan on exit is a known base bug, strict xfail).
…ythons without os.pidfd_open

CI's pinned uv only knows CPython 3.11.14, whose bundled SQLite has the WAL-reset
bug, so Hermes ran state.db in DELETE mode and every torture-chamber episode
skipped (green over zero coverage). The chaos cleanup also called os.pidfd_open,
which that build lacks, turning two strict xfails into errors.
…DELETE journal mode too

Both suites skipped outright on a WAL-reset-vulnerable SQLite, so the DELETE-mode population
(every user whose Python bundles SQLite 3.7.0-3.51.2, e.g. uv CPython 3.11.14's 3.50.4, and the
network/FUSE homes Hermes also falls back to DELETE on) had no multi-process integrity coverage.

- Parametrize both module fixtures over journal mode [wal, delete]. The delete arm goes through the
  production decision path: each child (_roles.py, including the `hermes` CLI run as __main__)
  pins the version hermes_state_wal.is_sqlite_wal_reset_vulnerable() reports to 3.50.4, and
  apply_wal_with_fallback picks DELETE itself; the harness never issues a journal-mode pragma.
  The wal arm skips only where Hermes would not run WAL; the delete arm runs everywhere.
- Every invariant that holds in both modes stays on in both: integrity_check, acknowledged writes
  exactly once, monotonic counts, repair never lowers rows, FTS parity, fd bounds,
  persisted == replayed. New in both: the store is still in the arm's mode and every SessionDB
  role reported the matching _wal_active (vacuity guard for the seam). WAL-only: the deleted
  -wal/-shm fd scan (the deleted main-file scan runs in both).
- DELETE mode blocks readers on writes and has no writer fairness: in that arm a SQLITE_BUSY
  refusal of a read/open/FTS pass is waited out and counted (reported in failure context), and
  plain writers are paced by 20 ms; in the wal arm busy still fails the role.
- chmod_flip runs 3-5 flips instead of 6-10 (same fault, process start-up dominated the time).

Red-proof (delete arm red, wal arm green): repair without live-writer checks/exclusion ->
"acked ... stored 0x"; a DELETE-mode commit misreported busy so the retry re-runs the insert ->
"acked ... stored 2x" (torture and compaction); default config enabling WAL on a vulnerable
SQLite -> "journal_mode is 'wal' (arm is delete)" + "_wal_active=True in the delete arm".
…ced handler errors, no-retry CI)

Independent review of #120171 found checks that could not fail. Each is now
proven red by a mutation that the old version reported as XFAIL or pass.

- chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used
  raises=AssertionError and RpcError subclasses it, so a gateway crash counted
  as the expected failure. Every invariant is now asserted normally; only the
  known leftovers (surviving tool tree and the tool_call it leaves without a
  result, both fixed by #120306) raise ToolOutlivedGateway, the only exception
  the xfail accepts. The DB check used to sit behind the orphan assert and
  never ran; running it exposed the dangling tool_call half of the same bug.
- history/test_prefix_stability: surface_switch's strict xfail tripped at the
  first prefix break, before usage and integrity. Messages/system prompt,
  usage and integrity are asserted first; the tools-array drift is checked
  last and raises ToolsArrayDrift, the only exception the xfail accepts.
- history/test_transcript_ledger: scripted steer/interrupt callables run on
  the fake provider's handler thread, where an assert only dropped the
  connection. Script records those failures and the test re-raises them after
  every turn; steer must land and the interrupted turn must report
  interrupted=True within 30 s.
- fakes/fake_llm_provider: Hang drops the connection at its deadline instead
  of leaving a kept-alive client waiting past it.
- parity: the API server port was picked, released, then bound by the child.
  Readiness now requires our child's pid from authenticated /health/detailed
  and retries on a fresh port when the child reports it in use. The fixture
  guard refused any HERMES_HOME under ~/.hermes, failing all parity tests
  whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root
  or a real profile.
- chaos/_gateway_harness: the gateway stays in pytest's process group, so
  the runner's kill of a timed-out file reaches it.
- sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable
  constructor failed: messages_fts"; the delete arm's busy tolerance keys on
  the result code. A failed episode's roles are stopped so the shared chamber
  and rig no longer fail every later episode.
- chaos, compaction, parity homes: updates.check=false (history already had
  it). The passive update check made a GitHub round-trip from every test
  surface, and on a blobless clone whose objects lag upstream its
  `git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that
  the 5 s timeout orphans; the orphan scans then failed on git processes.
- chaos/test_agent_turn_liveness: a PROBE failure now carries the provider
  call counts and the agent's stale-kill log, so a cross-turn breaker trip
  can be told apart from a slow probe.
- tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never
  retried into green; own uv cache entry (cache-suffix: e2e).
…20316 landed

Drop the strict xfail on the resumed-display test and re-enable the input
check for the micro_compaction ledger scenario. The ledger no-loss oracle now
exempts model-only rows (the merged user turn #120316 flags model_only): they
stand in for originals that stay displayed and were each tracked when sent.

A/B: removing the model_only flag in agent/micro_compaction.py turns both
micro cases red; with it, the history suites are green.
@teknium1
teknium1 merged commit 2b1bb70 into main Sep 23, 2026
35 checks passed
teknium1 added a commit that referenced this pull request Sep 23, 2026
…ced handler errors, no-retry CI)

Independent review of #120171 found checks that could not fail. Each is now
proven red by a mutation that the old version reported as XFAIL or pass.

- chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used
  raises=AssertionError and RpcError subclasses it, so a gateway crash counted
  as the expected failure. Every invariant is now asserted normally; only the
  known leftovers (surviving tool tree and the tool_call it leaves without a
  result, both fixed by #120306) raise ToolOutlivedGateway, the only exception
  the xfail accepts. The DB check used to sit behind the orphan assert and
  never ran; running it exposed the dangling tool_call half of the same bug.
- history/test_prefix_stability: surface_switch's strict xfail tripped at the
  first prefix break, before usage and integrity. Messages/system prompt,
  usage and integrity are asserted first; the tools-array drift is checked
  last and raises ToolsArrayDrift, the only exception the xfail accepts.
- history/test_transcript_ledger: scripted steer/interrupt callables run on
  the fake provider's handler thread, where an assert only dropped the
  connection. Script records those failures and the test re-raises them after
  every turn; steer must land and the interrupted turn must report
  interrupted=True within 30 s.
- fakes/fake_llm_provider: Hang drops the connection at its deadline instead
  of leaving a kept-alive client waiting past it.
- parity: the API server port was picked, released, then bound by the child.
  Readiness now requires our child's pid from authenticated /health/detailed
  and retries on a fresh port when the child reports it in use. The fixture
  guard refused any HERMES_HOME under ~/.hermes, failing all parity tests
  whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root
  or a real profile.
- chaos/_gateway_harness: the gateway stays in pytest's process group, so
  the runner's kill of a timed-out file reaches it.
- sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable
  constructor failed: messages_fts"; the delete arm's busy tolerance keys on
  the result code. A failed episode's roles are stopped so the shared chamber
  and rig no longer fail every later episode.
- chaos, compaction, parity homes: updates.check=false (history already had
  it). The passive update check made a GitHub round-trip from every test
  surface, and on a blobless clone whose objects lag upstream its
  `git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that
  the 5 s timeout orphans; the orphan scans then failed on git processes.
- chaos/test_agent_turn_liveness: a PROBE failure now carries the provider
  call counts and the agent's stale-kill log, so a cross-turn breaker trip
  can be told apart from a slow probe.
- tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never
  retried into green; own uv cache entry (cache-suffix: e2e).
@teknium1
teknium1 deleted the tests/core-e2e branch September 23, 2026 21:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/compression Context compression and continuation sessions area/sessions Session lifecycle, resume, persistence, history ci-reviewed applied to manually approve dangerous changes comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/cron Cron scheduler and job management comp/desktop Electron desktop app (apps/desktop/*) comp/gateway Gateway runner, session dispatch, delivery comp/tui Terminal UI (ui-tui/ + tui_gateway/) P2 Medium — degraded but workaround exists sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state tool/terminal Terminal execution and process management type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants