Skip to content

test: E2E suites for profile isolation, key routing, terminal transcripts, hermes update and live providers (core E2E wave 2) - #120326

Merged
teknium1 merged 27 commits into
mainfrom
tests/core-e2e-wave2
Sep 24, 2026
Merged

teknium1 merged 27 commits into
mainfrom
tests/core-e2e-wave2

Conversation

@teknium1

@teknium1 teknium1 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Four more end-to-end suites now guard classes of bugs users keep hitting: one profile's keys, model, cwd or .env reaching another profile; a request sent to the wrong host with the wrong key; a reply shown twice (or not at all) in the terminal; and an install broken by hermes update.

Rebuilt on main after #120171 (tests/core-e2e, merged as 2b1bb70): origin/main + this PR's own commits only.

What it adds

All suites use real processes, the recording loopback provider (tests/fakes/fake_llm_provider.py), a temp HOME, and no mocks of Hermes code.

Suite Class What it drives Invariant
tenancy/test_routing_truth_table.py C11 18 real hermes chat -q / -z configs (custom base_url, key_env vs literal, named and legacy providers, aliases, 401/429 fallback, 1- and 2-key pools, per-profile, aux, delegation) + one live TUI session walking 10 /model switches every request reaches exactly the selected host with exactly that host's key; anything else is a LEAK/MISSING failure; a CONNECT trap catches real egress
tenancy/test_two_tenant_gateway.py C7 one multiplexed hermes gateway run serving 3 profiles: interleaved api_server turns, cron fires, cross-key 401s, profile create/delete next to the live host, SIGTERM restart after every phase, no provider request, tool subprocess env, profile file (state.db + WAL, logs, config, .env) or backend log carries another tenant's canary
tenancy/test_two_tenant_desktop_backend.py C7 the same over one hermes serve (Desktop) on /api/ws: session.create {profile}, settings RPCs, model.save_key, restart + session.resume same, plus only the addressed profile's config.yaml / .env changes, byte-for-byte
terminal/test_terminal_transcript_pty.py C2 real hermes chat --cli and hermes --tui on a real PTY, rendered by an in-tree VT emulator: long streams, tool calls, reasoning, width changes mid-stream every reply exactly once and verbatim, in order; every persisted row rendered once; /exit leaves no processes
terminal/test_multiclient_session_model.py C14 seeded model check: 3 WS clients × create / switch / interrupt / delete-live / drop + reconnect / owner death + takeover on the real backend no duplicate or lost events, leases released, live sessions not deletable
terminal/test_interactive_roundtrip.py C20 clarify / deny / approve × immediate, switch-away-and-back, 2nd connection, reconnect, approval.respond, with a real rm -rf under manual approvals the answer reaches the tool, a deny never runs the command
upgrade/test_upgrade_path.py C6 real hermes update --yes from the previous release tag: clean, autostash, SIGKILL mid-pull, SIGKILL before deps, offline the retry heals to HEAD; the venv serves HEAD's tree; both profiles' state.db rows, .env, cron ids and migrated config survive; no service-manager restart escapes the sandbox
upgrade/test_fresh_process_entrypoints.py C6 imports all 1,785 shipped modules in clean processes, plus every [project.scripts] entry and serve's READY/SIGTERM handshake nothing fails at import; every entry point starts and exits cleanly
upgrade/test_config_roundtrip_properties.py C18 317 seeded cases over hermes config set, TUI config.set, dashboard PUT /api/config, migrations from every historical version, transient EMFILE a one-key save touches only that key; a read error never clobbers; migrations are idempotent
live/test_live_providers.py C9, C17 real provider APIs (OpenRouter, Anthropic, Nous, OpenAI, Gemini, xAI): /models parse plus a 3-turn tool conversation each, with dollar guards vendor schema drift, reasoning replay, streaming shape, prompt-cache hits and real auth caught nightly

The live suite sits behind a new live marker, excluded by default (addopts = "-m 'not integration and not live'"). Each case skips when its key is absent.

Unlanded fixes: merge-order-safe expected failures

No static xfail(strict=True) is left in these suites. A static strict xfail turns main red (XPASS) the moment its fix merges (this PR hit exactly that: 26 config round-trip cells XPASSed once main fixed their bugs). Every remaining expected failure is merge-order safe, one of two ways, both in tests/e2e/core/_pending_fixes.py:

No test reads source code.

Cell Bug Fix PR Gate
test_tui_gateway_model_switch_routing (leg 5) /model --provider X uses an alias's endpoint and key #120295 probe, strict, RoutingLeak
test_multiplexed_gateway_never_crosses_tenants routed cron work reuses the default profile's terminal env #120307 probe, strict, TenantLeak
test_desktop_backend_never_crosses_tenants gateway.run's import-time bridge leaks a routed .env (a race) #120307 probe, non-strict, TenantLeak
test_terminal_transcript_integrity[cli-resize_scrollback] #95375 resize re-prints scrollback #120321 message: a turn rendered/echoed ≥ 2x (DuplicateRender only)
test_kill_mid_pull_then_retry_heals[torn-tree] a torn checkout after a killed pull dies at import #120339 message: final hermes update failed … cannot import name
test_p2_env_lock_refusal_is_not_reported_as_success[*] #119928 refused .env write exits 0 #119929 message: exit=0 but .env unchanged / refused env write still rewrote config.yaml
test_models_listing_parses[gemini] (live) #62259 Bearer auth on native /v1beta/models #62267 / #116509 message: empty live listing

Probes for fixes that have since merged were deleted with their gaps (#120319, #120374, #120299); those cells are plain tests. Cells whose config/entrypoint bug main fixed are plain strict tests with their workarounds removed: #119844 long-scalar re-fold (config round-trip + test_clean_update_keeps_long_scalar_lines_verbatim), explicit-null strip, phantom agent: {}, transient-read clobber on TUI RPC / dashboard PUT, #54648 hermes-agent --help.

CI wiring (.github/workflows/tests.yml)

  • e2e job: npm ci --workspace ui-tui + builds ui-tui, and sets HERMES_E2E_REQUIRE_TUI=1 so a missing TUI build fails instead of skipping (scripts/run_tests.sh now forwards HERMES_E2E_REQUIRE_TUI, CI and GITHUB_ACTIONS through its env -i; before, the flag never reached pytest and the TUI cells skipped). It runs every tests/e2e file except core/upgrade.
  • New e2e-upgrade job: 60 min, 32-core, fetch-depth: 0 + tags (N-1 = git describe --tags HEAD~1). It installs bubblewrap and lifts the Ubuntu 24.04 userns restriction, then asserts the suite's own _bwrap_usable() probe (same flags), so every spawned updater runs in its own PID namespace with no systemd bus. N-1's venv is built from N-1's own uv.lock (uv sync --locked --extra all, the installer's tier 0), and HEAD/N-1 refs resolve lazily (no git at collection). It keeps the warm uv cache and uses a 3000 s file budget. All required checks pass already covers it through python-tests.
  • live-providers.yml: runs nightly (07:17 UTC), on v* tags and on dispatch. It only runs on the main repo, and it no-ops without LIVE_*_API_KEY secrets. It calls pytest directly because run_tests.sh strips credentials. The key values are scoped to the one Run live canaries step (job level carries only HAS_* booleans), so uv sync build backends and third-party actions never see them. It posts a usage/$ step summary.

Validation

Check Result
Rebased on origin/main 529d27e, every PR suite via run_tests.sh --include-integration, --file-retries 0 0 XPASS, 0 failed. upgrade: config round-trip 315 passed + 2 xfail, upgrade_path 6 + 1 xfail, fresh_process 14 + 0; tenancy: routing 18 + 1 xfail, gateway 0 + 1 xfail, desktop backend 0 + 1 xfail; terminal: interactive 10 + 0, multiclient 3 + 0, transcript PTY 4 + 1 xfail (TUI built, HERMES_E2E_REQUIRE_TUI=1); live: skipped without keys
Base (pre-fix head rebased on main) reproduces CI: 26 failed, all XPASS(strict) (24 config round-trip, hermes-agent-help, long-scalar update)
test_multiclient_session_model.py after the wait_first_delta per-connection index fix 10/10 runs green (30/30 cells)
HERMES_E2E_REQUIRE_TUI=1 through run_tests.sh with ui-tui/dist/entry.js removed the 2 TUI cells FAIL (before the forward: 2 skipped)
Live suite without keys -m live: 17 skipped; default: 17 deselected
ruff, check_no_tmp_literals, check-windows-footguns --all, check_public_surface, git diff --check clean

C20 approval_deny-rpc race: fixed on main by #120374; the cell is a plain test now.

Wave-1 flakes seen under load (not wave-2 files, not fixed here):

  • compaction/test_compaction_manual.py / test_compaction_auto.py: "the session system prompt was rewritten". This is a production bug. The workspace git snapshot is supposed to stay fixed for the session, but it is re-probed on rebuilds: manual /compress binds the cwd after the first build, so the cwd key changes from "" to the path; and a resumed agent never restores the saved snapshot. A checkout whose git status changes mid-session (for example, one more untracked file) then changes the prompt. It reproduces every time when a file is written into the repo after turn 0. A proposed ~30-line fix in agent/system_prompt.py + conversation_loop.py has red/green evidence, but no PR yet.
  • sqlite/test_torture_chamber.py[fts_corruption_fail_open]: a writer got database disk image is malformed at load ~200, and kill9_everything then saw its leftover gateway. Seen once and not triaged.

Each lane's own record: 10/10 consecutive green per file with retries disabled, and the red proofs are in the lane reports (tenancy 14, terminal 15 + 1 live bug, upgrade 28 + 5 fix-flips).

Not covered: MCP bearer tokens per profile, messaging adapters other than api_server, Windows/macOS update paths, gateway service restart after an update, OAuth device-code refresh, and credential pools with more than 2 keys. Each lane's NOT_COVERED.md lists them.

Infographic

wave-2 core E2E

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 2a55a8b — test(e2e): drop the #120307/#120295 gaps now on main; tighte

ℹ️ Info

CI-sensitive file review · View job

PR touches sensitive files, but the ci-reviewed label has been added, approving them.

Sensitive files changed:


debug info

CI timings

CI timings · View report · View job

Wall time 7m15s vs 6m8s (+18.2%). 12 job(s) slower, 7 faster,

  • OS-specific tests / macOS-only tests: -138.0s
  • Docs Site / docs-site-checks: -65.0s
  • Python tests / Run tests: +63.0s
  • Rust tests / cargo test (bootstrap installer): +31.0s
  • Python lints / Windows footguns (blocking): +23.0s

@alt-glitch alt-glitch added type/test Test coverage or test infrastructure P3 Low — cosmetic, nice to have comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/gateway Gateway runner, session dispatch, delivery comp/desktop Electron desktop app (apps/desktop/*) comp/tui Terminal UI (ui-tui/ + tui_gateway/) tool/terminal Terminal execution and process management area/profiles Multi-profile isolation, HERMES_HOME scoping area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation labels Sep 23, 2026
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
@teknium1
teknium1 force-pushed the tests/core-e2e branch 2 times, most recently from 5535d03 to 276f56e Compare September 23, 2026 18:10
@teknium1
teknium1 force-pushed the tests/core-e2e-wave2 branch from 92d6040 to b0b886b Compare September 23, 2026 18:17
Base automatically changed from tests/core-e2e to main September 23, 2026 21:54
@teknium1
teknium1 requested a review from a team September 23, 2026 21:54
@teknium1 teknium1 closed this Sep 23, 2026
@teknium1 teknium1 reopened this Sep 23, 2026
@teknium1
teknium1 force-pushed the tests/core-e2e-wave2 branch from b0b886b to 28145e7 Compare September 23, 2026 22:32
@teknium1
teknium1 force-pushed the tests/core-e2e-wave2 branch from 28145e7 to 35f3be7 Compare September 24, 2026 00:22
Class C11 (provider/model routing + credential resolution). One real
loopback OpenAI-compatible host per provider identity, each accepting only
its own key; every scenario configures all identities and only changes the
selection, and one invariant runs after every leg: requests land only on the
selected hosts, each with only that host's key, the user sees that host's
answer, and nothing egresses to a real inference API (CONNECT trap).

Real `hermes chat -q` / `hermes -z` subprocesses cover custom base_url
(key_env / literal), bare custom fail-fast, named providers: and legacy
custom_providers entries, startup -m aliases (#103933/#107191/#109440),
fallback after 401/429, credential pool rotation / single-key exhaustion,
per-profile routing, aux title + delegation routing. The real stdio
tui_gateway walks mid-session /model switches incl. pool recovery without
restart.

fake_llm_provider: api_key may be a list of accepted keys (pool on one
host) and opt-in record_get; defaults unchanged.

(cherry picked from commit bf1f07571d7777fe91b16f386e026c12bdbfe73b)
Class C7 (multiplex / tenancy isolation). A real 'hermes gateway run' with
gateway.multiplex_profiles serves default + alpha + beta, each with its own
loopback provider (accepting only its own key) and distinct canaries: provider
key under the same env var name, API-server key, .env marker, model, MEMORY.md,
SOUL.md, terminal cwd, cron prompt and an exported shell variable. Interleaved
api_server turns (/v1, /p/<name>/v1) each run 'env | sort; pwd' in the terminal
tool; a cron tick fires one job per profile; cross-profile API keys; profile
create/attach/delete next to the live host (PID must not change); SIGTERM
restart + a second cron tick. After every phase: no request, tool-subprocess
env snapshot, profile-home file (state.db + WAL, logs, sessions, config, .env)
or gateway output carries another tenant's canary.

Red on the parent of a445f346a8b (cron shared the launch profile's terminal
environment); see REPORT.md for the red-proof table.

(cherry picked from commit ba5bdae9aa9900d719be36bf9efa40037860cf0b)
Class C7 (multiplex / tenancy isolation). The real app backend, 'hermes serve'
with HERMES_DESKTOP=1 (in-process cron ticker), driven over /api/ws exactly like
the Desktop: session.create {profile}, interleaved prompt.submit turns, a cron
tick per profile, session-bound and profile-bound config.set, session-bound
model.save_key, profiles.create + 'hermes profile delete' next to the live
backend (PID unchanged), restart + session.resume + a second cron tick. Same
canary invariant as the gateway suite plus: each settings write changes only
the addressed profile's config.yaml/.env, and no RPC reply/event for a session
carries another tenant's canary.

Red on the parents of 6952f1d5425 (lazy gateway.run import latched a secondary's
terminal.* into the launch env), 5a14abea326 (save_key wrote the launch .env) and
1e3727666e5 (secondary session reported/persisted the launch model).

(cherry picked from commit 522858661a60abc47ece47d977d73c54067789bc)
…ckend

Class C14 (multi-client session ownership / routing / switching): prompts persisted
to a stale session after a switch, events rendered twice or leaking into another
chat, zombie leases ("already has a live owner"), a ws drop or client absence
killing the in-flight turn, a live session hard-deleted.

tests/e2e/core/terminal/_gateway_client.py spawns ONE real
`python -m hermes_cli.main serve --port 0` per module (Desktop's argv) in an
isolated HOME with only the recording fake provider, and drives it with
Desktop-shaped WebSocket JSON-RPC clients (gateway.ready, client.capabilities,
server->client request answers, abrupt TCP drop, reader pause for absence).

test_multiclient_session_model.py runs a seeded fuzzer (seeds 11/23/37) over 3
clients: create, switch(session.resume), fast and slow prompts, interrupt
mid-stream, delete of live (must be refused) and closed sessions, ws drop +
reconnect + resume mid-turn, private-chat owner death (idle and running) +
takeover, shared-chat owner death, client absence. A reference model checks after
every step and at the end: prompts persisted exactly once, in order, to the
session selected at send time; per-connection event seq strictly increasing (no
duplicate delivery); no events for sessions a connection never attached to;
canaries only in their own session; one live runtime per chat; orphaned runtimes
released in bounded time and resumable/promptable by another client; drop/absence
never cut a turn short (full reply persisted, message.complete status complete).

Red-proof (each fails, restored after): duplicate event write in write_json;
semantic revert of 67de938 (#98028/#100325, orphan reaper ignores turn
activity); semantic revert of de25545 (rebind instead of fan-out); lease never
released; session.delete of a live session not refused.

(cherry picked from commit fe078a891231d3ba459da4ca46bf2a06979f4444)
Class C20 (interactive prompts): a clarify question or dangerous-command approval
never shows, is lost on a session switch / reconnect / the other window, the answer
never reaches the tool, the approval outcome is ignored (denied command runs), or a
prompt stays pending after the turn.

test_interactive_roundtrip.py reuses the lane's real `hermes serve` harness with
approvals.mode=manual (not yolo). The fake provider issues clarify and
terminal(`rm -rf <victim>`) tool calls and records the tool result the agent sends
back. Matrix {clarify, approval-deny, approval-approve} x {answer at once, after
switching away and back (re-delivered via session.resume open_requests), from a
second connection, after an abrupt ws drop + reconnect, via the approval.respond
RPC}. Asserts the server->client request arrives in bounded time, the answer
reaches the tool (clarify user_response / "denied by user" + victim intact /
exit_code 0 + victim gone), the other window gets a card-settling signal,
approval.pending and open_requests are empty afterwards, and the turn completes.

Red-proof (each fails, restored after): clarify callback drops the answer;
approval response choice ignored (deny runs the command); open_requests not
re-delivered on resume; approval.respond RPC routing dropped.

(cherry picked from commit 372e0c04e56c54ceb237937bae7a990212fcfbe6)
… Ink TUI)

Class C2 (duplicate / vanishing / garbled transcript rendering) on the terminal
surfaces. Drives a real `hermes chat --cli` and a real `hermes --tui` (Node
frontend + tui_gateway child) as session leader of a real PTY against the
recording fake provider (long multi-chunk stream, tool-call turn, reasoning
turn, width resizes mid-stream via SIGWINCH), renders the byte stream through a
dependency-free VT emulator (_vt.py, cross-checked identical to pyte on
captured CLI and TUI streams) and asserts on the settled grid + scrollback:
every reply exactly once and verbatim (== scripted stream), every prompt echoed
once, conversation order, tool call once, reasoning at most once, rendered ==
persisted state.db rows, /exit -> 0 with no process left in the PTY session or
among tracked descendants (setsid'ed helpers included).

cli-resize_scrollback is a strict xfail (raises=DuplicateRender) for the open
live bug #95375: a width-changing resize on a 24-row terminal re-prints turns
that already sit in scrollback (warmup reply rendered 3x).

Red-proof: re-render final panel after stream, double-appended deltas,
tui_gateway double event emit, resize replay without clear, gateway outliving
the TUI (EOF+SIGTERM ignored), detached helper leak -> all red.

(cherry picked from commit 9e1ab342650fcbee791f0cc35a3ec679052e9492)
…tegrity (C6)

Class C6 (bricked installs / stale modules after hermes update): every module the
package ships is imported from a clean first-party sys.modules in a fresh sandboxed
interpreter; a stale-graph leg replays the pre-handoff v2026.9.14 updater graph and
imports the post-purge restart modules (file_signature burst #111942 class); and a
matrix of real entrypoints (--version, doctor, -z one-shot against the fake provider
with state.db integrity_check, serve READY on stdout + clean SIGTERM with no orphan,
console scripts) runs in fresh processes. Every process is bubblewrap-sandboxed (own
PID namespace, no user systemd bus, real ~/.hermes read-only) so no probe can reach a
live install.

(cherry picked from commit 6635ace39ae6c04d9391da7e1321277a28f8b219)
…oard (C18)

Class C18 (config/settings persistence): seeded-random property tests on the real
load/save/set path. Round-trip byte stability of untouched keys, config set changes
exactly one key (in-process, hermes config set subprocess, tui_gateway stdio JSON-RPC,
dashboard PUT /api/config), a failed/empty read never clobbers config.yaml (#113301
class), .env loading idempotent incl. self-references (#109902), every migration from
all historical _config_version values idempotent, and every top-level section tolerates
null. Open bugs found are pinned as strict xfail (TUI/dashboard empty-read collapse,
save_config dropping explicit nulls, #119844 refold, #119928).

(cherry picked from commit 66f808b2ec64a80a3c270e3e8d8e99c9d33ade9a)
…nterrupted-update legs (C6)

Class C6 (bricked installs, stale code, lost state after hermes update). A local bare
origin (--shared, no network) parks main at release N-1; a git-mode install with its
own uv venv gets user state written by the N-1 CLI itself (sessions, a named profile,
a cron job, a hand-edited config with comments / long quoted values / a legacy MCP
disabled flag), then the real `hermes update --yes` runs to HEAD.

Legs: clean, local edits + orphan autostash, SIGKILL mid-fast-forward (lock-only must
heal; torn tree pinned as a strict xfail: it bricks every entry point), SIGKILL when the
dependency sync starts (new code on the old venv must heal), origin unreachable (must
fail non-zero and change nothing). After each: exit code matches reality, the venv
serves HEAD's tree (every top-level package + HEAD-only modules import from outside
the checkout) and satisfies HEAD's dependency set, --version / doctor / one-shot turn
succeed in fresh processes, both profiles' state.db pass integrity_check with unchanged
rows and byte-identical messages, the cron job survives, and config.yaml equals HEAD's
own non-interactive migration of the pre-update bytes (plus independent value/comment
checks). Everything runs in the bubblewrap sandbox so the updater's all-profile gateway
scan and systemctl calls can never reach a live install.

(cherry picked from commit 9b275a1dfc644bdddf6ee77091a7126aeb536722)
…rker

Classes: C9 provider wire-format drift, C17 prompt-cache hits, C11 real
auth / credential routing / `/models` parse.

Mocks encode our own belief about each vendor's wire format; vendor-side
schema changes, reasoning-replay rules, streaming shape changes and cache
behaviour are only observable against the real APIs. This drives the REAL
AIAgent + real adapters (chat_completions, anthropic_messages via OpenRouter,
codex_responses for xAI/OpenAI, GeminiNativeClient) in a temp HERMES_HOME
through a scripted 3-turn conversation with a deterministic registered tool:

  turn 1  one forced tool call; valid JSON args; result round-trips
  turn 2  two tool calls in ONE assistant message; both results round-trip
  turn 3  no tools; answer from replayed history (tool results + reasoning
          replay) under a cache breakpoint

Invariants per turn: no 4xx on any agent-loop request (429 excepted), the
turn completes, no <think>/DSML/control-token/raw tool JSON in user text,
the credential resolved from env is the one sent and ONLY to the provider's
host, tool results are persisted to state.db. Anthropic-family routes also
assert 1..4 well-formed breakpoints (none on role:tool / inside
tool_result.content[]) and cache-read tokens > 0 on turn 3 (warm once).
A `/models` listing check per provider uses the live fetchers without the
curated fallback (Gemini is strict-xfail on open #62259).

Spend: cheap models, <= 3 turns, max_tokens 2048, max_iterations 6, a hard
per-test token and $ guard; usage + estimated $ printed per test and
appended to $HERMES_LIVE_USAGE_FILE. Every case skips cleanly without its
key. `live` is excluded by default addopts; select with `-m live`.

(cherry picked from commit 3bd7de6991bcb03bfe5d11945de6f2317c534b7d)
Runs tests/e2e/core/live (`-m live`) nightly, on `v*` tags and on
workflow_dispatch (optional -k filter). Secrets-gated on LIVE_*_API_KEY
repo secrets (each case skips without its key; the job no-ops when none
are configured), main-repo only, one run per ref (never cancels a release
gate), 25-minute timeout. Uses direct pytest because scripts/run_tests.sh
starts from `env -i` so no credential can reach a test. Publishes a
usage/cost table to the step summary and uploads junit + usage JSONL.

(cherry picked from commit 75c5656ff3905512bf93c1fd887bbd5e85396f29)
… their fix PRs

The C7/C11 suites were built on test/core-tenancy next to seven production
fixes that land as their own PRs. Without them, exactly four cells fail on
base: the bare-custom OpenRouter key leak, the TUI /model switch legs
(explicit --provider vs alias, Anthropic refresh onto a foreign host), the
gateway's shared 'default' terminal env for routed cron work, and the Desktop
backend's launch-profile model/.env. Each is strict, so it flips red once its
fix lands and the mark comes off in the same PR.
…ite in its own job

tests/e2e/core/terminal drives the real `hermes --tui` over a PTY, so the e2e
job now installs the Node workspaces and builds ui-tui, and
HERMES_E2E_REQUIRE_TUI=1 makes a missing build fail instead of skip.

tests/e2e/core/upgrade runs a real N-1 -> HEAD `hermes update`: it needs full
history + tags, bubblewrap (every updater runs sandboxed so it can never reach
a real gateway or systemd), the warm uv cache, and up to ~15 min for one file.
It gets its own 60-minute job instead of stretching the e2e job.
…the settle race

An approval.respond RPC can end the wait before the gateway attaches the hook
that withdraws the sent request, so it stays in open_requests. Red in 1 of 4
loaded union runs, deterministic with the gap widened, green with #120374.
A race cannot be strict; drop the mark when #120374 lands.
…hrough the hermetic env

run_tests.sh starts every run from env -i with an allowlist, so the e2e
job's HERMES_E2E_REQUIRE_TUI=1 never reached the terminal suite (a missing
ui-tui build skipped instead of failing) and the upgrade suite's CI branch
in sandbox_required_reason() was dead. With ui-tui/dist removed and
HERMES_E2E_REQUIRE_TUI=1: before, 2 skipped; after, 2 failed.
The six provider keys sat in job-level env, so every step saw them: uv sync
(and any sdist build backend it runs), setup-uv, checkout and the retry
action. The job now carries only secrets.X != '' booleans for the gate, and
the values are set on 'Run live canaries' alone. actionlint clean.
…safe probes

A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and
flips main red (XPASS) the moment its fix merges. Each gap now has a probe that
reproduces the defect's mechanism on the tree under test; the xfail applies only
while the probe reports the defect open, and only for the bug-specific exception
class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed,
RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
…ials error within 90 s

rc != 0 also counted the harness's SIGKILL (-9) after a 240 s hang as a fail-fast.
The cell now needs a real non-zero exit AND the resolver's
"provider 'custom' resolved without credentials" message, under a 90 s leg
timeout (the fixed tree exits in ~5 s).
…s own uv.lock

Collection no longer runs git (every CI shard imported the module and paid for
three git calls, and a git failure broke collection instead of skipping). The N-1
venv now comes from N-1's uv.lock via `uv sync --locked --extra all` into
install/venv with the user's uv config hidden, the installer's tier 0, instead of
a fresh resolve of .[all] against today's PyPI.
It scanned the connection's frames from 0, so any earlier turn's delta on the same
session matched at once and the interrupt/drop op fired before the new turn
streamed.
…suite's own probe

The terminal job only needs the Ink TUI, not every workspace (desktop/electron).
The e2e-upgrade bwrap check used different flags from _bwrap_usable (no --proc,
no --die-with-parent), so it could pass while the suite fell back to running the
real updater unsandboxed; it now asserts _helpers.BWRAP_OK itself.
…ng static xfails

#120319, #120374 and #120299 are on main, so their probes and every Gap
naming them go (the module's own rule); those cells are plain tests now.

The two static strict xfails with an open fix PR (#95375 cli
resize_scrollback, fix #120321; #62259 Gemini listing, fix #62267/#116509)
turned main red the moment the fix merged (XPASS). They now go through
_pending_fixes.known_failure: a run-time xfail only while the cell fails
with that gap's own message (a turn rendered more than once; an empty live
listing), any other failure stays red, and the fix just makes it pass.
… rest are message-gated

CI (the PR merged with main) went red with 26 XPASS(strict): main fixed
the long-scalar re-fold (#119844, 7f59505), the explicit-null strip
(93c5856), the phantom `agent: {}` (0698be8), the transient-read clobber
on the TUI RPC and dashboard PUT (db2f07c/061da070) and the
`hermes-agent --help` argv bug (#54648). Their marks and the workarounds
that kept them out of other cells are gone.

Still-open gaps XFAIL only on their own failure message
(_pending_fixes.known_failure): #119928 env-lock refusal reported as
success (fix #119929) and the torn-tree update dying at import (fix
#120339). The dashboard explicit-null cell was only xfailing because seed
605 no longer generates a null leaf; it uses seed 607 and passes.

_base_config_version imports N-1's DEFAULT_CONFIG in the N-1 venv instead
of string-matching `git show` source.
…own frame index

Turn.frame_idx was the submitter's frame count, but op_shared_owner_drop
reads the watcher's connection, whose history is shorter: the scan
started past the watcher's early deltas and timed out (seed=11,
9:shared_owner_drop:A, ~1 in 4 runs). Record every attached connection's
frame count at submit and scan from the one being read.
…wn_failure patterns

#120307 and #120295 merged, so their cells are plain tests and the probe
machinery in core/_pending_fixes.py has no consumer left (known_failure stays).
torn-tree also excuses 'No module named' (a later release pair can die on a new
module instead of a moved name); the gemini listing gate now requires the 401
from the models endpoint, so an outage or 5xx fails the cell.
@teknium1
teknium1 force-pushed the tests/core-e2e-wave2 branch from 35f3be7 to 2a55a8b Compare September 24, 2026 00:46
@teknium1
teknium1 merged commit 76aebb4 into main Sep 24, 2026
37 checks passed
@teknium1
teknium1 deleted the tests/core-e2e-wave2 branch September 24, 2026 00:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/profiles Multi-profile isolation, HERMES_HOME scoping area/sessions Session lifecycle, resume, persistence, history ci-reviewed applied to manually approve dangerous changes comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/desktop Electron desktop app (apps/desktop/*) comp/gateway Gateway runner, session dispatch, delivery comp/tui Terminal UI (ui-tui/ + tui_gateway/) P3 Low — cosmetic, nice to have sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state tool/terminal Terminal execution and process management type/test Test coverage or test infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants