test: E2E suites for profile isolation, key routing, terminal transcripts, hermes update and live providers (core E2E wave 2) - #120326
Merged
Merged
Conversation
૮ >ﻌ< ა ci reviewran on 2a55a8b — test(e2e): drop the #120307/#120295 gaps now on main; tighte ℹ️ InfoCI-sensitive file review · View jobPR touches sensitive files, but the Sensitive files changed: debug infoCI timingsCI timings · View report · View jobWall time 7m15s vs 6m8s (+18.2%). 12 job(s) slower, 7 faster,
|
teknium1
force-pushed
the
tests/core-e2e
branch
from
September 23, 2026 16:50
2cf408e to
f601e22
Compare
teknium1
force-pushed
the
tests/core-e2e
branch
2 times, most recently
from
September 23, 2026 18:10
5535d03 to
276f56e
Compare
teknium1
force-pushed
the
tests/core-e2e-wave2
branch
from
September 23, 2026 18:17
92d6040 to
b0b886b
Compare
teknium1
force-pushed
the
tests/core-e2e-wave2
branch
from
September 23, 2026 22:32
b0b886b to
28145e7
Compare
teknium1
force-pushed
the
tests/core-e2e-wave2
branch
from
September 24, 2026 00:22
28145e7 to
35f3be7
Compare
Class C11 (provider/model routing + credential resolution). One real loopback OpenAI-compatible host per provider identity, each accepting only its own key; every scenario configures all identities and only changes the selection, and one invariant runs after every leg: requests land only on the selected hosts, each with only that host's key, the user sees that host's answer, and nothing egresses to a real inference API (CONNECT trap). Real `hermes chat -q` / `hermes -z` subprocesses cover custom base_url (key_env / literal), bare custom fail-fast, named providers: and legacy custom_providers entries, startup -m aliases (#103933/#107191/#109440), fallback after 401/429, credential pool rotation / single-key exhaustion, per-profile routing, aux title + delegation routing. The real stdio tui_gateway walks mid-session /model switches incl. pool recovery without restart. fake_llm_provider: api_key may be a list of accepted keys (pool on one host) and opt-in record_get; defaults unchanged. (cherry picked from commit bf1f07571d7777fe91b16f386e026c12bdbfe73b)
Class C7 (multiplex / tenancy isolation). A real 'hermes gateway run' with gateway.multiplex_profiles serves default + alpha + beta, each with its own loopback provider (accepting only its own key) and distinct canaries: provider key under the same env var name, API-server key, .env marker, model, MEMORY.md, SOUL.md, terminal cwd, cron prompt and an exported shell variable. Interleaved api_server turns (/v1, /p/<name>/v1) each run 'env | sort; pwd' in the terminal tool; a cron tick fires one job per profile; cross-profile API keys; profile create/attach/delete next to the live host (PID must not change); SIGTERM restart + a second cron tick. After every phase: no request, tool-subprocess env snapshot, profile-home file (state.db + WAL, logs, sessions, config, .env) or gateway output carries another tenant's canary. Red on the parent of a445f346a8b (cron shared the launch profile's terminal environment); see REPORT.md for the red-proof table. (cherry picked from commit ba5bdae9aa9900d719be36bf9efa40037860cf0b)
Class C7 (multiplex / tenancy isolation). The real app backend, 'hermes serve'
with HERMES_DESKTOP=1 (in-process cron ticker), driven over /api/ws exactly like
the Desktop: session.create {profile}, interleaved prompt.submit turns, a cron
tick per profile, session-bound and profile-bound config.set, session-bound
model.save_key, profiles.create + 'hermes profile delete' next to the live
backend (PID unchanged), restart + session.resume + a second cron tick. Same
canary invariant as the gateway suite plus: each settings write changes only
the addressed profile's config.yaml/.env, and no RPC reply/event for a session
carries another tenant's canary.
Red on the parents of 6952f1d5425 (lazy gateway.run import latched a secondary's
terminal.* into the launch env), 5a14abea326 (save_key wrote the launch .env) and
1e3727666e5 (secondary session reported/persisted the launch model).
(cherry picked from commit 522858661a60abc47ece47d977d73c54067789bc)
…ckend
Class C14 (multi-client session ownership / routing / switching): prompts persisted
to a stale session after a switch, events rendered twice or leaking into another
chat, zombie leases ("already has a live owner"), a ws drop or client absence
killing the in-flight turn, a live session hard-deleted.
tests/e2e/core/terminal/_gateway_client.py spawns ONE real
`python -m hermes_cli.main serve --port 0` per module (Desktop's argv) in an
isolated HOME with only the recording fake provider, and drives it with
Desktop-shaped WebSocket JSON-RPC clients (gateway.ready, client.capabilities,
server->client request answers, abrupt TCP drop, reader pause for absence).
test_multiclient_session_model.py runs a seeded fuzzer (seeds 11/23/37) over 3
clients: create, switch(session.resume), fast and slow prompts, interrupt
mid-stream, delete of live (must be refused) and closed sessions, ws drop +
reconnect + resume mid-turn, private-chat owner death (idle and running) +
takeover, shared-chat owner death, client absence. A reference model checks after
every step and at the end: prompts persisted exactly once, in order, to the
session selected at send time; per-connection event seq strictly increasing (no
duplicate delivery); no events for sessions a connection never attached to;
canaries only in their own session; one live runtime per chat; orphaned runtimes
released in bounded time and resumable/promptable by another client; drop/absence
never cut a turn short (full reply persisted, message.complete status complete).
Red-proof (each fails, restored after): duplicate event write in write_json;
semantic revert of 67de938 (#98028/#100325, orphan reaper ignores turn
activity); semantic revert of de25545 (rebind instead of fan-out); lease never
released; session.delete of a live session not refused.
(cherry picked from commit fe078a891231d3ba459da4ca46bf2a06979f4444)
Class C20 (interactive prompts): a clarify question or dangerous-command approval
never shows, is lost on a session switch / reconnect / the other window, the answer
never reaches the tool, the approval outcome is ignored (denied command runs), or a
prompt stays pending after the turn.
test_interactive_roundtrip.py reuses the lane's real `hermes serve` harness with
approvals.mode=manual (not yolo). The fake provider issues clarify and
terminal(`rm -rf <victim>`) tool calls and records the tool result the agent sends
back. Matrix {clarify, approval-deny, approval-approve} x {answer at once, after
switching away and back (re-delivered via session.resume open_requests), from a
second connection, after an abrupt ws drop + reconnect, via the approval.respond
RPC}. Asserts the server->client request arrives in bounded time, the answer
reaches the tool (clarify user_response / "denied by user" + victim intact /
exit_code 0 + victim gone), the other window gets a card-settling signal,
approval.pending and open_requests are empty afterwards, and the turn completes.
Red-proof (each fails, restored after): clarify callback drops the answer;
approval response choice ignored (deny runs the command); open_requests not
re-delivered on resume; approval.respond RPC routing dropped.
(cherry picked from commit 372e0c04e56c54ceb237937bae7a990212fcfbe6)
… Ink TUI) Class C2 (duplicate / vanishing / garbled transcript rendering) on the terminal surfaces. Drives a real `hermes chat --cli` and a real `hermes --tui` (Node frontend + tui_gateway child) as session leader of a real PTY against the recording fake provider (long multi-chunk stream, tool-call turn, reasoning turn, width resizes mid-stream via SIGWINCH), renders the byte stream through a dependency-free VT emulator (_vt.py, cross-checked identical to pyte on captured CLI and TUI streams) and asserts on the settled grid + scrollback: every reply exactly once and verbatim (== scripted stream), every prompt echoed once, conversation order, tool call once, reasoning at most once, rendered == persisted state.db rows, /exit -> 0 with no process left in the PTY session or among tracked descendants (setsid'ed helpers included). cli-resize_scrollback is a strict xfail (raises=DuplicateRender) for the open live bug #95375: a width-changing resize on a 24-row terminal re-prints turns that already sit in scrollback (warmup reply rendered 3x). Red-proof: re-render final panel after stream, double-appended deltas, tui_gateway double event emit, resize replay without clear, gateway outliving the TUI (EOF+SIGTERM ignored), detached helper leak -> all red. (cherry picked from commit 9e1ab342650fcbee791f0cc35a3ec679052e9492)
…tegrity (C6) Class C6 (bricked installs / stale modules after hermes update): every module the package ships is imported from a clean first-party sys.modules in a fresh sandboxed interpreter; a stale-graph leg replays the pre-handoff v2026.9.14 updater graph and imports the post-purge restart modules (file_signature burst #111942 class); and a matrix of real entrypoints (--version, doctor, -z one-shot against the fake provider with state.db integrity_check, serve READY on stdout + clean SIGTERM with no orphan, console scripts) runs in fresh processes. Every process is bubblewrap-sandboxed (own PID namespace, no user systemd bus, real ~/.hermes read-only) so no probe can reach a live install. (cherry picked from commit 6635ace39ae6c04d9391da7e1321277a28f8b219)
…oard (C18) Class C18 (config/settings persistence): seeded-random property tests on the real load/save/set path. Round-trip byte stability of untouched keys, config set changes exactly one key (in-process, hermes config set subprocess, tui_gateway stdio JSON-RPC, dashboard PUT /api/config), a failed/empty read never clobbers config.yaml (#113301 class), .env loading idempotent incl. self-references (#109902), every migration from all historical _config_version values idempotent, and every top-level section tolerates null. Open bugs found are pinned as strict xfail (TUI/dashboard empty-read collapse, save_config dropping explicit nulls, #119844 refold, #119928). (cherry picked from commit 66f808b2ec64a80a3c270e3e8d8e99c9d33ade9a)
…nterrupted-update legs (C6) Class C6 (bricked installs, stale code, lost state after hermes update). A local bare origin (--shared, no network) parks main at release N-1; a git-mode install with its own uv venv gets user state written by the N-1 CLI itself (sessions, a named profile, a cron job, a hand-edited config with comments / long quoted values / a legacy MCP disabled flag), then the real `hermes update --yes` runs to HEAD. Legs: clean, local edits + orphan autostash, SIGKILL mid-fast-forward (lock-only must heal; torn tree pinned as a strict xfail: it bricks every entry point), SIGKILL when the dependency sync starts (new code on the old venv must heal), origin unreachable (must fail non-zero and change nothing). After each: exit code matches reality, the venv serves HEAD's tree (every top-level package + HEAD-only modules import from outside the checkout) and satisfies HEAD's dependency set, --version / doctor / one-shot turn succeed in fresh processes, both profiles' state.db pass integrity_check with unchanged rows and byte-identical messages, the cron job survives, and config.yaml equals HEAD's own non-interactive migration of the pre-update bytes (plus independent value/comment checks). Everything runs in the bubblewrap sandbox so the updater's all-profile gateway scan and systemctl calls can never reach a live install. (cherry picked from commit 9b275a1dfc644bdddf6ee77091a7126aeb536722)
…rker
Classes: C9 provider wire-format drift, C17 prompt-cache hits, C11 real
auth / credential routing / `/models` parse.
Mocks encode our own belief about each vendor's wire format; vendor-side
schema changes, reasoning-replay rules, streaming shape changes and cache
behaviour are only observable against the real APIs. This drives the REAL
AIAgent + real adapters (chat_completions, anthropic_messages via OpenRouter,
codex_responses for xAI/OpenAI, GeminiNativeClient) in a temp HERMES_HOME
through a scripted 3-turn conversation with a deterministic registered tool:
turn 1 one forced tool call; valid JSON args; result round-trips
turn 2 two tool calls in ONE assistant message; both results round-trip
turn 3 no tools; answer from replayed history (tool results + reasoning
replay) under a cache breakpoint
Invariants per turn: no 4xx on any agent-loop request (429 excepted), the
turn completes, no <think>/DSML/control-token/raw tool JSON in user text,
the credential resolved from env is the one sent and ONLY to the provider's
host, tool results are persisted to state.db. Anthropic-family routes also
assert 1..4 well-formed breakpoints (none on role:tool / inside
tool_result.content[]) and cache-read tokens > 0 on turn 3 (warm once).
A `/models` listing check per provider uses the live fetchers without the
curated fallback (Gemini is strict-xfail on open #62259).
Spend: cheap models, <= 3 turns, max_tokens 2048, max_iterations 6, a hard
per-test token and $ guard; usage + estimated $ printed per test and
appended to $HERMES_LIVE_USAGE_FILE. Every case skips cleanly without its
key. `live` is excluded by default addopts; select with `-m live`.
(cherry picked from commit 3bd7de6991bcb03bfe5d11945de6f2317c534b7d)
Runs tests/e2e/core/live (`-m live`) nightly, on `v*` tags and on workflow_dispatch (optional -k filter). Secrets-gated on LIVE_*_API_KEY repo secrets (each case skips without its key; the job no-ops when none are configured), main-repo only, one run per ref (never cancels a release gate), 25-minute timeout. Uses direct pytest because scripts/run_tests.sh starts from `env -i` so no credential can reach a test. Publishes a usage/cost table to the step summary and uploads junit + usage JSONL. (cherry picked from commit 75c5656ff3905512bf93c1fd887bbd5e85396f29)
… their fix PRs The C7/C11 suites were built on test/core-tenancy next to seven production fixes that land as their own PRs. Without them, exactly four cells fail on base: the bare-custom OpenRouter key leak, the TUI /model switch legs (explicit --provider vs alias, Anthropic refresh onto a foreign host), the gateway's shared 'default' terminal env for routed cron work, and the Desktop backend's launch-profile model/.env. Each is strict, so it flips red once its fix lands and the mark comes off in the same PR.
…ite in its own job tests/e2e/core/terminal drives the real `hermes --tui` over a PTY, so the e2e job now installs the Node workspaces and builds ui-tui, and HERMES_E2E_REQUIRE_TUI=1 makes a missing build fail instead of skip. tests/e2e/core/upgrade runs a real N-1 -> HEAD `hermes update`: it needs full history + tags, bubblewrap (every updater runs sandboxed so it can never reach a real gateway or systemd), the warm uv cache, and up to ~15 min for one file. It gets its own 60-minute job instead of stretching the e2e job.
…the settle race An approval.respond RPC can end the wait before the gateway attaches the hook that withdraws the sent request, so it stays in open_requests. Red in 1 of 4 loaded union runs, deterministic with the gap widened, green with #120374. A race cannot be strict; drop the mark when #120374 lands.
…hrough the hermetic env run_tests.sh starts every run from env -i with an allowlist, so the e2e job's HERMES_E2E_REQUIRE_TUI=1 never reached the terminal suite (a missing ui-tui build skipped instead of failing) and the upgrade suite's CI branch in sandbox_required_reason() was dead. With ui-tui/dist removed and HERMES_E2E_REQUIRE_TUI=1: before, 2 skipped; after, 2 failed.
The six provider keys sat in job-level env, so every step saw them: uv sync (and any sdist build backend it runs), setup-uv, checkout and the retry action. The job now carries only secrets.X != '' booleans for the gate, and the values are set on 'Run live canaries' alone. actionlint clean.
…safe probes A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and flips main red (XPASS) the moment its fix merges. Each gap now has a probe that reproduces the defect's mechanism on the tree under test; the xfail applies only while the probe reports the defect open, and only for the bug-specific exception class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed, RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
…ials error within 90 s rc != 0 also counted the harness's SIGKILL (-9) after a 240 s hang as a fail-fast. The cell now needs a real non-zero exit AND the resolver's "provider 'custom' resolved without credentials" message, under a 90 s leg timeout (the fixed tree exits in ~5 s).
…s own uv.lock Collection no longer runs git (every CI shard imported the module and paid for three git calls, and a git failure broke collection instead of skipping). The N-1 venv now comes from N-1's uv.lock via `uv sync --locked --extra all` into install/venv with the user's uv config hidden, the installer's tier 0, instead of a fresh resolve of .[all] against today's PyPI.
It scanned the connection's frames from 0, so any earlier turn's delta on the same session matched at once and the interrupt/drop op fired before the new turn streamed.
…suite's own probe The terminal job only needs the Ink TUI, not every workspace (desktop/electron). The e2e-upgrade bwrap check used different flags from _bwrap_usable (no --proc, no --die-with-parent), so it could pass while the suite fell back to running the real updater unsandboxed; it now asserts _helpers.BWRAP_OK itself.
…ng static xfails #120319, #120374 and #120299 are on main, so their probes and every Gap naming them go (the module's own rule); those cells are plain tests now. The two static strict xfails with an open fix PR (#95375 cli resize_scrollback, fix #120321; #62259 Gemini listing, fix #62267/#116509) turned main red the moment the fix merged (XPASS). They now go through _pending_fixes.known_failure: a run-time xfail only while the cell fails with that gap's own message (a turn rendered more than once; an empty live listing), any other failure stays red, and the fix just makes it pass.
… rest are message-gated CI (the PR merged with main) went red with 26 XPASS(strict): main fixed the long-scalar re-fold (#119844, 7f59505), the explicit-null strip (93c5856), the phantom `agent: {}` (0698be8), the transient-read clobber on the TUI RPC and dashboard PUT (db2f07c/061da070) and the `hermes-agent --help` argv bug (#54648). Their marks and the workarounds that kept them out of other cells are gone. Still-open gaps XFAIL only on their own failure message (_pending_fixes.known_failure): #119928 env-lock refusal reported as success (fix #119929) and the torn-tree update dying at import (fix #120339). The dashboard explicit-null cell was only xfailing because seed 605 no longer generates a null leaf; it uses seed 607 and passes. _base_config_version imports N-1's DEFAULT_CONFIG in the N-1 venv instead of string-matching `git show` source.
…own frame index Turn.frame_idx was the submitter's frame count, but op_shared_owner_drop reads the watcher's connection, whose history is shorter: the scan started past the watcher's early deltas and timed out (seed=11, 9:shared_owner_drop:A, ~1 in 4 runs). Record every attached connection's frame count at submit and scan from the one being read.
…wn_failure patterns #120307 and #120295 merged, so their cells are plain tests and the probe machinery in core/_pending_fixes.py has no consumer left (known_failure stays). torn-tree also excuses 'No module named' (a later release pair can die on a new module instead of a moved name); the gemini listing gate now requires the 401 from the models endpoint, so an outage or 5xx fails the cell.
teknium1
force-pushed
the
tests/core-e2e-wave2
branch
from
September 24, 2026 00:46
35f3be7 to
2a55a8b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Four more end-to-end suites now guard classes of bugs users keep hitting: one profile's keys, model, cwd or
.envreaching another profile; a request sent to the wrong host with the wrong key; a reply shown twice (or not at all) in the terminal; and an install broken byhermes update.Rebuilt on
mainafter #120171 (tests/core-e2e, merged as 2b1bb70):origin/main+ this PR's own commits only.What it adds
All suites use real processes, the recording loopback provider (
tests/fakes/fake_llm_provider.py), a tempHOME, and no mocks of Hermes code.tenancy/test_routing_truth_table.pyhermes chat -q/-zconfigs (custom base_url, key_env vs literal, named and legacy providers, aliases, 401/429 fallback, 1- and 2-key pools, per-profile, aux, delegation) + one live TUI session walking 10/modelswitchesLEAK/MISSINGfailure; a CONNECT trap catches real egresstenancy/test_two_tenant_gateway.pyhermes gateway runserving 3 profiles: interleaved api_server turns, cron fires, cross-key 401s, profile create/delete next to the live host, SIGTERM restart.env) or backend log carries another tenant's canarytenancy/test_two_tenant_desktop_backend.pyhermes serve(Desktop) on/api/ws:session.create {profile}, settings RPCs,model.save_key, restart +session.resumeconfig.yaml/.envchanges, byte-for-byteterminal/test_terminal_transcript_pty.pyhermes chat --cliandhermes --tuion a real PTY, rendered by an in-tree VT emulator: long streams, tool calls, reasoning, width changes mid-stream/exitleaves no processesterminal/test_multiclient_session_model.pyterminal/test_interactive_roundtrip.pyapproval.respond, with a realrm -rfunder manual approvalsupgrade/test_upgrade_path.pyhermes update --yesfrom the previous release tag: clean, autostash, SIGKILL mid-pull, SIGKILL before deps, offline.env, cron ids and migrated config survive; no service-manager restart escapes the sandboxupgrade/test_fresh_process_entrypoints.py[project.scripts]entry andserve's READY/SIGTERM handshakeupgrade/test_config_roundtrip_properties.pyhermes config set, TUIconfig.set, dashboardPUT /api/config, migrations from every historical version, transient EMFILElive/test_live_providers.py/modelsparse plus a 3-turn tool conversation each, with dollar guardsThe live suite sits behind a new
livemarker, excluded by default (addopts = "-m 'not integration and not live'"). Each case skips when its key is absent.Unlanded fixes: merge-order-safe expected failures
No static
xfail(strict=True)is left in these suites. A static strict xfail turnsmainred (XPASS) the moment its fix merges (this PR hit exactly that: 26 config round-trip cells XPASSed once main fixed their bugs). Every remaining expected failure is merge-order safe, one of two ways, both intests/e2e/core/_pending_fixes.py:Gap+expect_gaps, same pattern as Gateway replies, cron deliveries and the boot outbox are proven exactly-once through faults, crashes and DST (delivery E2E suites + CI) #120344): a few lines reproduce the defect's mechanism in a throwaway interpreter with its ownHOME; the xfail applies only while the probe reports it open, and only for the exception class the cell raises when it sees THAT leak. A broken probe fails the cell.known_failure): a run-time xfail only while the cell fails with the gap's own assertion message; any other failure stays red, and once the fix lands the cell simply passes.No test reads source code.
test_tui_gateway_model_switch_routing(leg 5)/model --provider Xuses an alias's endpoint and keyRoutingLeaktest_multiplexed_gateway_never_crosses_tenantsTenantLeaktest_desktop_backend_never_crosses_tenants.env(a race)TenantLeaktest_terminal_transcript_integrity[cli-resize_scrollback]DuplicateRenderonly)test_kill_mid_pull_then_retry_heals[torn-tree]final hermes update failed … cannot import nametest_p2_env_lock_refusal_is_not_reported_as_success[*].envwrite exits 0exit=0 but .env unchanged/refused env write still rewrote config.yamltest_models_listing_parses[gemini](live)/v1beta/modelsProbes for fixes that have since merged were deleted with their gaps (#120319, #120374, #120299); those cells are plain tests. Cells whose config/entrypoint bug main fixed are plain strict tests with their workarounds removed: #119844 long-scalar re-fold (config round-trip +
test_clean_update_keeps_long_scalar_lines_verbatim), explicit-nullstrip, phantomagent: {}, transient-read clobber on TUI RPC / dashboard PUT, #54648hermes-agent --help.CI wiring (
.github/workflows/tests.yml)e2ejob:npm ci --workspace ui-tui+ buildsui-tui, and setsHERMES_E2E_REQUIRE_TUI=1so a missing TUI build fails instead of skipping (scripts/run_tests.shnow forwardsHERMES_E2E_REQUIRE_TUI,CIandGITHUB_ACTIONSthrough itsenv -i; before, the flag never reached pytest and the TUI cells skipped). It runs everytests/e2efile exceptcore/upgrade.e2e-upgradejob: 60 min, 32-core,fetch-depth: 0+ tags (N-1 =git describe --tags HEAD~1). It installs bubblewrap and lifts the Ubuntu 24.04 userns restriction, then asserts the suite's own_bwrap_usable()probe (same flags), so every spawned updater runs in its own PID namespace with no systemd bus. N-1's venv is built from N-1's ownuv.lock(uv sync --locked --extra all, the installer's tier 0), and HEAD/N-1 refs resolve lazily (no git at collection). It keeps the warm uv cache and uses a 3000 s file budget.All required checks passalready covers it throughpython-tests.live-providers.yml: runs nightly (07:17 UTC), onv*tags and on dispatch. It only runs on the main repo, and it no-ops withoutLIVE_*_API_KEYsecrets. It calls pytest directly becauserun_tests.shstrips credentials. The key values are scoped to the oneRun live canariesstep (job level carries onlyHAS_*booleans), souv syncbuild backends and third-party actions never see them. It posts a usage/$ step summary.Validation
origin/main529d27e, every PR suite viarun_tests.sh --include-integration,--file-retries 0HERMES_E2E_REQUIRE_TUI=1); live: skipped without keysXPASS(strict)(24 config round-trip,hermes-agent-help, long-scalar update)test_multiclient_session_model.pyafter thewait_first_deltaper-connection index fixHERMES_E2E_REQUIRE_TUI=1throughrun_tests.shwithui-tui/dist/entry.jsremoved-m live: 17 skipped; default: 17 deselectedruff,check_no_tmp_literals,check-windows-footguns --all,check_public_surface,git diff --checkC20
approval_deny-rpcrace: fixed on main by #120374; the cell is a plain test now.Wave-1 flakes seen under load (not wave-2 files, not fixed here):
compaction/test_compaction_manual.py/test_compaction_auto.py: "the session system prompt was rewritten". This is a production bug. The workspace git snapshot is supposed to stay fixed for the session, but it is re-probed on rebuilds: manual/compressbinds the cwd after the first build, so the cwd key changes from""to the path; and a resumed agent never restores the saved snapshot. A checkout whose git status changes mid-session (for example, one more untracked file) then changes the prompt. It reproduces every time when a file is written into the repo after turn 0. A proposed ~30-line fix inagent/system_prompt.py+conversation_loop.pyhas red/green evidence, but no PR yet.sqlite/test_torture_chamber.py[fts_corruption_fail_open]: a writer gotdatabase disk image is malformedat load ~200, andkill9_everythingthen saw its leftover gateway. Seen once and not triaged.Each lane's own record: 10/10 consecutive green per file with retries disabled, and the red proofs are in the lane reports (tenancy 14, terminal 15 + 1 live bug, upgrade 28 + 5 fix-flips).
Not covered: MCP bearer tokens per profile, messaging adapters other than api_server, Windows/macOS update paths, gateway service restart after an update, OAuth device-code refresh, and credential pools with more than 2 keys. Each lane's
NOT_COVERED.mdlists them.Infographic