test: E2E known-bug pins no longer turn main red when their fix lands - #121499
Merged
Merged
Conversation
૮ >ﻌ< ა ci reviewran on 9176750 — chore: retrigger CI (ci.yaml startup failure on previous hea debug infoCI timingsCI timings · View report · View jobWall time 7m33s vs 7m29s (+0.9%). 1 job(s) slower, 12 faster, 1 unchanged.
|
teknium1
added a commit
that referenced
this pull request
Sep 24, 2026
…lass Review of #121499: a slow-but-healthy dashboard install, partial pricing, a nested permissions schema error, and a refusal that omits the holder pid no longer XFAIL as the pinned bug.
teknium1
force-pushed
the
tests/e2e3-known-merge-order-safe
branch
2 times, most recently
from
September 24, 2026 13:07
5548b38 to
f64c8d5
Compare
teknium1
added a commit
that referenced
this pull request
Sep 24, 2026
…lass Review of #121499: a slow-but-healthy dashboard install, partial pricing, a nested permissions schema error, and a refusal that omits the holder pid no longer XFAIL as the pinned bug.
…lass Review of #121499: a slow-but-healthy dashboard install, partial pricing, a nested permissions schema error, and a refusal that omits the holder pid no longer XFAIL as the pinned bug.
teknium1
force-pushed
the
tests/e2e3-known-merge-order-safe
branch
from
September 24, 2026 13:59
7c96689 to
9176750
Compare
teknium1
added a commit
that referenced
this pull request
Sep 24, 2026
…lass Review of #121499: a slow-but-healthy dashboard install, partial pricing, a nested permissions schema error, and a refusal that omits the holder pid no longer XFAIL as the pinned bug.
iankoratskydcn
added a commit
to iankoratskydcn/hermes-agent
that referenced
this pull request
Sep 24, 2026
* fix(import): refuse a damaged backup archive before touching the home
`hermes import` only checked the central directory (`is_zipfile`,
`namelist`), so an archive with one member whose deflate stream or CRC is
rotten passed validation and blew up mid-restore with a zlib.error
traceback -- after config.yaml and everything before the bad member had
already been replaced, with later members never written (#121258).
Add a pre-flight pass in run_import that streams every member through
1 MiB reads (zipfile verifies the CRC at EOF) and collects every
BadZipFile / zlib.error / EOFError. If any member is damaged the command
prints a capped list and returns 1 with the home untouched. Stdlib
`ZipFile.testzip()` is deliberately not used: it lets zlib.error escape
and names at most the first bad member.
Widen the per-member catch from the previous commit with EOFError so a
member that rots between the two passes still becomes a "skipped" warning
+ `Import incomplete` / exit 1 instead of a traceback. Rework that
commit's test to drive the per-member path (the archive now never reaches
it with a corrupt member), and let `_break_member` serve the pre-flight
read before failing the restore's own read.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
* fix(import): pre-flight refuses every member read error, ignores skipped members
The integrity pre-flight caught only BadZipFile/zlib.error/EOFError, so a
bzip2 member with a bad stream (OSError "Invalid data stream"), an lzma
LZMAError or a media read error still escaped as a traceback instead of the
clean "archive is damaged" refusal. Name the archive-read errors once
(_ZIP_MEMBER_READ_ERRORS) and use the same tuple, plus OSError, in the
pre-flight and in the per-member catch of the restore (PermissionError is an
OSError, so it no longer needs listing).
The pre-flight also decompressed members the restore never writes
(gateway.pid and the other runtime files, archived SQLite sidecars), so a rot
in one of those refused an otherwise restorable backup. Move the skip rule
into _import_skipped() and use it in both places so they cannot drift.
run_import's docstring now states the real contract: 1 for a damaged archive
or an incomplete restore, None on success or a declined prompt.
* test(import): fold damaged-archive tests into two invariants
Keep one test per invariant: "a damaged archive is refused before anything
is written" (parametrized over a bad deflate stream, a stored CRC mismatch,
a bad bzip2 stream and a cut-short deflate stream, each with 12 bad members
so the capped all-members listing is checked in the same run) and "a member
that rots after the pre-flight is reported as incomplete" (now also carrying
a damaged gateway.pid, which the pre-flight must not refuse on). The separate
listing test and the fake-zip EOFError unit test are folded in; the rot test
raises EOFError through the per-member catch, so dropping EOFError from
_ZIP_MEMBER_READ_ERRORS still fails a test.
* refactor(import): share the restore's member filter with the integrity pre-flight
* test(e2e): #121258 and #119953 are fixed; drop their strict known-failure markers
The upgrade lane's import E2E would XPASS(strict) on this branch: a rotten member
is refused before any write with exit 1, and a skipped member reports
'Import incomplete' with exit 1.
* fix(compression): proactive prune and micro-compaction keep turns another surface appended
Both commits call archive_and_compact with no watermark, which archives
every active row. A turn another surface appended to the same session
after this process loaded it (a Desktop session continued from Telegram)
was archived with the rest as compacted (active=0, compacted=1): marked
summarized away though no summary holds it. The display, REST with
include_compacted and session_search still show it, but it is gone from
the model's history, so the agent forgets a turn the user can still see.
Micro-compaction also runs a slow aux summary call before its commit, so
a turn that arrived during it went the same way.
Both now cap the archive at the newest row the process held, the rule the
in-place compaction commit applies, so those rows take the concurrent-
append path and are cloned after the new set. Micro-compaction captures
the held history and the store's watermark before the summary call and
caps at commit. The cap moves into held_archive_watermark(session_db,
session_id, ...), which the compressor paths can call; _held_watermark
stays as the in-place commit's wrapper.
A store without get_active_message_watermark or get_message_role keeps
today's archive-everything commit, as a store without archive_and_compact
already skips the prune. Micro-compaction's DB sync runs under a broad
except, so an unguarded call there would have failed silently.
Held lists without row ids (the gateway's replay dicts) keep today's
behaviour, as in the in-place commit. Both features are opt-in
(compression.proactive_prune_tokens > 0, compression.micro_compact).
(cherry picked from commit 8d67bee760db0fc65d4330df15198ddfd5f24480)
* fix(compression): a watermark the store cannot answer for skips the pass
_micro_start_watermark collapsed two different answers into None: "this
store has no watermark API", where the commit keeps today's archive-
everything behaviour as it always has, and "the read raised". None means
archive every active row, so a transient store failure committed exactly
the unbounded archive this watermark was added to prevent — a turn another
surface appended vanished from active model history when the read failed
and the write then succeeded.
The helper now returns (status, watermark), and a failed read skips the
pass rather than committing without a bound. The proactive prune path was
already safe: its watermark is resolved inside the try that wraps the
commit, so a raised read aborts the commit instead of widening it.
The new test drives a real SessionDB whose watermark read raises: the pass
must not run and the foreign turn must stay live. Red without the guard.
(cherry picked from commit dd5179e3c72552e6e7737575cc305459728fc549)
* fix(compression): a stale prune/micro-compaction generation aborts instead of publishing beside the winner
Prune and micro-compaction hold no compression lease. When the newest exact row of the history they
hold is no longer active, another compaction (a /compress on this or another surface, or an earlier
pass) has already committed. `held_archive_watermark` then falls back to the start watermark, which is
right for the in-place commit (its lease rules out overlap) but for a lease-less caller archives the
winner's rows and clones them back as a "concurrent tail": two summary generations live in one session.
`_archive_watermark_for` now asks `held_archive_watermark` to raise `StaleHeldHistory` on that branch.
The prune commit returns its input unchanged (`prune:stale_generation`), micro-compaction checks before
the slow summary call and skips the pass (telemetry `stale_generation`), and its commit re-checks so a
compaction that lands during the summary call is not overwritten either. The in-place commit keeps its
fallback.
Raised on #120634 by @ehz0ah (agent/conversation_compression.py:3623 thread) and independently by
@JoaoMarcos44 in #120821.
Co-authored-by: ehz0ah <haozhe4547@gmail.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
* fix(compression): a micro-compaction pass that loses the race is a true no-op
A compaction landing during the micro summary call made _sync_micro_compact_to_db
skip its write, but _micro_compact still returned the spliced list, advanced the
cursor and reported "absorbed". finalize_turn's _persist_session flush then
appended the unmarked summary row as live, beside the winning generation.
The commit-time re-check also never fired in the common case: held_archive_watermark
returned early when newest_held >= the start watermark, which is always true when
nothing was appended. A lease-less caller (stale_raises=True) now checks the newest
held row's liveness regardless; the in-place lease path is unchanged.
_sync_micro_compact_to_db now reports a stale abort, and _micro_compact then returns
the original list, restores the rolling summary/cursor (and a defragged marker), and
emits stale_generation telemetry. _micro_start_watermark returns
(skip_outcome, watermark) so the two identical pre-flight skip branches collapse.
* fix(compression): a failed watermark read takes prune's logged no-op path
_archive_watermark_for sat outside prune's try, so a raising
get_active_message_watermark/get_message_role escaped prune_tool_results_only and
surfaced only at debug level in turn_preflight, skipping the "keeping the original
transcript" warning. Read it inside the existing try, with StaleHeldHistory caught
ahead of the generic handler. Also add the missing E302 blank line.
* test(compression): parametrize the held-watermark invariants over prune and micro
The stack added five near-duplicate tests (two invariants x two call sites plus a
watermark-failure case). Collapse them into the two invariants AGENTS.md asks for,
parametrized over prune/micro on a real SessionDB: a late-appended turn is kept
exactly once, and a stale generation aborts leaving the winner the only live
version. The stale case now also covers a compaction landing during the micro
summary call, followed by a _persist_session-equivalent flush.
* test(compression): pin the micro pass no-op when the watermark read fails
* fix(memory): a journey node id names the card's text, not just its index
The edit request still carried only the displayed index, so the entry it
meant was rebuilt from whatever sat at that index when the mutation ran.
A writer prepending an entry between the graph being drawn and the edit
being submitted shifted the list: selecting "beta" then edited "alpha"
and left "beta" untouched.
The node id now carries a digest of the card's text
(memory:<source>:<index>:<fingerprint>). Resolution prefers the position
while it still holds that text, and otherwise finds the text, refusing
when it is gone or ambiguous. The locked re-resolution in _mutate_memory
(43d3d4e851) goes through the same _locate_memory, so it now names the
clicked card too. Ids from an older graph have no fingerprint and resolve
by position, as before.
(cherry picked from commit 4bdf50f5dadcbe84efe019c50ac5d54ef9548e6e)
* fix(memory): key the journey card lookups to the fingerprinted node id
Two places map a node id back to its memory card by rebuilding the bare
`memory:<source>:<index>` string: the timeline chart rows and the desktop
starmap's tooltip/body map. With the fingerprint in the id both missed
every card, so a memory row rendered an empty body.
Both now key the card under whichever shapes apply, so an imported or
pre-fingerprint graph keeps working. The CLI help and the memory doc now
describe the id as `journey list` prints it.
(cherry picked from commit fa0c984c5d7d7ad0820fa2f2bde088e532f4e897)
* fix(memory): build journey cards with the memory store's own parser
_memory_cards re-split MEMORY.md with plain utf-8 while MemoryStore reads
utf-8-sig, so a Notepad BOM stayed glued to the first entry: its card
title rendered with a U+FEFF, and — now that the node id carries a
digest of the card's text — its fingerprint never matched the store's,
so every edit/delete of that card failed with "stale — refresh the
graph" and refreshing could not fix it. Iterating
MemoryStore._read_file makes the fingerprint contract true by
construction and drops the duplicated delimiter literal.
* fix(memory): a journey edit honours the memory tool's char limit
_mutate_memory discarded the limit _mutate hands its closure and wrote a
replacement of any length. The oversize entry then read as external drift
to _detect_external_drift, so the memory tool's own remove/replace refused
with a fresh .bak on every call until the file was fixed by hand. Refuse
the edit the way the tool's replace does, with the same wording.
* fix(memory): a journey delete is never refused by the char limit
ba59c5292b applied the memory tool's char cap to every Journey mutation,
including delete. MemoryStore._edit checks the cap only on replace: a
file already over its total (lowered memory_char_limit, well-formed hand
edits, external writers) is a supported state and deleting entries is
how it gets back under. Journey is the pruning UI, so its delete was
refused with "Replacement would put memory at N/limit ... Shorten the
new content" where main deleted fine. Guard the check with
`replacement is not None`, as the tool does.
The existing char-limit test now also deletes from an over-cap USER.md
(red before this change, green after).
* refactor(memory): resolve a journey fingerprint by text alone
The occurrence hint in _resolve_fingerprint (_local_index_hint, "the
occurrence the user clicked") came from the contributor's dedupe=False
mutate, which this salvage deliberately dropped: MemoryStore._mutate
collapses byte-identical entries and _apply goes back to the first match
with entries.index(text), so the hint could never pick an occurrence. Its
only live effect was a spurious "stale" refusal for a duplicated card
after an earlier entry was removed, although the text was still there.
Match by fingerprint -> first entry with that text, delete
_local_index_hint (and with it the second copy of the profile-index
formula), and read _memory_cards() only on the legacy index-id path, so
a fingerprinted id no longer re-reads and hashes both files for a hint.
Legacy ids still resolve by position.
The render lookup goes back to a single memory_node_id key: every
render_frames caller builds the graph fresh, and build_learning_graph
always sets the fingerprint. The desktop star-map keeps its dual key
(imported graphs). The memory_fingerprint docstring said a memory add
prepends; it appends; what shifts a card is an earlier entry removed.
* test(memory): trim the journey identity tests to their invariants
AGENTS.md asks for at most two invariant tests per change and no
change-detectors; the salvage brought nine. Keep what goes red on main:
- one parametrized test: an edit/delete lands on the clicked card after
an earlier entry was removed, and repeating it with the same id is
refused as stale instead of hitting whatever sits at that index
(absorbs the vanished-target test);
- the char-limit test (now also covering delete) and the BOM block in
test_memory_writes_match_memory_tool_format (only guard for the
shared parser);
- the render body test, minus its `count(":") == 3` id-shape assert.
Dropped: the id-layout change-detector, detail/profile/long-entry and
legacy-id variants (legacy index ids stay covered by the existing
memory:memory:N / memory:profile:2 tests). The render test also patched
a nonexistent hermes_constants._cached_default_hermes_root with
raising=False (a no-op) and set HERMES_HOME by hand; it now uses the
hermetic conftest home like its siblings. Test wording no longer claims
a memory add prepends (it appends).
* chore: map rodricksz4h5's commit email for attribution
* fix(auth): keep Nous credentials when Vercel's edge checkpoint blocks a token refresh
Vercel's Security Checkpoint in front of portal.nousresearch.com answers
non-browser POST /api/oauth/token with a text/plain 403
(x-vercel-mitigated: deny) or a 429 challenge page. Those responses come
from the edge, not from the token endpoint, yet `_refresh_access_token`
folded the non-JSON 403 into the 401/403 -> invalid_grant default, so
`_refresh_nous_or_quarantine` wiped a still-valid refresh token and told
every affected user to re-login (#120602).
Classify a 403/429 carrying `x-vercel-mitigated` as `upstream_blocked`
(the code #115812 established for WAF blocks on the inference path):
retryable, relogin_required False, Retry-After forwarded through the
existing `parse_retry_after_seconds` helper. Nothing else moves: a 401
stays terminal even behind the header, a header-less non-JSON 403 keeps
the b8ce8875b0 stance (dead grant -> re-login), and 5xx/429/404 without
the header are unchanged.
Tests: three new rows in the classification matrix (403 deny, 429
challenge -> non-terminal; 401 + header -> terminal), a Retry-After
forwarding test, and a runtime-resolver test asserting the on-disk
access/refresh tokens survive a 403/429 checkpoint. All five new cases
fail on the previous commit.
* fix(auth): keep device-code polling alive through edge/WAF non-JSON errors
The generic RFC 8628 device-token poll loop (shared by the Nous Portal and
xAI flows) aborted the whole login when the token endpoint returned a
non-JSON error body. Vercel fronts the Nous Portal and answers rate-limited
clients with a text/plain 403 (x-vercel-mitigated: deny) or 429 — no JSON
body, so such a response can never carry authorization_pending/slow_down.
One mitigation response mid-approval killed a device login the user may
still be approving in the browser.
Treat non-JSON 403/408/429/5xx as transient: back off (doubling from the
current interval, floor 5s, cap 60s; Retry-After honored when present) and
keep polling until the device code expires. Statuses outside that set keep
the existing abort behavior and JSON OAuth errors keep each caller's exact
error contract.
Co-Authored-By: Claude Code <noreply@anthropic.com>
(cherry picked from commit 13ff660c5ebcbade203faad5e63c12de0603af5f)
* chore(contributors): map zzragida@gmail.com -> zzragida
* fix(auth): bound device-poll edge backoff and gate 403 on x-vercel-mitigated
The edge/WAF backoff in the shared Nous+xAI device-code poll loop had three
gaps found in review:
- Retry-After was parsed with a bare int() and never capped, so a
`Retry-After: 3600` slept an hour past a 5-15 minute device code. Parse it
with the shared agent.retry_utils.parse_retry_after_seconds and bound every
sleep by min(60, time left before the device-code deadline).
- The backoff was written into current_interval, so after a block normal
authorization_pending polls kept the inflated interval and slow_down grew
from it. Keep it in its own edge_backoff, reset on any OAuth JSON response.
- Any non-JSON 403 was treated as transient, disagreeing with the refresh
classifier from the previous commit. Only a 403 carrying
x-vercel-mitigated is the edge speaking; a header-less non-JSON 403 raises
as before. 408/429/5xx stay transient.
Tests reduced to the two invariants: recovery after edge blocks (each sleep
<= 60, back to the server interval afterwards) and a persistent block ends at
the deadline without oversleeping.
* fix(auth): stop promising an automatic retry on edge-blocked Nous refresh
The upstream_blocked message said "Hermes will retry", but nothing schedules
a retry of the refresh; the next request simply tries again. Reword it to
what actually happens (credentials kept, try again shortly).
Fold the separate Retry-After test into the existing parametrized refresh
classification row (header + one assertion) and drop the assertions that
pinned exact error wording.
* test(auth): fold the edge-block refresh case into the 503 matrix and pin the deadline clamp
* chore: map 686f6c61 for salvage of #120235
* fix(state): detect Windows database holders before maintenance
(cherry picked from commit b006ae2dcf6b210d3db3dfbe78c5aac040044644)
* test(state): identify real Windows SQLite holder process
(cherry picked from commit 7524006919906df3b0303b38c95a9b37ad5f27a2)
* test(state): cover Windows sidecars and force fence
(cherry picked from commit 1fa30c2a8d2331393ac9f81abd181f5cb42d0e97)
* test(state): use real Windows SQLite holder pid
(cherry picked from commit 1b1244a653e3a3abd8273a86b06fa58de253a322)
* test(state): drive the Restart Manager scan through an injected rstrtmgr
The two Windows holder tests faked the `_IS_WINDOWS` module constant to
reach the rstrtmgr lane, which the hermes_platform convention forbids
(no platform faking; gate with windows_only or inject the helper). They
also only checked routing, never the ctypes control flow.
Replace both with one test that calls `_windows_restart_manager_holders`
directly against an injected `ctypes.WinDLL`: db + -wal registered (-shm
absent), ERROR_MORE_DATA sizing followed by the data call, our own pid
excluded, session ended, and an RmStartSession error surfacing as OSError
so `foreign_state_db_holders` can fail closed. The real-Windows witness
stays in test_state_db_holders_windows_live.py (windows_only).
Fixes the false all-clear that let optimize-storage/prune run against a
live WAL holder (#120205), consumed by held_store_refusal,
live_writer_holds_db, doctor and the FTS WAL guard.
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
* fix(state): hoist Restart Manager ctypes metadata to module level
`_windows_restart_manager_holders` rebuilt two ctypes Structure classes and
called `ctypes.POINTER()` on the fresh class every scan. POINTER memoises in
the process-global `ctypes._pointer_type_cache`, so each call pinned another
class set forever: probe measured 3 -> 221 cache entries over 200 calls. The
scan runs from auto-VACUUM admission and the FTS-rebuild guard in a long-lived
gateway, so the growth was monotonic. Structures, constants and argtypes now
live at module level (ctypes.wintypes imports on every platform); the
`ctypes.WinDLL("rstrtmgr")` load stays per call so the unit test can still
inject a fake rstrtmgr. Same probe after: 22 -> 22.
While here: one `_sqlite_family(base)` helper replaces the fourth hand-rolled
(db, -wal, -shm) triple in this module and the three existing ones; the
Windows lane computes abspath once instead of per holder; the RmGetList
sizing call and the retry loop collapse into one bounded loop (pass 1 sizes
with a NULL buffer, later passes re-size), which keeps the existing fake's
`apps is None` contract.
* fix(sessions): let set-journal-mode --force waive only a failed holder scan
Retiring the win32 gate left `--force` with no reader, so on a Windows host
where the Restart Manager cannot start a session (restricted/service context)
the fail-closed `(-1, scan failed)` sentinel made `set-journal-mode`
permanently unrunnable, while optimize/optimize-storage/prune kept a working
override. `--force` now drops only pid <= 0 sentinel entries — a process the
scan actually found is still refused — and the help text says exactly that.
The holder test parametrised over `force` now asserts something real: force
plus a live holder is refused, force plus a failed scan proceeds (and without
force the failed scan is refused). Rewrite the user-guide paragraph that still
described a POSIX-only scan and a no-scan Windows path. Drop the dead
`import time`. The doctor holder test compares the child-reported pid instead
of `Popen.pid`, which is the venv launcher on Windows, so un-skipping it there
does not assert a pid equality that cannot hold.
* perf(state): build Restart Manager ctypes metadata lazily, bind argtypes by name
Importing hermes_state loaded ctypes and built the Restart Manager structures
on every POSIX process that never scans with Restart Manager; build them once
on first scan instead. Bind argtypes by function name rather than zipping dict
order against a positional tuple.
* ci: drop the venv-e2e workflow edit; tests-os.yml already runs the live test
The workflow edit trips the review-label gate. The new test is marked
windows_only, so tests-os.yml's '-m windows_only' job on windows-latest
already collects it.
* test(e2e): #120205 is fixed; drop the strict known-failure marker on the Windows state-db guard
CI's Windows E2E job XPASSed (strict) on this branch: the write guard now
refuses while a live gateway holds state.db.
* test: provider-catalog OAuth E2E (device-code cadence, refresh rotation) + multi-dialect loopback fake
* test: provider-catalog E2E matrix — credential routing, base URL, dialect turn, usage/cost, listing, /model switch, fallback over every discovered provider
Rows come from real plugin discovery (providers.list_providers in a child), sharded by
name hash into 3 files. Strict KNOWN entries: #121347 (xai base_url), #121359 (fallback
base_url for anthropic/openrouter), #121387 (picker listing ignores base_url), #121388
(nebius switch validation). OAuth cells switch to a strict message-gated helper.
* test: provider-catalog E2E: merge-order-safe KNOWN gates, exact-path fake, bounded vendor stalls
Review fixes on the provider-catalog matrix:
- KNOWN bugs go through tests.e2e.core._pending_fixes.known_failure again (strict_known and the
"now green -> drop KNOWN" asserts are gone), so a fix PR landing first simply turns its cells
green. Each KNOWN is gated on the bug's OWN observed signature via a dedicated CatalogGap raised
only at the gated assertion: xai = no inference at the fake + api.x.ai CONNECT; listing = vendor
host hit with no probe error; nebius switch = vendor-host validation error; fallback = fallback
fake untouched + its vendor host CONNECTed; OAuth patterns anchored. Unlisted red cells still fail.
- CatalogFake answers only exact routes (configured base path + dialect endpoint / listing path),
404 otherwise; reached_own_endpoint and the listing cells assert the exact path.
- Turns run with agent.auto_recovery_cycles: 0 (the documented post-exhaustion ladder parked the
xai/fallback rows for minutes: 14 CONNECTs over 118 s, bounded by design, not a retry bug) and a
watchdog kills a child 8 s after a vendor-host CONNECT with no inference at its fake.
- An unknown api_mode fails the row instead of skipping; only explicit auth types and the named
keyless provider skip.
- Usage cell asserts the exact sum the fake reported for answered main-turn calls.
- Listing 404 / hang degradation cells for the rows that already list from the configured endpoint.
- Shard body moved into the helper; fallback key check uses the dialect's auth header; the three
unconditional-skip OAuth params dropped (kept under NOT COVERED).
* test: provider-catalog listing: accept both exact listing paths under the Anthropic override; relay-id rendering only on control rows
Merge-order check with #121398 (picker honours base_url) showed the stricter cell turned red on a
correct fix: minimax lists at <base>/models under the /anthropic override, and deepinfra /
commandcode-anthropic post-filter the relay's generic ids. The KNOWN-gated cell now asserts only
the configured endpoint at an exact listing path and no vendor host; rendering the relay's own ids
is a separate cell on the rows that already list from the configured endpoint. Also encoding= on
the OAuth auth.json read/write (windows footgun).
* test: provider-catalog E2E: drop the wall-clock vendor-CONNECT watchdog (killed a healthy deepinfra turn on CI)
CI (run 35996862679) failed test_catalog_matrix_0.py::test_provider_row[deepinfra] with rc -9
"killed 8.0s after CONNECT api.deepinfra.com:443 with no inference at the fake"; it passed locally.
Root cause is the harness, not the product and not a cache leak:
- The child homes are identical locally and on CI: no models_dev_cache.json anywhere, models.dev
CONNECT refused by the sentinel in both places (checked the probe home + egress timeline).
- deepinfra CONNECTs its own vendor host once during startup (catalog fetch, refused at once and
neg-cached), then does ~2 s of CPU-bound init (imports, config load, scratch prune over
psutil.process_iter, system prompt, title write) before its first POST to the fake. On the
64-way CI runner that gap stretched past the 8 s grace and the watchdog killed a turn that
would have succeeded. Scaled repro: grace 1.5 s locally reproduces the exact CI red (same six
cells, same five local-probe GETs, rc -9).
- The watchdog was added to bound the auto-recovery ladder, which write_home already disables
(agent.auto_recovery_cycles: 0): xai now ends on its own (7 refused CONNECTs, rc 2, ~18 s) and
its KNOWN signature (fake_inference=0 + api.x.ai CONNECT) is unchanged.
run_hermes keeps only the TURN_TIMEOUT hard bound (proc.wait(timeout=...)); the fallback file
drops the same watchdog arguments (deepinfra as a fallback had the identical latent race).
Unused imports in _catalog_helpers.py removed.
Catalog suite (6 files): 3x serial + 2 parallel copies, XDG caches pointed at an empty scratch
dir: 125 passed / 9 skipped / 0 failed every run, xfails unchanged (matrix_0 1, oauth 2,
listing 37, fallback 2).
* fix(desktop): honour arrowed slash-picker highlight on bare '/' query
implicitSlashAcceptIndex() returned null immediately when no command name
had been typed yet (!typed), so the activeExplicit branch was never
reached. Pressing Enter after arrow-key selection on a bare '/' sent a
bare '/' instead of the highlighted command (#98535).
Fix: move the activeExplicit guard above the !typed early-return. A
deliberately arrowed highlight always wins regardless of query content —
'Enter means I want this one' should hold even before any characters are
typed.
Regression tests added:
- bare '/' + arrow-key → returns the highlighted index
- bare '/' + no arrow-key → still returns null (auto-accept not engaged)
* test(desktop): cover the live empty slash query
detectTrigger('/') passes query '', not '/'. An arrowed highlight still
accepts that row, and an unhighlighted bare slash still returns null.
* fix(desktop): announce a deferred /compress completion
A manual /compress whose session.compress reply is `status: 'pending'`
(#97948 — the gateway's compute-host wait expired while the host kept
compressing) shows one 8s "still running in the background" toast and
returns. The summary and the success toast are rendered only in the
synchronous result branch that `return` skips, so when the host finishes
the `compacted`/`ready` edge merely clears the spinner: no transcript
line, no completion toast, and hydrateFromStoredSession swaps the
transcript underneath the user with no acknowledgement it ever ran.
Claim the session on the pending reply and announce the completion on the
terminal edge, reusing the notice id the handler already notified under so
the pending toast is replaced in place rather than stacked. The transcript
line is appended only after the hydrate resolves — it replaces the
transcript wholesale and would otherwise drop the line. The claim is
consumed once, so a later auto-compaction cannot replay the notice and an
auto-compaction the user never asked for stays silent.
* fix(desktop): refocus find bar input on repeated Cmd-F
A second open while the bar is already visible rewrote EMPTY, which
cleared the query, and the focus effect only watched active, so it
never ran again. Bump focusRequest instead, keep the typed query and
the captured scope, and select the existing text on refocus.
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
* test(desktop): packaged-app smoke — asarUnpack contract + packaged binary boots to a first chat (#121097)
* test(desktop): remote-backend topology — image bytes not client paths (#120730), rename across backend restart (#121192)
* test(desktop): lineage/sidebar integrity after compaction and branching (#121148 #121088 #121062 #121096)
* test(desktop): format + drop empty KNOWN scaffolding in lineage spec
* test(desktop): real compaction lineages — rotated rows (#121148) and in-turn compaction handoff (#121088) as merge-order-safe KNOWN cells
lineage-sidebar's #121148/#121088 steps never compacted (/compress was
refused as would_grow on a tiny transcript), so they could not fail.
lineage-rotation drives real auto-compaction with compression.in_place:
false (state.db-verified sealed parent + continuation) and asserts one
sidebar row per lineage after rotation, reload and cold relaunch; the live
nested-branch shape is KNOWN #121148. lineage-compaction-prompt compacts
inside an acknowledged redirect turn so the refresh passes through
preserveLocalPendingTurnMessages; the duplicate prompt is KNOWN #121088.
known.ts: expectNoSymptom marks a test expected-failing at run time only on
the bug's own assertion (a fix merging first stays green). lineage-sidebar
keeps the branch/switch scenario, drops the #121096 claim and now checks the
in-page duplicate sampler.
* test(desktop): #120730 Bot-Mode shape (local primary + remote secondary) and a backend that really cannot see client paths
remote-secondary: a bot on a remote secondary connection (connections.json)
with the local backend as primary; a client-only image must reach the remote
as image.attach_bytes. startRemoteBackend({hide}) covers client folders with
an empty tmpfs in a private mount namespace (unprivileged userns; annotated
'fidelity' when unavailable), used by both remote specs. remote-topology:
the post-restart state.db poll (which passed before the new backend opened
the DB) is replaced by a new-pid check + the restarted backend serving the
title after reload.
* test(desktop): packaged boot fails fast (60s) with main log tail; README lists the new specs and KNOWN policy
* fix(desktop): take keyboard focus when the terminal pane is revealed
Revealing the terminal restored the layout but left the composer holding
the keyboard. Terminals stay mounted while hidden, so the activation
focus effect never re-runs. Ctrl+backtick, the palette row, and the
statusbar pill now claim focus on the reveal direction and re-assert if
the composer steals it. Hiding does not. No second focus chord.
* fix(desktop): ship a Windows voice chord that is not the sidebar toggle
Off macOS, ctrl folds to mod, so a Ctrl+B voice default is the sidebar
chord. Ship Ctrl+Alt+V instead, and let a binding be cleared to an empty
combo so sidebar mod+b can be unbound. Point the Voice settings hint at
the voice conversation action, not dictation.
* test: resize E2E submits each question only after its echo, so a starved CLI cannot eat the Enter
test_resizes_keep_each_transcript_line_once_in_tmux_scrollback typed the
question, slept a fixed 0.5 s, then sent Enter. The classic CLI treats an
Enter processed within 50 ms of the last buffer change as a pasted newline
(_RAPID_INPUT_ENTER_WINDOW_S, #10994), and that clock is processing time,
not arrival time. On a starved runner the CLI can still be chewing on the
typed batch when the Enter lands in the PTY; it then processes the Enter
right after the last character, inserts a newline, and the question is
never submitted. CI (64-worker e2e run) failed exactly so: the composer
held "question zq1q please" plus a blank row and 't1w059' never appeared.
Repro: a scratch copy of hermes_cli/cli_tui_mixin.py that blocks the event
loop for 1.0 s when the composer's first key is processed. Base fails 4/4
with the CI signature; the fix passes 12/12 (stall hit on all 3 turns).
Fix: wait for the typed question to render (every key processed) before
the 0.5 s gap and the Enter, the same handshake tests/e2e/core/terminal/
_pty.py::submit already uses.
* test: resize E2E dumps provider request count, agent/errors log tails and a CLI thread dump on a missed wait
* test: classic-CLI E2Es type only into a live composer and leave the Enter real time
Root cause of the CI-only reds of test_resizes_keep_each_transcript_line_once_in_tmux_scrollback
("'t1w059' never appeared", provider saw 0 requests): the test typed its first question two
seconds after "Welcome to Hermes", but that line is printed before prompt_toolkit's app starts.
On a starved runner the app was not yet up, so the keys landed in the still-cooked tty: the
kernel echoed them (the plain "question zq1q please" row above the status bar), the echo
satisfied wait_for(question), and the later Enter completed the line in the kernel buffer. When
the app took raw mode it read text + Enter in one batch, and the rapid-input guard (#10994)
kept the Enter as a pasted newline: the question sat unsent in a two-line draft, "ctx --".
- Wait for the app-painted status bar before typing, and for the question on the "❯" composer
row (the app's own render, never the tty echo) before the Enter.
- The Windows ConPTY composer test wrote "\r" the instant the echo painted, inside the same
50 ms rapid-input window: give the Enter 0.5 s after the last typed key.
- Keep remain-on-exit so the SIGABRT faulthandler dump of the diagnostics stays readable.
Reproduced locally by delaying CLI startup 3 s after the welcome line: the pre-fix harness
fails with the exact CI screen and main=0; the fixed harness passes. No network is on the
turn path: every outbound call (update check -> api.github.com, tirith install -> GitHub
releases) runs on a daemon thread; black-holing all non-loopback DNS leaves the first turn
at 1.4 s (tmux) / 2.7 s (chat -q).
* test: make E2E known-bug pins merge-order safe
* test: anchor four KNOWN patterns on the bug itself, not its symptom class
Review of #121499: a slow-but-healthy dashboard install, partial pricing,
a nested permissions schema error, and a refusal that omits the holder pid
no longer XFAIL as the pinned bug.
* test: drop the now-empty strict KNOWN mechanism from the userstate import E2E
#119953 and #121258 landed on main (9966de47602) and emptied its KNOWN table;
the strict marker helper left behind is dead.
* fix(desktop): re-auth a lapsed saved Cloud gateway from Settings
A saved Hermes Cloud connection on a local-primary device had no
sign-in surface when its gateway session lapsed. The dial rejects
with a reauth-shaped error whose copy says 'Open Settings → Gateway
and sign in again', but the Settings OAuth row only renders in
remote mode, and both 'Use gateway' paths called selectConnection
only — the silent portal cascade ran solely on first connect, team
change, or boot-primary failure (overlay). Portal liveness also
read 'Signed in' while the agent session was dead.
One recovery ladder now exists in lib/cloud-agent-session.ts
(reestablishCloudAgentSession: drop lapsed cookies → ensure portal
session → silent per-agent cascade) and is shared by Settings
(retry the switch once after re-establishing) and the boot overlay
(same sequence, deduplicated, so the two cannot drift).
Tests: three invariants in gateway-settings.test.tsx — cascade +
retry on reauth rejection, no cascade for non-reauth failures,
interactive portal login when that session lapsed too. Proven red
on base for the two cascade cases.
* fix(desktop): reject missing Cloud dashboard URL
* fix(desktop): let the cloud ladder own the reauth logout
The boot overlay dropped this gateway's cookies before branching, and
reestablishCloudAgentSession drops them again as its first rung, so the
cloud path fired oauthLogoutConnectionConfig twice per sign-in. Move the
logout into the native-OAuth branch that still needs it; the overlay's
cloud test now pins exactly one call.
* fix(desktop): coalesce zone-editor split preview to one frame
Repeated pointer events were writing the split preview on every move, and
Shift did not flip the line until the next move. Paint at most once per
frame, skip an unchanged preview, and recompute orientation on Shift while
the pointer is still inside. The grid is not written until the click commits.
* fix(desktop): load local skin before gateway connects
* fix(agent): surface why every init-time fallback entry refused
(cherry picked from commit e0113f673c1d2eefb2a217974893a1d801d3034f)
* style(test): make the fallback-refusal warning suite formatter-stable
(cherry picked from commit 1c82a53ba958365303edb3a4b967575d7b3ae8ad)
* fix(agent): neutral wording for the init-fallback refusal summary
The explicit-provider branch raises missing_provider_credentials_message,
not the generic 'No LLM provider configured' error, so the WARNING that
summarises refused fallback entries must not assert that verdict. Pin it
with a test on the explicit-provider path.
Refs #119533
* refactor(agent): bind fallback provider once; name bare exceptions
The init-time fallback loop re-derived str(_fb.get("provider")) four
times per refused entry even though _fallback_entries guarantees the
key; bind it once per iteration and reuse it for logs, the refusal
list and the moa check.
A bare exception (e.g. KeyError()) stringifies to "", so the refusal
WARNING rendered "kimi ()" with no reason at all; fall back to the
exception type name.
* test(agent): fold fallback-refusal suite into two invariants
The explicit-provider test duplicated the refused-ladder test except
for the primary slug, and its startswith() check pinned old log
wording rather than behaviour. Parametrize the refused-ladder test over
an openrouter and an anthropic primary (each asserting its own raise)
and keep the recovered-ladder negative control.
Drop the per-test HERMES_HOME setenv and tmp_path params: conftest
already isolates HERMES_HOME per test and the path never reads it.
* fix(agent): name an exhausted primary pool in the init refusal warning
Issue #119533 is the misleading "No LLM provider configured" when the
primary's credential pool is burned. The refusal WARNING only fired when
fallback entries existed and never said why the primary failed, so the
reporter's trigger with no ladder configured still left no trace.
resolve_provider_client returns a bare (None, None), but the pool state
is a cheap local read: ask load_pool(primary) whether it has credentials
yet none available, include "credential pool exhausted" in the WARNING,
and emit it even with no fallback entries. A pool-less, ladder-less
primary (genuine first-run) stays silent to avoid duplicating the raise.
* test(agent): make the fake pool's has_available keyword-only like the real one
A positional call would pass the fake but raise TypeError against the real
CredentialPool, which the diagnostic's suppress() would then swallow.
* fix(plugins): route inject_message to the TUI session_key queue
Ink TUI and desktop never registered an inject host, and sharing
set_gateway_message_injector with a live messaging gateway would let
the last writer win. A separate host queues the reported session_key
onto that session's prompt queue and leaves other keys for the gateway.
Fixes #87412
* fix(gateway): a busy follow-up queued as the turn ends is no longer left without an owner
The base adapter sends a message to the runner's busy handler while the
previous turn still holds the session guard. The handler awaits before it
decides: the compression-lock read in the stock interrupt mode, or the
profile secret-scope load when profiles are multiplexed. If the previous
turn reaches _finish_session_task meanwhile, it finds the slot empty,
releases the guard and exits. The follow-up is then queued with no task
to drain it, either by the runner FIFO or by the base adapter when the
handler returns False. It waits until the next inbound message, whose new
turn runs BEFORE it (M1, M3, M2). If nothing else arrives, it never runs.
_handle_message_while_active now checks the guard after the handler
returns. If the guard is gone, it starts what the handler queued. If the
handler left the event to the base path and nothing is queued, it starts
the event itself.
A handler that raises after it stored the event (for example when
composing the busy ack fails) now counts as having handled it, based on
the event's _gateway_accepted receipt. Otherwise the new path would start
the stored event and the base path would queue it again, so it would run
twice. This also stops the older no-race case from merging the same text
into the slot twice ("M2\nM2").
tests/gateway/test_busy_followup_guard_release.py covers the stock
interrupt and multiplexed queue-text routes, and the case where the
handler raises after it stored the follow-up.
Known behavior, unchanged or out of scope:
- During a restart drain with busy_input_mode queue/steer, a follow-up
started this way hits the drain gate ("not accepting new work")
instead of waiting in the slot for the shutdown flush. The non-race
path does the same (#82381).
- A message can still run before an earlier follow-up if it arrives
after the guard is released but before that follow-up's handler
returns (for example during its busy-ack send). Nothing is lost.
- The interrupt decision still uses the running agent read before the
handler awaits.
- The "Interrupting current task" ack can arrive after the turn has
already ended.
(cherry picked from commit bb9f15c7fa29cdffc56790ea3118d8baac2195f1)
* chore: map delltrak's commit email to their GitHub login
* test(gateway): pin debounce/direct-merge follow-ups started when the session releases mid-route (#121393)
Queue-mode (debounce) and direct-merge follow-ups whose busy handler
returns False after the running turn released the guard must start a
fresh turn; a live guard keeps the queue-behind behavior.
Test file taken verbatim from 33801e8fb7 (PR #121464). Its base.py hunk is
dropped: the same guard re-check is already provided (more broadly) by the
preceding commit from PR #121003, and the two hunks conflict.
(test file cherry picked from commit 33801e8fb7d0c1553511e470f410b636ed9426b8)
* style(tests): drop trailing blank line at EOF in busy follow-up release test
* refactor(gateway): read the declared _gateway_accepted field directly in busy routing
_gateway_accepted is a declared MessageEvent field (event.py), so the
defensive getattr is unnecessary; wake.py already reads it directly. Drop
the dead 'handled = False' initializer: both try/except branches assign
it and a CancelledError propagates before it is read.
* test(gateway): reuse text-merge helpers and parametrize the mid-route release test
Drop the local telegram stub (conftest._ensure_telegram_mock installs it)
and the copied _make_event/_DummyAdapter/_make_adapter in favour of the
helpers in test_active_session_text_merge.py. Merge the debounce and
direct-merge cases into one test parametrized over busy_text_mode.
* fix(streaming): back off between stream reconnects; treat Anthropic connection errors as transient
Stream-level reconnects (HERMES_STREAM_RETRIES) fired back-to-back with no
delay after a drop, hammering a provider that had just dropped us. Wait
1s * 2**attempt (capped at 4s) in _retry_after_drop, polling the interrupt
flag every 0.1s so /stop still exits immediately. Both retry continuations
(mid-tool-call and pre-delivery) go through _retry_after_drop.
anthropic.APIConnectionError (the SDK's wrapper for connect/read drops,
incl. stale-kill aborts) was not in the transient tuple, so Anthropic-path
drops failed the turn instead of reconnecting.
Refs #60029
Co-authored-by: isheng <ishengeqi@163.com>
* refactor(streaming): reuse jittered_backoff for stream reconnect delay; no SDK import in error path
Review follow-up (#60029): restart the stale clock after the backoff so the
stale monitor cannot strike a stream that has not reopened yet. the private 1/2/4 s schedule duplicated
agent.retry_utils.jittered_backoff and escaped the tests/agent autouse stub
that zeroes it, so unrelated retry tests slept for real. The Anthropic
connection-error type is now read from sys.modules (an Anthropic error implies
the SDK is loaded) instead of importing it inside the handler. Tests record the
requested delay instead of patching the global time module.
* fix(desktop): silence live-region elapsed timers
Signed-off-by: Kevin Yin <182213728+yinkev@users.noreply.github.com>
(cherry picked from commit 27bc6e46a54b69d74baac8a0ec82ffb6cf69d902)
* fix(desktop): keep turn-activity live region mounted across idle gaps (#46225)
TurnActivityIndicator returned null whenever the turn went quiet, so each
working/idle flip remounted a role=status aria-live node and screen readers
re-announced it. Latch once shown; while idle keep the row as an empty,
sr-only live region (still in the a11y tree) instead of unmounting, and only
mount the pulse/timer while active.
Co-authored-by: LINGJIAKUN <200388706+LINGJIAKUN@users.noreply.github.com>
* chore: map contributor emails for streaming salvage (providers-misc)
* fix(kimi): exclude brotli from Accept-Encoding to work around httpx/brotlicffi streaming decode bug (#59556)
httpx's brotlicffi backend (pinned for Discord attachment decoding) fails
on Kimi API's content-encoding: br SSE streaming responses with:
brotli: decoder process called with data when can_accept_more_data() is False
This surfaces as 'API call failed after 3 retries: Connection error'
on Windows Desktop where the Electron-packaged Python venv includes
brotlicffi (#59556). DeepSeek and Xiaomi are unaffected — only
moonshot endpoints (api.moonshot.ai, api.moonshot.cn) trigger the bug.
The fix forces Accept-Encoding: gzip so the Kimi API falls back to gzip
compression, which httpx handles reliably on all platforms. This is the
same workaround already applied in tools/skills_hub.py for sitemap
fetches and the Hermes index download.
Closes associated issues: #28043, #48428
(cherry picked from commit 9c130133b0676138858b73847d1d7c7549339484)
* test(kimi): pin gzip-only Accept-Encoding on both Moonshot profiles (#28043)
* fix(lsp): raise LSP subprocess StreamReader limit to 16 MiB
(cherry picked from commit 0768510a14693cc5cf521373e330858a5aba576f)
* refactor(lsp): inline the one-caller subprocess stream-limit helper into LSPClient (#31417)
Co-authored-by: leavedrop <1433810735@qq.com>
* fix(agent): record the serving downstream provider in stream-drop diagnostics
Relay routing is re-rolled per request and the winning downstream is reported only
inside the delta chunk bodies, so the header snapshot in agent/stream_diag.py could
not attribute a mid-stream drop to a provider. The per-attempt diag now carries
serving_provider (first non-empty chunk-body provider), log_stream_retry prints it,
and the post_api_request payload exposes it as upstream_provider for plugins auditing
route compliance. Fixes #90216.
(cherry picked from commit f0cf4fe4f4b47e384532008b1bf56ad259c87dc2)
* chore(stream_diag): log the swallowed provider-note failure at DEBUG
Review follow-up on #118759: the bare `except Exception: pass` in
stream_diag_note_serving_provider is intentional (best-effort annotation that
must never break streaming) but indistinguishable from a missed error-handling
gap. Keep the swallow, add a DEBUG log with exc_info so the intent is legible.
(cherry picked from commit 6674471642c4bdb81ed212cd1cee0242647b60ed)
* fix(stream_diag): feed the existing chunk-body upstream_provider into diag + hook payload (#90216)
Drop the duplicate chunk-reading helper, run_agent facade forward and
_last_serving_provider agent state; the chat-completions loop already captures
chunk.provider, so stamp it on the per-attempt diag there and read the hook's
upstream_provider from the assembled response.provider.
* test(stream_diag): trim serving-provider tests to the log-line and hook-payload contracts (#90216)
* test(lsp): scope-drop the module-wide guard bypass and derive the oversize stderr line from _STREAM_LIMIT (#31417)
The new stderr tests only send extra stderr; they need no real signals, so the
module-wide live_system_guard_bypass mark is removed. The mock server now takes
the line size from the test (_STREAM_LIMIT + 1) so the overrun path follows the
limit. Also restores the blank lines before test_shutdown_never_signals_*.
* chore: map contributor emails for streaming salvage (reasoning-strip)
* fix(agent): strip stray arg tags to line boundary, not end of text
(cherry picked from commit 443db6b1b7076fcb08fb6ddacf11baa902ac544c)
* test(agent): cover arg-tag line-boundary edge cases
Co-authored-by: crazyief <crazyief@users.noreply.github.com>
(cherry picked from commit 7e970b6d8f02414d33a4741a54e860fe9d61df72)
* test(agent): trim arg-tag line-boundary tests to invariants (#102303)
Keep the parametrized cut-fragment contract (prose after a fragment
survives) and the bracketed-prose contract across both strippers; drop
the per-whitespace-variant EndOfLine class.
Co-authored-by: soroush5 <mrsoroushahmadi@gmail.com>
* fix(agent): flatten list-shaped thinking block payloads in extract_reasoning
Non-strict OpenAI-compatible backends (Mistral via custom provider) can
deliver a typed thinking block's text as a JSON array instead of a
string; extract_reasoning called .strip() on the raw value and crashed
the whole API call with AttributeError: 'list' object has no attribute
'strip' (#106006). Flatten the payload first, same as the
reasoning/reasoning_content fields above it.
Fixes #106006
(cherry picked from commit cadec563988ee6a84ea75f063c7f35f8a25713e5)
* test(agent): trim extract_reasoning thinking-block tests to invariants (#106006)
* refactor(cli): reuse storage tool-call strip patterns in _strip_reasoning_tags
The display stripper kept byte-identical inline copies of
_STRAY_TOOL_CALL_CLOSER_PATTERN and _UNTERMINATED_TOOL_CALL_PATTERN
(same flags), which had to be edited in lockstep (#102303). Import the
compiled patterns lazily instead.
* fix(agent): strip arg fragments glued to a bare tool name (#101899)
The line-boundary arg-tag rule (#102303) required the tag at line
start, so a cut call with no <tool_call> opener but the tool name on
the same line (process_manage<arg_key>action</arg_key>...) leaked raw
markup. Allow an identifier-only prefix before the tag; the prefix has
no spaces, so prose like 'The <arg_key> holds...' still survives.
* fix(agent): dedupe repeated stream continuation tails
(cherry picked from commit 87c93960cd9e8dd502ffe82a10febdefa3cf64c4)
* fix(agent): scope continuation dedupe to stream stubs
(cherry picked from commit 82acba16e08cce9f613598325d7c566838fb002e)
* fix(gateway): mark shutdown/restart notifications as interim sends
Grafted from #98445 onto gateway/run_shutdown.py (the shutdown notifier moved
out of gateway/run.py). Both the active-chat and home-channel sends now carry
_interim_metadata so a stream-is-the-message adapter never seals an in-flight
answer with the shutdown advisory.
Fixes #98432
* test(gateway): update restart-notification metadata assertions for interim marker
The home-channel/thread assertions in
test_restart_home_channel_notification_not_deduped_across_threads
expected raw thread metadata and None; shutdown sends now carry
_interim_send per #98432. Align both assertions.
(cherry picked from commit c4ec949ae77ae5611684061a4afb8d96af335e26)
* test(agent): trim continuation-assembly tests to 2 invariants; fix mock import path
The picked intentional-repetition loop test imported tests.run_agent.test_run_agent,
which does not exist on main (helpers live in tests/agent/test_run_agent.py).
* test(gateway): shutdown notices carry the interim marker (#98432)
Ported from 12845c7b1a (#98445): home-channel + active-chat sends assert
_interim_send, and the cached-thread-source assertion expects the marker.
* fix(gateway): mark cron interrupt notice as an interim send
Same send shape as the shutdown notices grafted from #98445: it fires during
drain while turns may still stream, so it must not seal a live answer (#98432).
* refactor(gateway): apply shutdown-notice interim marker once in _send_notice_logged
Also type continuation parts as (text, is_stub) tuples end-to-end and bound
the overlap scan to the tail of the joined text.
* fix(agent): unmask opaque streaming 5xx with one non-streaming retry
Some OpenAI-compatible gateways validate requests only on their
non-streaming path; the same request streamed returns a generic
"500 something went wrong" (observed on AssemblyAI's LLM Gateway with
tools + reasoning_effort). The user sees three identical opaque 500s
and no actionable message.
On a pre-delta 5xx, re-issue the request once non-streaming:
- probe succeeds -> deliver it and latch non-streaming for the session
(the existing _adopt_final_response path)
- probe 4xx -> surface the provider's real validation error instead of
the opaque 5xx
- any other probe failure -> keep the original error
The probe never runs after partial delivery, for non-5xx errors, or
when the failure carries no HTTP status.
(cherry picked from commit 37e683913513fbf7b714ce7c0625d387bfadba83)
* fix(agent): review fixes for stream 5xx unmask probe
Cross-vendor review findings, all validated against source:
- re-raise InterruptedError/KeyboardInterrupt from the probe (was
swallowed into the opaque-500 error path, breaking user interrupts)
- do NOT latch _disable_streaming on probe success: restore the prior
preference so a transient gateway 500 doesn't permanently disable
streaming for the session (one-turn recovery)
- one probe per 60s window latched on the agent (fresh _StreamingCall
per outer retry would re-probe every attempt against a failing backend)
- gate the probe to the chat-completions wire: _adopt_final_response
replays chat-completions shapes only
(cherry picked from commit 3e6de57c7cfc7a1bb4eb834916f7743c5f91b900)
* fix(agent): unmask probe owns its stream lifecycle, never leaks the latch
- The recovered non-streaming delivery opens and closes its own
on_stream_start/end pair. The failed streaming attempt already emitted its
terminal on_stream_end(finished=False), so adoption's deltas were reaching
consumers after a terminal error event.
- Adoption now runs under its own except/finally: a response that cannot be
replayed restores the saved _disable_streaming (instead of leaving streaming
latched off for the session) and keeps the original 5xx, rather than escaping
_handle_stream_error inside _call()'s except block.
- Test sentinels record probe invocation before raising. The production code
swallows exceptions from the probe call (that is where the 4xx/5xx
classification lives), so the negative-invariant tests returned "not handled"
either way and could not fail — mutation-checked, 8 of 11 stayed green with
the guard they cover removed. Every gate now goes red when its guard is
removed.
(cherry picked from commit 8ec812dac68aace6325158fab44151fa482bacbc)
* fix(agent): suppress the unmask probe once /stop is pending
The retry loop checks _interrupt_requested before opening another streaming
attempt, but the probe runs inside _handle_stream_error — outside that check —
so a 5xx that landed after the user pressed /stop still spent one more request
before the interrupt was honoured. Bail out before the probe window is stamped:
the original error propagates and the loop's own interrupt check raises
InterruptedError instead of buying one more call.
(cherry picked from commit a39b7782006c018d276a6a834432f59e66b1e820)
* test(agent): trim stream 5xx unmask tests to 2 invariants (#107948)
* fix(agent): treat jiter SSE parse ValueErrors as transient provider errors (#65147)
Extends PROVIDER_STREAM_PARSE_MARKERS (single home) with jiter's serde
vocabulary and excludes those ValueErrors from _is_local_validation_error so
the turn-level retry/fallback path runs instead of aborting as a local bug.
Grafted from #65154 (its conversation_loop regex is superseded by the markers).
Co-authored-by: Simplicio, Wesley (ext) <wesley.simplicio.ext@siemens-energy.com>
* fix(agent): unmask probe suspends the stream stale watchdog, replays without latching
The non-streaming 5xx unmask probe runs on the stream worker thread while
StreamingWaitMonitor keeps polling last_chunk_time, which nothing refreshes
during the probe. A probe longer than the stale timeout fired
_kill_stale_stream every window: a false "Reconnecting..." status and a
_bump_stale_streak strike toward HERMES_STREAM_STALE_GIVEUP. Suspend the
stale check (timeout = inf) for the probe's duration; the probe keeps its
own non-streaming watchdog.
The probe also went through _adopt_final_response, which latches
_disable_streaming and logs "switching to non-streaming for this session",
then undid the latch in a finally. Split the pure replay into
_replay_final_response and call that instead.
Also: status via error_classifier._extract_status_code, monotonic probe
window seeded in agent_init, docstring says "per 60s window".
* fix(anthropic): require message stop for streams
(cherry picked from commit 8bca0d97c602624f6acaaa3073a8fbb43f718e3c)
* fix(bedrock): reject streams missing messageStop
(cherry picked from commit b35be7ec5f28772f13c2ae70b54d71b67496ab69)
* fix(anthropic): keep live delta streaming with the message_stop gate (#121320)
Drop the contributor's text/thinking delta buffering: it withheld every
token until stream end, killing live streaming. The message_stop gate
alone is enough — a dropped stream raises EmptyStreamError and the retry
loop's existing partial-delivery handling applies. Restore the partial
tool-name test to main's live-streaming contract.
* fix(bedrock): raise EmptyStreamError on missing messageStop; tolerate it in relay finalizer (#109988)
RuntimeError is not classified by the stream retry loop; EmptyStreamError is.
The Relay finalizer replays the same events, so it now returns None instead
of raising and the live consumer surfaces the error.
* test(anthropic): drop buffered-text EOF case; live text keeps partial-delivery contract
* fix(bedrock): keep IAM-denied converse() fallback working with the messageStop gate (#109988)
* fix(anthropic): drop dead Iterable guard; fold eventless-stream message into message_stop gate (#121320)
The real SDK MessageStream is always iterable, so the isinstance(stream,
Iterable) fallback only served get_final_message-only shims, which the
message_stop gate now rejects anyway. Iterate the stream directly; such
shims are unsupported.
In _call_anthropic an eventless stream never reaches get_final_message()
any more (no message_stop -> gate raises first), so the AssertionError
branch was dead. Use saw_stream_event to pick the more specific
'empty stream with no events' message inside the gate instead.
* test(anthropic): end fake streams in message_stop; cover bedrock relay finalizer
Fakes that model a completed Messages stream now emit message_stop so they
pass the #121320 gate. test_text_only_message_without_stop_reason_passes
asserted the pre-gate contract; it now expects EmptyStreamError when a
text-only stream ends without message_stop. Adds one assertion that
_finalize_bedrock_relay_events returns None without messageStop (#109988).
* test(e2e): retire the #121320/#109988 known-bug pins now that the bugs are fixed
tests/e2e/core/_pending_fixes.py: once a fix lands, delete its known_failure entry. With message_stop/messageStop now required, the stream-drop cells pass, so their KNOWN entries go and the assertions stand on their own.
* fix: preserve max_tokens boost on truncated tool-call retries (#72770)
Re-port of f3c6307d22 (PR #72802) formula hunk into agent/turn_truncation.py:
ladder from the requested cap when max_tokens is unset, and cap at 2x that cap
so a request already at >=32768 is not re-sent unchanged. Bundled
skills/email/himalaya files from the PR are dropped.
* test(agent): truncated tool-call retry keeps boost above requested cap (#72770)
* fix(agent): re-assert tool availability on partial-stream continuation
Re-port of 6f01f7476b (PR #72273) onto the _LENGTH_CONTINUATION_NETWORK_STUB
constant that _get_continuation_prompt now returns: state the cut was a
transport interruption, that tools remain available, and drop 'Finish the
answer directly' which read as a text-only instruction.
* fix(agent): tools-available note on dropped-tools continuation; keep legacy stub recognized (#74990)
The dropped-tools partial-stream continuation now also states the cut was a
transport interruption and tools remain available. The pre-rewording
network-stub text stays in the compressor's synthetic-turn set so
crash-persisted nudges from older sessions are not mistaken for user turns.
* fix(agent): refund the API call when an empty response activates fallback (#77305)
Fallback activation after empty-response retries re-entered the outer loop
without refunding the provisional call, so a provider hop consumed an
iteration. Refund via _refund_api_call and carry the count back through
EmptyResponseVerdict -> FinalResponseVerdict -> s.api_call_count.
Salvaged from #88867.
Co-authored-by: RelaxJonh <92573950+RelaxJonh@users.noreply.github.com>
* fix(agent): share the truncation-retry output boost; clamp to model limit (#72770, #79715)
The 2^n output-budget ladder was open-coded in three places. Only the
tool-call retry got the #72770 fix; the length-continuation retry
(ceiling max(32768, cap)) and the Codex-incomplete retry (ceiling
max(32768, base)) still re-sent the same budget once the cap reached
32768.
boosted_output_cap() now serves all three sites: it ladders from the
budget actually in use, its ceiling is max(32768, 2×cap), and it is
clamped to the model's known output limit (_get_anthropic_max_output on
anthropic_messages). When the request already sits at that limit it is
not doubled, so the retry no longer earns a provider 400 (#79715).
The length-continuation fix mirrors #72802.
Co-authored-by: webtecnica <contato@webtecnica.com.br>
* refactor(agent): module-level _refund_api_call import; int verdict count (#77305)
turn_context_compaction imports neither conversation_loop nor
turn_empty_response, so the function-local import is not guarding a
cycle. The helper's docstring now als…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
E2E cells that pin open bugs now xfail at run time only on the bug's own failure message, so a fix PR landing before or after the suite can no longer turn main (or the fix PR) red with an XPASS.
Why
Every KNOWN pin under
tests/e2e/was a collection-timexfail(strict=True, raises=...)(viaknown_marks,known_param,apply_known,xfail_known,known(...)decorators, or a "now green → fail" table). Strict xfail XPASSes the moment the fix merges, which is what broke #121291 when #121264 landed. The repo's convention,tests/e2e/core/_pending_fixes.py::known_failure, is a run-time xfail keyed on the failure message: XFAIL only on that signature, any other failure propagates, a clean pass stays a pass.Changes
_pending_fixes.py: newknown_gate(KNOWN, key, raises=...). It returnsknown_failure(*KNOWN[key], raises=...)when the table names the cell, and a no-op otherwise, so one wrapped block serves gated and plain cells.key -> (pattern, "#issue reason"). The pattern anchors on the bug's own message, never.*. It was checked against the real[observed: …]text (junit XML) and against decoy failures.KnownGap,BoundaryBreach,KnownSymptom,KnownBugSymptom,KnownBugError,Issue120527/Issue120937,Gap,PolledFasterThanAllowed, …), it is passed asraises=and is still raised only at the bug assertion.HarnessErrorand plain harnessAssertionErrors stay outside every gate.known_marks(windows, providers-openai),known,known_param(security),apply_known(mcp_plugins),xfail_known(dashboard). The emptyKNOWN_RED"now green" table (parity), the emptyKNOWN_BROKEN(history) andKNOWN_BUGS(chaos) mark tables, and a dead emptyKNOWN(windows unicode, vertex recovery) are gone too.81 cells converted:
Native provider files (#121345, now merged) are included.
Deleted as already fixed: the delivery
expect_gap(request, 120315, …)probe gate on the 5STREAM_OVERFLOW_GAP_CELLS. #120315 has merged, its probe reportsfixedon main, and those cells pass as plain regression tests. With that,expect_gap(the laststrict=Truexfail helper) had no callers and was removed.Left as is (not strict pins):
delivery/test_cron_virtual_clock_soak.pyGAPS (fix(cron): keep server-local jobs on their wall clock across DST changes #119970, still open) is a run-timepytest.xfailreachable only while a behavioural probe reproduces the defect on the tree under test. Once the fix is in the tree the probe saysfixedand the scenario is a plain test, so it is merge-order safe by construction and not strict.security/test_attach_handshake.py:149active_session_registry_snapshot(..., strict=True)is a product API argument, not an xfail.known_failureusers (live, upgrade config/upgrade path, delivery messaging, copilot-acp crash count gated on Copilot ACP: a crashed CLI is sometimes reported as a timeout and retried past api_max_retries #121467) were already merge-order safe and are unchanged.Proof: three representative cells
Each fix-shaped product edit was temporary and restored byte-exactly;
git diff -- agent hermes_cli toolsis empty.security/test_workspace_escape.py::test_quarantined_auth_store_stays_read_denied[read_quarantined_auth_copy](gated on #121278)[observed: read_quarantined_auth_copy: read_file {'path': '../.hermes/auth.json.corrupt'}: the tool result carries the protected file's content]"auth.json.corrupt"toagent/file_safety.py::_CREDENTIAL_FILE_NAMES→ PASSED (2✓, both quarantine cells)token + "-BROKEN") → FAILED (AssertionError: … no longer holds its canary, raised inside the gate, not swallowed)kanban/test_kanban_decompose_billing.py::test_huge_decompose_reply_is_bounded(gated on #118607)[observed: 500-child reply created 500 child rows (bound 64)]hermes_cli/kanban_decompose.py::_apply_fanoutrefuses > 64 children → PASSEDkanban decomposeCLI timeout 240 s → 0.2 s (child killed) → FAILEDsubprocess.TimeoutExpiredmcp_plugins/test_mcp_streamable_http.py::test_server_crash_mid_call_fails_that_call_and_the_next_turn_reconnects[read-only](gated on #121042)[observed: read-only: next turn did not reach the restarted server: … The connection has been re-established …]tools/mcp_tool_registration.py::_annotation_read_only_hintalso reads mcp 2.xread_only_hint→ PASSEDKnownSymptom: … The connection has not recovered yet(same type, different signature, not swallowed)Further per-lane proofs (all green on fix → PASS, red on an injected fault):
providers/test_oauth_device_flow.py[honors_server_interval](Nous device-code login always polls at 1s, ignoring the server'sinterval: 5→ 429 kills the login #121163): removing the 1 s poll-interval cap made it pass; a login timeout turned it red.providers/test_openai_responses.py[http_200_soft_failure](Codex HTTP-200invalid_encrypted_contentsoft failures bypass replay recovery and trigger fallback #120399): replay recovery before fallback made it pass; a decoy fallback reply with the same exception type and a different message was red.providers/test_native_codex_app_server.py::test_native_compaction_is_not_persisted_as_assistant_text(codex_app_server: native contextCompaction item persisted as a raw-JSON assistant message #121301): dropping thecontextCompactionprojection made it pass; a same-typeKnownSymptomtimeout with a different subject was red.providers/test_native_gemini_schema.py::test_v1_ref_parameter_keeps_its_shape(Gemini adapter drops$ref/$defsinstead of inlining them, producing empty tool parameter schemas #99438): inlining$refbefore the legacy translator made it pass; a wrong-messageKnownSymptomand aTimeoutErrorwere both red.Test runs (
scripts/run_tests.sh --include-integration <dir>, every touched dir)XPASS is now impossible: no
strict=Truexfail remains undertests/e2e/. In the first upgrade dir run,test_update_userstate_import.pyandtest_profile_update_cron.pyerrored because this local worktree had no.venv. The import file is shown re-run above with.venvlinked;test_profile_update_cron.py(untouched by this PR) was not re-run.Windows: the 5 patterns are anchored on the static text of each
expect()message. They were checked in Python against Windows-shaped messages and decoys (e.g. thepkg\sub/repr, a System32bash.exevs a Git Bash path). They have not yet run on a Windows host; this PR's tests-os lane is their first real run.Hygiene:
ruff checkclean,scripts/check_no_tmp_literals.py tests/e2eclean,scripts/check-windows-footguns.py --allclean.