Repository navigation
fix(runtime): recover reset streams and deduplicate terminal model errors - #16
Merged
Merged
Conversation
…re, and recover from resets Reported 2026-09-06: a Telegram turn completed an earlier model call, then exhausted three attempts with `httpx.ReadError: [Errno 104] Connection reset by peer`, and the user saw the same "the configured model endpoint is not running" warning twice. Three defects, each reproduced on this base: 1. Duplicate notice. `agent/turn_recovery.py`'s terminal paths emit the failure envelope through `status_callback` AND return it as `final_response`; `_prepare_gateway_status_message` and `_sanitize_gateway_final_response` then map both onto the same provider-error reply, so chat surfaces post it twice. Chat surfaces now take the provider-failure category from the final response only. Programmatic surfaces (local/api/webhook) keep both raw channels, and the noise/compression/warning filters are untouched. 2. Unsupported diagnosis. One connection reply covered every connection shape. A mid-response reset is not evidence that the endpoint is down — the same turn had just been answered by it. The reply table now separates an interrupted transfer, a refused/unroutable connect (which keeps the endpoint-down wording it was written for, NousResearch#86570), and a cause-flattened `APIConnectionError`, which names both possibilities and asserts neither. Auth / policy / rate-limit classification and redaction are unchanged. 3. Reset recovery was bypassed. `try_recover_primary_transport` — the one place that retires the poisoned httpx pool and rebuilds the client before giving up — is gated on a hand-maintained type list that omitted `ReadError`, the shape 64 of 67 failed attempts actually took. The canonical classifier already listed it, so the gate is now derived from `TRANSPORT_ERROR_TYPES` and cannot drift again. This changes WHICH failures get the single rebuild, not how many attempts anything gets: recovery still fires at most once per API-call block, only after the classifier called the error retryable and the retry budget is spent. No retry-count, model, credential or TLS changes. Tests: `tests/run_agent/test_readerror_reset_recovery.py` injects a real TCP RST from a local listener, so the `httpx.ReadError` driving recovery is produced by httpx rather than constructed, and the rebuilt client goes on to complete a real request. The gateway tests drive the real terminal path and assert the delivery contract (exactly one user-visible bubble) rather than any wording. This improves local recovery and messaging. It does not, and cannot, remove the upstream/intermediary resets themselves; the logs cannot attribute those. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
૮ >ﻌ< ა ci reviewrunning on 497cc47 — fix(gateway,agent): one accurate notice per terminal connect Still running 2 jobs:
|
Lei-k
marked this pull request as ready for review
September 6, 2026 08:24
This was referenced Sep 6, 2026
teknium1
pushed a commit
to NousResearch/hermes-agent
that referenced
this pull request
Sep 20, 2026
Split the single connection row of the gateway's shaped provider-error reply into three: an interrupted established connection (reset/EOF/RemoteProtocolError), a refused/unroutable endpoint (the case #86570 wrote the "not running or is unreachable" wording for), and a cause-free SDK ``APIConnectionError`` that supports neither diagnosis. A reset says nothing about whether the endpoint is up, so telling the user to restart a server that just answered sends them to debug the wrong thing (#116323). Selectively adapted from Lei-k#16 (497cc47) via PR #109701, rebased onto the current reply contract (rate-limit > auth > policy > connection, every reply names a slash command, no operator jargon).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #15.
Candidate:
497cc47556b10eba94c65147cceedeba622aa087Base:
64fc0d989347df727d58f4d29a74c0359978bd4dEvidence
Honest wider-test boundary
Deployment boundary
This PR is against fork main 0.21.0. Production remains on 0.20.6; it has NOT been restarted, upgraded, or patched in place.
A separate minimal 0.20.6 backport was prepared and independently reviewed: 191 focused tests plus 283 and 14 nearby tests passed (2 platform skips). Its isolated candidate image executed successfully with no network or production mounts; the embedded source hashes match the reviewed backport. A minimal authenticated request through that backport also succeeded. None of this claims production deployment or production Telegram end-to-end verification.
Human review/merge policy remains intact. Production rollout requires its own approved safe window and read-back; no automatic merge.