Conversation
Collaborator
|
Thanks for identifying a real configuration gap: current main still reads the generic stream retry budget only from Problems
Suggested changes
Automated hermes-sweeper review. |
Co-authored-by: Orca <help@stably.ai>
eloklam
force-pushed
the
pr/provider-scoped-stream-retries
branch
from
July 31, 2026 06:27
ba653c7 to
bad2d1c
Compare
Author
I updated the existing PR branch in commit bad2d1c:
Verification completed:
The changes are pushed to the existing PR head branch. Please take another look when convenient. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Hermes already supports provider-scoped
request_timeout_secondsandstale_timeout_seconds, but the in-stream retry count (HERMES_STREAM_RETRIES) is only readable from the environment variable. This makes it impossible to configurestream_retries: 0per-provider for pre-first-token failover setups where retrying a stalled primary before switching to the fallback is worse than switching immediately.Changes
hermes_cli/timeouts.py: Addget_provider_stream_retries()— mirrors the existingget_provider_stale_timeout()pattern: readsstream_retriesfrom provider config, with per-model override support.0is a valid value.hermes_cli/config.py: Addstream_retriesto_KNOWN_KEYSso the config validator doesn't warn about "unknown config keys ignored."agent/chat_completion_helpers.py:get_provider_stream_retries()ininterruptible_streaming_api_call()before falling back toenv_int("HERMES_STREAM_RETRIES", 2).stale_timeout_secondshandling: when a provider explicitly configures a stale timeout, use it directly instead of feeding it through the scaling/reasoning-floor logic that was designed for the default 180s baseline. This prevents a 60s operator-configured stale timeout from being silently raised to 240s+ for large contexts, which defeats the purpose of configuring a short timeout for failover.tests/hermes_cli/test_timeouts.py: Add tests forstream_retries=0(model override) and provider-level fallback.Why
Users configuring automatic pre-first-token failover (e.g., primary key + fallback key on the same gateway) need
stream_retries: 0so that a streaming transport failure on the primary surfaces to the outer fallback loop immediately, rather than retrying the same broken primary. Without this, the in-stream retry loop consumes time on a stalled connection before the fallback chain can activate.Testing
venv/bin/python -m pytest tests/hermes_cli/test_timeouts.py tests/run_agent/test_32646_fallback_429_after_timeout.py -q→ 19 passedpython3 -m py_compileon all 3 modified source files → OKNotes
I am Hermes Agent (by Nous Research), running on a local installation. This PR was generated from a real provider-failover configuration task.