Sync upstream hermes-agent main into NousAI-Assistant (154 commits, through bab9a85b67) — clean; includes PR #56's deferred commits - #57
Merged
Conversation
…imeout resolver (NousResearch#85125 Phase 1) One shared foundation for the timeout/hang backlog instead of per-incident site-local fixes: - agent/deadline.py: run_bounded_async (thread-timer deadline that survives a blocked event loop, generalizing the telegram adapter primitive), run_bounded_sync, clamp_timeout (kills the NousResearch#83220 time_t OverflowError class at the boundary), resolve_timeout (config.yaml timeouts: section > legacy env bridge > default), kill_process_tree (whole-tree termination for the NousResearch#71148 orphan class), DeadlineExpired (our deadline, mechanically distinct from provider timeouts). - tool_executor._resolve_concurrent_tool_timeout migrates onto the resolver; exact legacy env-var contract preserved (default 420, 0 disables). - timeouts: accepted as a known config root; documented in cli-config.yaml.example. Pure addition otherwise — no behavior change, no new env vars, no cache impact. Later phases (NousResearch#85125) migrate tool-execution, MCP, and subprocess call sites onto these primitives.
- run_bounded_async: cancel + abandon the inner task when the CALLER is cancelled (leak the telegram original also had) - kill_process_tree: check taskkill exit code (Windows contract parity), suppress console flash via windows_hide_flags, and sweep a psutil descendant snapshot taken before signalling — reaches grandchildren in their own setsid sessions and the non-group-leader case (NousResearch#71148 class) - resolve_timeout: reject bool (YAML true would become a 1s deadline) and NaN config values with fall-through instead of resolving unbounded - BoundedResult: kw_only to prevent positional transposition - tests: real clamped-value time_t regression proof, own-session descendant kill, external-cancellation task cleanup, bool/NaN config fall-through; pin already-dead-pid contract
The os.killpg call sits below an early 'if sys.platform == win32: return' so it can never execute on Windows; the scanner is line-based and needs the inline marker.
…roviders Custom providers (custom:xxx) serve their own pricing; models.dev stores OpenRouter prices for the same model ids. The cost guard fired on that foreign pricing and blocked composer/CLI model switches on custom providers with a wildly wrong warning (NousResearch#54348). expensive_model_warning now only trusts model_info/models.dev pricing when the provider maps to a models.dev provider and the info's provider_id matches, and only consults the pricing-entry lookup when the billing route is known. Salvaged from NousResearch#54422; the PR's desktop-hook half predates the use-model-controls rewrite and is superseded by the hook's existing rollback handling.
The custom-provider pricing-trust fix makes provider="test" (not a models.dev provider) correctly silent — use anthropic in the fixture.
Run the expensive-model warning for explicit startup `-m` / `--provider` overrides before the chat loop starts, and fail closed for non-interactive invocations that select an expensive or known-confusing model. Also classify Nous paid-model 404s that say credits are required as billing exhaustion so they fail fast with billing guidance. Tested: - scripts/run_tests.sh tests/hermes_cli/test_cli_startup_model_cost_guard.py tests/hermes_cli/test_model_cost_guard.py tests/agent/test_error_classifier.py -- --tb=short -q
…over the light oneshot fast-path Follow-ups on top of the salvaged NousResearch#70324: - _confirm_startup_expensive_model_override evaluates the unified registry (combined_selection_warning) so id-keyed guards like the data-training-tier warning fire at startup too, not just the cost guard. - The Termux-adjacent light oneshot fast-path (added after the PR branched) ran _run_and_exit_oneshot without the guard — same bug class, third sibling site now covered.
…ousResearch#85954) Clicking an agent in a multi-profile roster pays the entire backend spawn + WebSocket dial cost on first open — several seconds of 'loading' (Bot Mode report). Expose the existing pool-only primitive (openGatewayForProfile: opens/pools the socket WITHOUT activating it, already no-ops for the primary and shared-remote routes) as host.warmProfile(name) so rosters can pre-dial after mount and the first click lands on a live socket. Fire-and-forget by design; failures stay silent — the real open path re-runs its own ensure.
…rust and gpt-5.5-pro confusion nudge (NousResearch#85970) 54cc39a (distrust foreign pricing for custom providers) tested with openai/gpt-5.5-pro fixtures; 83d373a (salvaged NousResearch#70324) made that exact id warn unconditionally as a known-confusion model. Each was green alone; together the distrust tests fail on every main run (slice 6). Use a neutral fixture id for the distrust tests and add a regression test pinning the composed behavior: the id-keyed nudge survives custom-provider pricing distrust.
…e/configure (NousResearch#85963) Three widenings for capabilities UIs (Bot Mode's bot builder): 1. profiles.create share_auth (default false): skip the auth.json COPY so the new profile reads OAuth/token state through the existing global-root fallback and refreshes write through to it. A copy forks token state — the first refresh on either side invalidates the other for single-use refresh tokens; sharing keeps ONE live token pool for the main profile and every bot. Static .env keys still copy (no refresh semantics). Receipt: mirrored.auth = 'shared'. 2. profiles.describe reports mcp_servers [{name, enabled, transport}] from the profile's config. 3. profiles.configure accepts enabled_mcp_servers (replace semantics): toggles via the standard disabled flag; enabling a server the profile lacks copies its definition from the launch profile's catalog (names never invented). Launch catalog read BEFORE the home override flips config resolution. E2E: describe keys include mcp_servers; create with share_auth -> mirrored.auth='shared' + no auth.json in the profile dir; configure applied.mcp_servers=true.
Salvage of NousResearch#79604 (webtecnica) + NousResearch#85721 (pierrenode), combined and rebased onto current main with simplify-code findings folded in. NousResearch#79604: update_session_model() wrote the model name to sessions.model but never persisted the provider into model_config. On resume, the runtime recombined the persisted model with the config.yaml primary provider (which may not serve that model), producing auth errors. Fix: add optional provider parameter to update_session_model, merged into model_config via the shared _merge_model_config_json helper (not hand-rolled SQL). Wire both gateway /model call sites to pass result.target_provider. NousResearch#85721: session_gateway_runtime() had no billing_provider fallback. A CLI session that never ran /model has no gateway_runtime or top-level provider in model_config — billing_provider (written on every session's first accounted API call) is the only durable record. Fix: add billing_provider as the last-resort fallback in session_gateway_runtime(), filtering bare billing buckets (auto/custom) that are not routable identities. Simplify-code findings addressed: - Use _merge_model_config_json instead of 40 lines of branched SQL - Share _BARE_BILLING_PROVIDERS from hermes_state.py (was duplicated as a set in tui_gateway/server.py) - Merge None-filtering from NousResearch#85920 with the billing_provider fallback into one coherent return path Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
…file test test_slash_worker_accepts_profile_home mocks hermes_constants with get_hermes_home=MagicMock(return_value="/tmp/hermes_test"), a str. In production get_hermes_home() returns a Path, and hermes_state.py's module-level DEFAULT_DB_PATH = get_hermes_home() / "state.db" does path division. Under the str mock that becomes str / str, so importing tui_gateway.server inside the patch raises TypeError and the test fails on every main run (slice 4). Wrap the mock return in Path(...) so it matches the real return type. Test-only; no production code change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… notifier HERMES_EXEC_ASK (and gateway platform markers without a notify callback) were short-circuiting interactive CLI into silent pending_approval, so the Approve/Deny panel never appeared. Prefer the registered CLI callback when present, and set HERMES_EXEC_ASK only in start_gateway so importing gateway.run from CLI tools cannot poison the process.
Regression tests for silent pending_approval when ask-mode leaks into interactive CLI, plus a Path-typed hermes_constants mock so the slash-worker profile_home test survives per-file isolation.
The subprocess import test was creating .tmp-hermes-exec-ask-import/ in the repo root without cleanup. Switch to pytest's tmp_path fixture so the temp directory is auto-cleaned and never appears as untracked.
…ritical path The apps/desktop check:test:ui job is the slowest required check (~5m40s wall; vitest self-report: tests 101s, environment 345s, import 302s). With 425 isolated jsdom test files, per-file env boot and module-graph re-import dominate — the tests themselves are ~100s. Replace the single check:test:ui script with three check:test:ui:shard-NofM scripts using vitest's built-in --shard. The workspace-discovery matrix in js-tests.yml already fans out every check:* script as its own job, so no workflow changes are needed. Measured locally (8-core, same suite): unsharded 234s shard 1/3 + 2/3 + 3/3 70s / 80s / 73s (141+141+141 files, all pass) Projected CI gate path: ~5m42s -> ~2m20s per shard in parallel. --no-isolate was evaluated and rejected: 3.5x faster but 103 test files (533 tests) fail without isolation. pool=threads was a wash; vmThreads was slower. The local 'npm run check' aggregate now calls test:ui directly (identical unsharded behavior as before).
…ipts Review finding (reuse): every other check:* script in this file delegates to its base script, keeping the runner command defined once. The shard scripts inlined 'vitest run --project ui' three times, so a later change to test:ui (flags, project rename) would silently drift from what CI runs. Route them through 'npm run test:ui -- --shard=N/3' instead (npm forwards post---- args). Verified: npm run check:test:ui:shard-1of3 passes (142 files) with identical file distribution.
…nner Closes the silent-skip hole reviewers flagged: with the index/count hardcoded in three sibling strings, a copy-paste slip (shard-2of3 running --shard=1/3) or a partial 3->4 migration would silently skip a third of the 428-file suite while CI stays green. scripts/run-ui-shard.mjs parses N/M from npm_lifecycle_event (the script NAME is the single source of truth), validates the package's shard family is exactly 1..M for one M, and delegates through 'npm run test:ui' so the vitest command stays single-sourced. Mutation-verified: shard-9of3 name -> exit 1 'index out of range'; adding shard-4of4 beside the 3-family -> exit 1 'must form exactly 1..M'; correct invocation runs shard 2/3 (142 files) identically to before.
…runner
Review findings: spawnSync('npm', ...) without shell fails on Windows
(npm is npm.cmd; Node >=18.20 throws EINVAL — same handling as
test-desktop.mjs and stage-native-deps.mjs), and a spawn-level failure
exited 1 with no diagnostic. CI is ubuntu-only but the desktop workspace
supports local Windows dev.
NousResearch#52970) (NousResearch#85711) * fix(dashboard): suppress Ctrl+C shutdown traceback * fix(dashboard): extend clean Ctrl+C exit to the Windows serve branch The Windows loop-factory branch (and its pre-0.36 asyncio.run fallback) runs under the same uvicorn capture_signals() re-raise as the POSIX path, so console Ctrl+C leaked the identical KeyboardInterrupt traceback there. Guard both serve calls with the same clean-exit contract, keeping the import-resolution try/except comment accurate (genuine serve-time errors still propagate). Also ports the reworded POSIX-test docstring (the serve path is no longer 'byte-for-byte unchanged'), wraps the POSIX KI test in pytest.fail so a regression reports red instead of aborting the pytest session, and adds the windows_only sibling test. Extends NousResearch#52970 to the whole bug class. * chore: map contributor email for @wangs1203 * test(dashboard): actually exercise the pre-0.36 Windows fallback KI contract Copilot review caught that patching uvicorn._compat.asyncio_run with raising=False makes the import succeed, so _runner is non-None and the extra asyncio.run patch never covered the fallback. Split it out: a dedicated windows_only test halts the _compat import (None in sys.modules) so the fallback branch is genuinely selected, then asserts the same clean-KI contract on bare asyncio.run. --------- Co-authored-by: Emiya·Leon <wangs.coder@gmail.com>
… tarball cache
Every job in the js-tests matrix (~10 jobs/run, 13 after the UI-suite
sharding) runs a full 'npm ci' that deletes and re-extracts the entire
workspace node_modules and reruns all postinstalls — including the
Electron binary fetch (~100MB) — because setup-node's 'cache: npm' only
caches the ~/.npm tarball cache.
Cache the installed tree itself with actions/cache (the SHA-pinned
v4.2.4 already used by e2e-desktop.yml), keyed on the exact lockfile
hash, and skip 'npm ci' on a hit:
- key includes runner.os + node26 + npm12 so a toolchain bump never
reuses a stale tree
- NO restore-keys: a partial hit would leave a stale tree ('npm ci'
skipped means nothing would repair it), so anything but an exact
lockfile match reinstalls from scratch
- distinct keys for the discovery job (--ignore-scripts tree) and the
check jobs (with-scripts tree + ~/.cache/electron), which differ in
postinstall artifacts
Measured from run 31783969717: the npm-ci step is 30-45s per check job.
On warm cache this drops to a few seconds of restore, saving roughly
5-8 runner-minutes per PR run and ~1GB of registry traffic, and taking
~35s off every job on the merge-gate critical path.
…npm upgrade Review findings on the caching commit: - ~/.cache/electron was dead weight: with npm ci skipped on an exact cache hit, the download cache is never read (electron's unpacked binary lives in node_modules/electron/dist, inside the cached tree); it only inflated every saved archive by ~110MB. - 'npm i -g npm@12' ran unconditionally in all 14 matrix jobs (~5-15s each); now a no-op when the bundled npm is already 12.x, which also keeps the installed major aligned with the npm12 cache-key tag. yaml + actionlint pass.
…g.yaml existence (NousResearch#86212) profiles.create inherits the launch profile's provider+model when the caller doesn't pin one — but the gate was 'config.yaml doesn't exist yet'. Voice-section mirroring (NousResearch#85755) runs FIRST and legitimately creates config.yaml (tts/stt), so inheritance silently skipped for every non-clone profile since: the bot's editor showed 'Inherit (launch profile)' while the profile actually had NO model section, and the first message failed with 'No inference provider configured' even though the main agent was authenticated and working (Bot Mode tester report, screenshots). Gate on what we actually care about: the profile's own raw config lacking a complete model section (provider+default). Clones bring their own section and stay untouched; explicit pins unchanged. E2E: create receipt now model_inherited=true and the fresh profile's config.yaml carries the launch profile's provider/model.
…earch#86227) Three fixes for what profiles.describe & friends expose (tester report with screenshot): 1. describe's toolsets used the RAW registry (get_all_toolsets) — leaking internal platform composites (hermes-discord, hermes-cron, feishu_drive, discord_admin, desktop_ui, ...) that are gated, platform-restricted, or deliberately hidden from users — and reported everything enabled whenever the profile had no pin. Now: the same filtered universe the `hermes tools` checklist offers (_get_effective_configurable_toolsets, platform-filtered), with enablement resolved via _get_platform_tools like the runtime. 2. skills.manage accepts optional `profile`: list/install scoped to that profile's skills dir via the home override, so editors can manage a bot's skills (incl. hub installs) from the main window. search/browse/inspect (the hub catalog) unchanged. 3. New mcp.catalog method: the bundled MCP catalog with per-profile installed/enabled state + required env keys, so capability UIs can offer the full menu and route un-setup entries through setup instead of silently listing dead servers. E2E: describe now returns only user-facing toolsets with honest enabled flags; mcp.catalog returns the 5-entry catalog; skills list scopes to the named profile.
…ibility A captured native-compaction checkpoint lives in the persisted codex_reasoning_items sidecar, but the wire restructure that follows it (prune_pre_checkpoint_items) ran unconditionally: the native gate only decided whether context_management went into the request, and no signal from it ever reached _chat_messages_to_responses_input. So a single checkpoint kept deleting every pre-checkpoint item from all later requests — after a mid-session swap out of the gpt-5.6 family, after compression.enabled: false, after the rejection kill switch, and after a session resume that reloads the sidecar from state.db. The model receiving the opaque blob was no longer the one able to decode it, and nothing was logged. Thread a single native_compaction_eligible boolean, derived from the same value that gates the context_management field, into the converter. When ineligible: do not replay type: "compaction" items and do not prune. Safe because native compaction never truncates Hermes' local history, so the fallback still carries the full conversation. All Responses call sites are covered: build_kwargs and convert_messages derive the flag via _native_compaction_active, the auxiliary/compression client is explicitly ineligible, and the converter defaults to False (pre-feature wire) so future call sites are safe by construction. Fixes NousResearch#85914 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… enabled (NousResearch#86239) Follow-up to NousResearch#86227: _DEFAULT_OFF_TOOLSETS entries (a2a, spotify, discord, video, x_search, ...) and the region-specific yuanbao are global opt-ins configured via `hermes tools`/Settings; showing them unchecked in every per-profile capabilities editor is noise, and showing a2a at all confuses users who never enabled agent-to-agent serving (tester report). Enabled ones still show — hiding an ACTIVE toolset would misrepresent the profile.
…arch#86243) The /docs/skills hub page gains an embed mode for host apps that iframe it as a skill PICKER (first consumer: Bot Mode's agent editor). With ?embed=picker: - docs chrome (navbar/footer/hero) is hidden - every card gains '+ Add to this Agent', which posts {type: 'hermes-skill-pick', name, identifier, installCmd, source} to the parent window The page never installs anything — the host validates event.origin and performs the install through its own gateway (skills.manage). Normal page rendering is untouched (no query param = no changes).
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…ault (NousResearch#86414) * fix(desktop): the main agent's model pick persists as the profile default Reported: the default bot switches to the OpenAI API account instead of the user's subscription, and doesn't retain the previous selection. Root cause: the composer model picker always sent the switch as --session scope, even for the PRIMARY profile's main agent. So the pick never wrote config.yaml model.provider — and with model.provider unset, resolve_provider('auto') falls through to a leftover OPENAI_API_KEY env var and picks OpenAI/OpenRouter. The subscription the user selected was only ever a per-session override that evaporated on the next session. Fix: when the pick targets the primary profile's main agent (touchesPrimary), send --global so it persists to config.yaml (model.default + model.provider) via the existing model-switch persist path. A SET model.provider already outranks the OPENAI_API_KEY env var in resolve_provider (tier 2 vs tier 3), so the main agent now keeps the chosen provider across restarts. Secondary chat tiles stay --session so picking a model in one chat never rewrites the profile default (the cross-session-contamination guard the old comment protected). No change to resolve_provider's priority chain, so NousResearch#29285 (an explicit env key beating a STALE oauth login) is untouched — we simply make the user's explicit main-agent selection the config default it always should have been. * MoA presets stay session-scoped; update tests for primary-persist intent Fix CI (ui shard 3of3): the primary main-agent pick now persists via --global, but MoA (mixture-of-agents) presets must NOT — a transient orchestration choice can't become the global gateway default. Exclude provider==='moa' from the persist path (stays --session). Update the primary-picker test to assert --global (the new intent) and keep the MoA + secondary-tile tests asserting --session (the guards that prove the narrowing). 19/19 green locally.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
test_run_prompt_submit_requeues_all_unstarted_notifications_with_real_threading
failed twice in one hour on CI slices for two UNRELATED PRs (86371,
86374) with `assert set() == {proc_batch_2, proc_batch_3}`. Root cause:
session.init/create tests earlier in the file start real per-session
notification poller daemon threads and never stop them. Those pollers
outlive their test and keep polling the PROCESS-GLOBAL
process_registry.completion_queue, stealing-and-requeuing the target
test's events mid-assertion so its bounded drain loop can starve.
Reproduced: with 30 leaked foreign-session pollers injected via a
sabotage conftest, the target test fails standalone ~1 in 3 runs with
the exact CI assertion; with the reap fixture active it passed 8/8
under the same sabotage.
Fix:
- tui_gateway/server.py: _start_notification_poller registers
(stop_event, thread) in module-level _notification_pollers (pruned of
dead threads on each spawn; threads get a stable
tui-notif-poller-<sid> name for debugging).
- tests/test_tui_gateway_server.py: autouse fixture sets every
registered live poller's stop event after each test and joins them
under ONE shared 3s budget (the poller loop wakes at least every
0.5s), so no poller survives into the next test. No per-thread
timeout, no session-dict mutation — a first draft that mutated
session state and joined per-thread hung the file; full-file runtime
with this version is 15.2s vs 13.2s baseline.
…ousResearch#86473) Adds the full MCP setup surface as profile-scoped gateway RPCs so a desktop client (Bot Mode's bot editor, the core Capabilities tab) can add/configure/test/authenticate/remove MCP servers for ANY profile, not just the launch profile: - mcp.servers.list (profile) -> configured servers (transport, auth, oauth_tokens_present, enabled, tool names; no secret values) - mcp.servers.add (profile, name, config|preset, bearer_token?) -> reuses mcp_config._apply_mcp_preset / _save_mcp_server / _save_bearer_auth_token - mcp.servers.set_api_key (profile, name, value, env_var?) -> http auth header template or stdio env ref, via save_env_value - mcp.servers.test (profile, name) -> _probe_single_server + oauth state - mcp.servers.remove (profile, name) - mcp.servers.oauth.start/poll (profile, name[, session_id]) -> mirrors the PROVIDER oauth session/poll model (not the FastAPI dashboard flow): a background worker drives the same interactive machinery 'hermes mcp login' uses, capturing the browser redirect on a local loopback listener. Client opens auth_url via openExternal and polls until status=='approved'. All handlers are profile-scoped via set_hermes_home_override in try/finally (mirrors skills.manage). Shared helpers live in tui_gateway/mcp_rpc_helpers.py and are aliased onto server.py's namespace so the rebound handler bodies (HandlerRegistry.install) can resolve them — a plain def in methods_tools is unreachable post-rebind. Reuses hermes_cli/mcp_config.py throughout; no config logic duplicated; no raw yaml near config.yaml (config-read-guard safe). Tests: tests/tui_gateway/test_mcp_profile_rpcs.py, 8 E2E against real temp HERMES_HOME profiles asserting add/list/set_api_key/remove land in the RIGHT profile's config.yaml and not the launch profile's. 8/8. Registration + live mcp.servers.list verified in an imported gateway. Co-authored-by: Teknium <teknium1@users.noreply.github.com>
…cess env The browser-use CLI runs under its own Python (uv tool / uvx), which can differ from Hermes's venv interpreter. PYTHONPATH/PYTHONHOME inherited from the agent process point at Hermes's venv site-packages, and a child interpreter honors them ahead of its own — so the CLI imported compiled C-extensions (pydantic_core) built for the wrong interpreter and crashed with ABI mismatch / ModuleNotFoundError (issues 83427, 84841, 86006, 86104; hits the desktop backend on py3.14 and any shell exporting PYTHONPATH). Strip both vars in _base_subprocess_env() — the CLI manages its own environment and never needs Hermes's import path. Salvaged from PR 83471 by Benjamin (@n1majne3), the earliest of two independent fixes (also PR 84022 by @jklance16, PYTHONPATH-only); regression test covers both vars and preserves unrelated env.
Follow-up on the NousResearch#83854 salvage: prepend $HERMES_HOME/bin ahead of the venv and user-local bin dirs, matching the managed-first Browser Use CLI resolution policy — the worker resolves the same canonical binary the agent process does.
…ly discarding them
Earlier releases accepted api_mode: openai on custom provider entries.
The canonical transport set is now {chat_completions, codex_responses,
anthropic_messages, bedrock_converse, codex_app_server}, and an
unrecognized value was silently ignored at both consumption sites
(_normalize_custom_provider_entry passes the raw string through and
agent_init's accepted-set check drops it; _parse_api_mode returns None),
falling through to hostname-based detection.
For hosts with a detection rule the provider silently switches
transports after an update. Observed live: a custom entry for
api.actual.inc with api_mode: openai (valid when written) flipped to
codex_responses via the hostname rule, and every reasoning-bearing
request to the relay's /v1/responses failed with a wrapped non-JSON
error while /v1/chat/completions worked throughout.
Fix: one shared alias map (_canonical_api_mode) consulted by both
sites. openai/openai_chat -> chat_completions, responses ->
codex_responses, anthropic/messages -> anthropic_messages, bedrock ->
bedrock_converse. Canonical names and unknown values pass through
unchanged, so invalid-config behavior is untouched.
Tests: alias map contract (every alias lands in _VALID_API_MODES),
normalizer canonicalization incl. the transport: key alias, and the
runtime gate accepting legacy spellings while still rejecting unknowns.
…probes Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the auxiliary clients (NousResearch#56681), but the endpoint discovery and pricing probes did not. Both probe families resolved TLS from process-wide env vars only: - the requests-based metadata/pricing probe (agent/model_metadata.py::_resolve_requests_verify) - the urllib-based /models catalog probe (hermes_cli/models.py::probe_api_models) A custom endpoint whose chain verifies against the provider's configured bundle, but not the process SSL_CERT_FILE, then logged a spurious CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked. Pointing a global CA env var at the bundle fixes it but changes verification for every provider, defeating the point of a per-provider setting. This threads the selected provider's TLS settings into both probe paths, reusing get_custom_provider_tls_settings so there is no second precedence chain: - _resolve_requests_verify(base_url) looks up the provider's ssl_verify / ssl_ca_cert before falling back to the env vars. Callers with no base_url keep the exact env-only behavior. - probe_api_models builds an ssl.SSLContext from the provider settings and passes it through open_credentialed_url, which gains an ssl_context seam on the cloned secure opener. Unmatched or public endpoints pass None and keep urllib's default policy. Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe families (provider CA, ssl_verify:false, unmatched, missing file, config lookup failure) plus end-to-end assertions that the resolved verify value and SSLContext actually reach the request seam. Verified against the neighboring metadata, pricing, TLS, and urllib-security suites (266 tests) with no regressions.
Upstream added stream=True to the /models metadata probe; widen the fake_get signature so the captured verify assertion still runs.
…direct guard ActualProfile.fetch_models() overrides ProviderProfile's default implementation with its own Actual-specific base_url resolution (ACTUAL_BASE_URL env var, hosted-vs-local normalization), but called raw urllib.request.urlopen(req, timeout=timeout) directly instead of the base class's open_credentialed_url(). Every other provider either uses the base class default or forwards to it via super() and gets SafeCredentialRedirectHandler for free — Actual is the only provider that attaches a Bearer token to its own Request object and opens it with the stdlib's default redirect handling, which forwards every header, including Authorization, across a cross-origin redirect. Actual's own feature surface makes the trigger realistic: ACTUAL_BASE_URL is a first-class, documented way to point this provider at a self-hosted or local-offline endpoint (see the local-loopback no-auth path already handled elsewhere in this provider), so a misconfigured or compromised endpoint 302-ing to another host leaks ACTUAL_API_KEY to it. Fix: import and call the same open_credentialed_url() the base class uses, keeping Actual's own URL-resolution logic unchanged. Adds an end-to-end regression test using two real local HTTP servers (no mocking of the security module itself) — one redirects, the other records the Authorization header it receives — mirroring test_urllib_security.py's own redirect tests. Also repoints the existing fetch_models test's mock from urllib.request.urlopen to hermes_cli.urllib_security.open_credentialed_url, since fetch_models no longer calls the former. Mutation-verified: the new redirect test fails on pre-fix code with the Authorization header observed at the redirect target.
…gration) (NousResearch#86506) delegation.max_iterations is the per-subagent tool-call budget. The old default of 50 truncated substantial delegated work: leaf agents spend ~15-20 turns on reconnaissance before producing output, then ran out of budget mid-task and returned 'completed but unfinished' summaries. 250 gives real delegated work room to finish. Changes: - config_defaults.py: delegation.max_iterations 50 -> 250; _config_version 35 -> 36 - tools/delegate_tool.py: DEFAULT_MAX_ITERATIONS fallback 50 -> 250 (kept in sync with the shipped default to prevent drift) - config_migrations.py: _migrate_to_36 lifts configs still pinned at exactly the OLD default 50 -> 250 on update, so existing installs inherit the new headroom. Any other explicit value (deliberate override) is preserved; unset inherits 250 at read time. - cli-config.yaml.example: doc the new default The cap is per-child and children run concurrently (max_concurrent_children default 3), so this raises worst-case fan-out cost; delegation.child_timeout_seconds (default 0 = off) remains available as a wall-clock guardrail, and users can still pin a lower max_iterations explicitly. Verified: migration lifts 50->250, preserves a deliberate 120, leaves unset untouched (3/3); DEFAULT_CONFIG reads version=36, max_iterations=250, fallback=250.
…turn The statusbar gauge painted only what the backend reported as measured occupancy, which a session has none of until a turn runs in this process. Turning the gauge on mid-conversation, or resuming a chat, therefore showed nothing until the next message. Fetch session.context_breakdown as soon as the gauge is on screen instead of when its popover opens. It is the same read-only estimate the popover already used (chars/4 over the live prompt, tools and transcript — no provider call), and it reports the measured figure once the backend has one. The popover becomes presentational and reads the gauge s merged usage, so the bar and the panel cannot disagree.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…hrough bab9a85) Clean merge - no conflicts. Includes the 6 commits deferred by PR #56's partial sync: upstream resolved the model-cost-guard contradiction itself in 16b54e2 (test-side fix, via their PR NousResearch#85970), confirming the partial-sync call. Replay (git merge-tree) produced zero conflicts and the merged tree passed the recorded semantic checks: drift grep for 'Hermes Desktop' in apps/desktop/{src,electron} is empty, and all brand carve-out values are intact (upstream's apps/desktop/package.json edits - UI-test sharding scripts, get-windows to optionalDependencies - auto-merged around the branded fields). CI files touched (1), benign: js-tests.yml adds exact-key node_modules caching (SHA-pinned actions/cache, no restore-keys) and skips the npm@12 global install when already on 12.x. No new actions, secrets, triggers, or permissions. package-lock.json change mirrors the get-windows optionalDependencies move - no version changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT
૮ >ﻌ< ა ci reviewran on 1fa00fc — Sync upstream hermes-agent main into NousAI-Assistant (154 c
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Squashing a sync PR flattens upstream history and breaks future syncs. Merge with the "Create a merge commit" method only.
What this is
Routine upstream sync:
nousresearch/hermes-agent:main→NousAI-Assistant, picking up where PR #56's partial sync left off.9a8af40192(exactly PR Partial sync of upstream hermes-agent main (17 commits, through 9a8af40192) — stops before upstream-red cost-guard commit #56's sync point) through tipbab9a85b67— includes the 6 commits PR Partial sync of upstream hermes-agent main (17 commits, through 9a8af40192) — stops before upstream-red cost-guard commit #56 deferredgit merge-treereplay clean (single merge base, deep fetch verified), committed tree byte-identical to the replay treeThe deferred cost-guard breakage is resolved
Upstream fixed its own contradiction in
16b54e2a0f(via their PR NousResearch#85970) — a test-side fix: the fixtures were rewritten tovendor/priced-modelbecauseopenai/gpt-5.5-procarries an intentionally id-keyed confusion nudge that is independent of pricing trust (now explicitly documented in a dedicated test). Verified locally on the merged tree: behavior and tests are consistent. (This also validates PR #56's partial-sync choice over guessing a lockstep implementation fix, which would have guessed wrong.)CI-sensitive file review (
ci-reviewedapplied per recorded policy).github/workflows/js-tests.yml(+48/−2): node_modules caching keyed exactly on the lockfile via SHA-pinnedactions/cache(no restore-keys, so no stale-tree risk;npm ciskipped only on exact hits), and the global npm@12 install becomes a no-op when already on 12.x. No new actions, permissions, triggers, or outbound calls.package-lock.json: only mirrorsget-windows 9.3.0moving fromdependenciestooptionalDependenciesinapps/desktop/package.json— no version changes, no new packages.Semantic checks (all pass)
"Hermes Desktop"inapps/desktop/{src,electron}: emptyapps/desktop/package.jsonedits (UI-test 3-way sharding scripts,get-windowsoptional) auto-merged cleanly around the branded fieldsfalse &&guards still in ci.yml)Upstream highlights
/looprecurring-command feature (hermes_cli/loops.py + gateway/TUI wiring + docs)🤖 Generated with Claude Code
https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT
Generated by Claude Code