Skip to content

Sync upstream hermes-agent main into NousAI-Assistant (154 commits, through bab9a85b67) — clean; includes PR #56's deferred commits - #57

Merged
uaixo merged 155 commits into
NousAI-Assistantfrom
claude/nousai-sync-upstream-20260815
Aug 15, 2026
Merged

uaixo merged 155 commits into
NousAI-Assistantfrom
claude/nousai-sync-upstream-20260815

Conversation

@uaixo

@uaixo uaixo commented Aug 15, 2026

Copy link
Copy Markdown
Owner

⚠️ MERGE WITH MERGE COMMIT — DO NOT SQUASH

Squashing a sync PR flattens upstream history and breaks future syncs. Merge with the "Create a merge commit" method only.

What this is

Routine upstream sync: nousresearch/hermes-agent:main → NousAI-Assistant, picking up where PR #56's partial sync left off.

The deferred cost-guard breakage is resolved

Upstream fixed its own contradiction in 16b54e2a0f (via their PR NousResearch#85970) — a test-side fix: the fixtures were rewritten to vendor/priced-model because openai/gpt-5.5-pro carries an intentionally id-keyed confusion nudge that is independent of pricing trust (now explicitly documented in a dedicated test). Verified locally on the merged tree: behavior and tests are consistent. (This also validates PR #56's partial-sync choice over guessing a lockstep implementation fix, which would have guessed wrong.)

CI-sensitive file review (ci-reviewed applied per recorded policy)

  • .github/workflows/js-tests.yml (+48/−2): node_modules caching keyed exactly on the lockfile via SHA-pinned actions/cache (no restore-keys, so no stale-tree risk; npm ci skipped only on exact hits), and the global npm@12 install becomes a no-op when already on 12.x. No new actions, permissions, triggers, or outbound calls.
  • package-lock.json: only mirrors get-windows 9.3.0 moving from dependencies to optionalDependencies in apps/desktop/package.json — no version changes, no new packages.

Semantic checks (all pass)

  • Drift grep "Hermes Desktop" in apps/desktop/{src,electron}: empty
  • Brand carve-out values intact; upstream's apps/desktop/package.json edits (UI-test 3-way sharding scripts, get-windows optional) auto-merged cleanly around the branded fields
  • Desktop E2E remains disabled upstream (both false && guards still in ci.yml)

Upstream highlights

  • New /loop recurring-command feature (hermes_cli/loops.py + gateway/TUI wiring + docs)
  • Turn-lease work: cross-process turn leases, session turn-lease store, sequential tool timeouts, deadline module
  • Desktop: external-terminal launcher, context-usage breakdown panel, UI-test sharding
  • MCP OAuth session handling in the TUI gateway; custom-provider CA probes; binary-document write guard
  • 19 new contributor mappings (from upstream)

🤖 Generated with Claude Code

https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT


Generated by Claude Code

kshitijk4poor and others added 30 commits August 13, 2026 13:33
…imeout resolver (NousResearch#85125 Phase 1)

One shared foundation for the timeout/hang backlog instead of per-incident
site-local fixes:

- agent/deadline.py: run_bounded_async (thread-timer deadline that survives
  a blocked event loop, generalizing the telegram adapter primitive),
  run_bounded_sync, clamp_timeout (kills the NousResearch#83220 time_t OverflowError
  class at the boundary), resolve_timeout (config.yaml timeouts: section >
  legacy env bridge > default), kill_process_tree (whole-tree termination
  for the NousResearch#71148 orphan class), DeadlineExpired (our deadline, mechanically
  distinct from provider timeouts).
- tool_executor._resolve_concurrent_tool_timeout migrates onto the resolver;
  exact legacy env-var contract preserved (default 420, 0 disables).
- timeouts: accepted as a known config root; documented in
  cli-config.yaml.example.

Pure addition otherwise — no behavior change, no new env vars, no cache
impact. Later phases (NousResearch#85125) migrate tool-execution, MCP, and subprocess
call sites onto these primitives.
- run_bounded_async: cancel + abandon the inner task when the CALLER is
  cancelled (leak the telegram original also had)
- kill_process_tree: check taskkill exit code (Windows contract parity),
  suppress console flash via windows_hide_flags, and sweep a psutil
  descendant snapshot taken before signalling — reaches grandchildren in
  their own setsid sessions and the non-group-leader case (NousResearch#71148 class)
- resolve_timeout: reject bool (YAML true would become a 1s deadline) and
  NaN config values with fall-through instead of resolving unbounded
- BoundedResult: kw_only to prevent positional transposition
- tests: real clamped-value time_t regression proof, own-session
  descendant kill, external-cancellation task cleanup, bool/NaN config
  fall-through; pin already-dead-pid contract
The os.killpg call sits below an early 'if sys.platform == win32: return'
so it can never execute on Windows; the scanner is line-based and needs
the inline marker.
…roviders

Custom providers (custom:xxx) serve their own pricing; models.dev stores
OpenRouter prices for the same model ids. The cost guard fired on that
foreign pricing and blocked composer/CLI model switches on custom
providers with a wildly wrong warning (NousResearch#54348).

expensive_model_warning now only trusts model_info/models.dev pricing
when the provider maps to a models.dev provider and the info's
provider_id matches, and only consults the pricing-entry lookup when
the billing route is known. Salvaged from NousResearch#54422; the PR's desktop-hook
half predates the use-model-controls rewrite and is superseded by the
hook's existing rollback handling.
The custom-provider pricing-trust fix makes provider="test" (not a
models.dev provider) correctly silent — use anthropic in the fixture.
Run the expensive-model warning for explicit startup `-m` / `--provider`
overrides before the chat loop starts, and fail closed for non-interactive
invocations that select an expensive or known-confusing model.

Also classify Nous paid-model 404s that say credits are required as billing
exhaustion so they fail fast with billing guidance.

Tested:
- scripts/run_tests.sh tests/hermes_cli/test_cli_startup_model_cost_guard.py tests/hermes_cli/test_model_cost_guard.py tests/agent/test_error_classifier.py -- --tb=short -q
…over the light oneshot fast-path

Follow-ups on top of the salvaged NousResearch#70324:
- _confirm_startup_expensive_model_override evaluates the unified
  registry (combined_selection_warning) so id-keyed guards like the
  data-training-tier warning fire at startup too, not just the cost guard.
- The Termux-adjacent light oneshot fast-path (added after the PR
  branched) ran _run_and_exit_oneshot without the guard — same bug
  class, third sibling site now covered.
…ousResearch#85954)

Clicking an agent in a multi-profile roster pays the entire backend
spawn + WebSocket dial cost on first open — several seconds of
'loading' (Bot Mode report). Expose the existing pool-only primitive
(openGatewayForProfile: opens/pools the socket WITHOUT activating it,
already no-ops for the primary and shared-remote routes) as
host.warmProfile(name) so rosters can pre-dial after mount and the
first click lands on a live socket. Fire-and-forget by design;
failures stay silent — the real open path re-runs its own ensure.
…rust and gpt-5.5-pro confusion nudge (NousResearch#85970)

54cc39a (distrust foreign pricing for custom providers) tested with
openai/gpt-5.5-pro fixtures; 83d373a (salvaged NousResearch#70324) made that exact
id warn unconditionally as a known-confusion model. Each was green alone;
together the distrust tests fail on every main run (slice 6).

Use a neutral fixture id for the distrust tests and add a regression test
pinning the composed behavior: the id-keyed nudge survives custom-provider
pricing distrust.
…e/configure (NousResearch#85963)

Three widenings for capabilities UIs (Bot Mode's bot builder):

1. profiles.create share_auth (default false): skip the auth.json
   COPY so the new profile reads OAuth/token state through the
   existing global-root fallback and refreshes write through to it.
   A copy forks token state — the first refresh on either side
   invalidates the other for single-use refresh tokens; sharing keeps
   ONE live token pool for the main profile and every bot. Static
   .env keys still copy (no refresh semantics). Receipt:
   mirrored.auth = 'shared'.

2. profiles.describe reports mcp_servers
   [{name, enabled, transport}] from the profile's config.

3. profiles.configure accepts enabled_mcp_servers (replace
   semantics): toggles via the standard disabled flag; enabling a
   server the profile lacks copies its definition from the launch
   profile's catalog (names never invented). Launch catalog read
   BEFORE the home override flips config resolution.

E2E: describe keys include mcp_servers; create with share_auth ->
mirrored.auth='shared' + no auth.json in the profile dir; configure
applied.mcp_servers=true.
Salvage of NousResearch#79604 (webtecnica) + NousResearch#85721 (pierrenode), combined and
rebased onto current main with simplify-code findings folded in.

NousResearch#79604: update_session_model() wrote the model name to sessions.model
but never persisted the provider into model_config. On resume, the
runtime recombined the persisted model with the config.yaml primary
provider (which may not serve that model), producing auth errors.
Fix: add optional provider parameter to update_session_model, merged
into model_config via the shared _merge_model_config_json helper (not
hand-rolled SQL). Wire both gateway /model call sites to pass
result.target_provider.

NousResearch#85721: session_gateway_runtime() had no billing_provider fallback.
A CLI session that never ran /model has no gateway_runtime or
top-level provider in model_config — billing_provider (written on
every session's first accounted API call) is the only durable record.
Fix: add billing_provider as the last-resort fallback in
session_gateway_runtime(), filtering bare billing buckets (auto/custom)
that are not routable identities.

Simplify-code findings addressed:
- Use _merge_model_config_json instead of 40 lines of branched SQL
- Share _BARE_BILLING_PROVIDERS from hermes_state.py (was duplicated
  as a set in tui_gateway/server.py)
- Merge None-filtering from NousResearch#85920 with the billing_provider fallback
  into one coherent return path

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
…file test

test_slash_worker_accepts_profile_home mocks hermes_constants with
get_hermes_home=MagicMock(return_value="/tmp/hermes_test"), a str. In
production get_hermes_home() returns a Path, and hermes_state.py's
module-level DEFAULT_DB_PATH = get_hermes_home() / "state.db" does path
division. Under the str mock that becomes str / str, so importing
tui_gateway.server inside the patch raises TypeError and the test fails on
every main run (slice 4). Wrap the mock return in Path(...) so it matches the
real return type. Test-only; no production code change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… notifier

HERMES_EXEC_ASK (and gateway platform markers without a notify callback)
were short-circuiting interactive CLI into silent pending_approval, so
the Approve/Deny panel never appeared. Prefer the registered CLI callback
when present, and set HERMES_EXEC_ASK only in start_gateway so importing
gateway.run from CLI tools cannot poison the process.
Regression tests for silent pending_approval when ask-mode leaks into
interactive CLI, plus a Path-typed hermes_constants mock so the
slash-worker profile_home test survives per-file isolation.
The subprocess import test was creating .tmp-hermes-exec-ask-import/
in the repo root without cleanup. Switch to pytest's tmp_path fixture
so the temp directory is auto-cleaned and never appears as untracked.
…ritical path

The apps/desktop check:test:ui job is the slowest required check
(~5m40s wall; vitest self-report: tests 101s, environment 345s,
import 302s). With 425 isolated jsdom test files, per-file env boot
and module-graph re-import dominate — the tests themselves are ~100s.

Replace the single check:test:ui script with three check:test:ui:shard-NofM
scripts using vitest's built-in --shard. The workspace-discovery matrix in
js-tests.yml already fans out every check:* script as its own job, so no
workflow changes are needed.

Measured locally (8-core, same suite):
  unsharded            234s
  shard 1/3 + 2/3 + 3/3  70s / 80s / 73s   (141+141+141 files, all pass)

Projected CI gate path: ~5m42s -> ~2m20s per shard in parallel.

--no-isolate was evaluated and rejected: 3.5x faster but 103 test files
(533 tests) fail without isolation. pool=threads was a wash; vmThreads
was slower. The local 'npm run check' aggregate now calls test:ui
directly (identical unsharded behavior as before).
…ipts

Review finding (reuse): every other check:* script in this file delegates
to its base script, keeping the runner command defined once. The shard
scripts inlined 'vitest run --project ui' three times, so a later change
to test:ui (flags, project rename) would silently drift from what CI runs.
Route them through 'npm run test:ui -- --shard=N/3' instead (npm forwards
post---- args).

Verified: npm run check:test:ui:shard-1of3 passes (142 files) with
identical file distribution.
…nner

Closes the silent-skip hole reviewers flagged: with the index/count
hardcoded in three sibling strings, a copy-paste slip (shard-2of3 running
--shard=1/3) or a partial 3->4 migration would silently skip a third of
the 428-file suite while CI stays green.

scripts/run-ui-shard.mjs parses N/M from npm_lifecycle_event (the script
NAME is the single source of truth), validates the package's shard family
is exactly 1..M for one M, and delegates through 'npm run test:ui' so the
vitest command stays single-sourced.

Mutation-verified: shard-9of3 name -> exit 1 'index out of range';
adding shard-4of4 beside the 3-family -> exit 1 'must form exactly 1..M';
correct invocation runs shard 2/3 (142 files) identically to before.
…runner

Review findings: spawnSync('npm', ...) without shell fails on Windows
(npm is npm.cmd; Node >=18.20 throws EINVAL — same handling as
test-desktop.mjs and stage-native-deps.mjs), and a spawn-level failure
exited 1 with no diagnostic. CI is ubuntu-only but the desktop workspace
supports local Windows dev.
NousResearch#52970) (NousResearch#85711)

* fix(dashboard): suppress Ctrl+C shutdown traceback

* fix(dashboard): extend clean Ctrl+C exit to the Windows serve branch

The Windows loop-factory branch (and its pre-0.36 asyncio.run fallback)
runs under the same uvicorn capture_signals() re-raise as the POSIX
path, so console Ctrl+C leaked the identical KeyboardInterrupt
traceback there. Guard both serve calls with the same clean-exit
contract, keeping the import-resolution try/except comment accurate
(genuine serve-time errors still propagate).

Also ports the reworded POSIX-test docstring (the serve path is no
longer 'byte-for-byte unchanged'), wraps the POSIX KI test in
pytest.fail so a regression reports red instead of aborting the pytest
session, and adds the windows_only sibling test.

Extends NousResearch#52970 to the whole bug class.

* chore: map contributor email for @wangs1203

* test(dashboard): actually exercise the pre-0.36 Windows fallback KI contract

Copilot review caught that patching uvicorn._compat.asyncio_run with
raising=False makes the import succeed, so _runner is non-None and the
extra asyncio.run patch never covered the fallback. Split it out: a
dedicated windows_only test halts the _compat import (None in
sys.modules) so the fallback branch is genuinely selected, then asserts
the same clean-KI contract on bare asyncio.run.

---------

Co-authored-by: Emiya·Leon <wangs.coder@gmail.com>
… tarball cache

Every job in the js-tests matrix (~10 jobs/run, 13 after the UI-suite
sharding) runs a full 'npm ci' that deletes and re-extracts the entire
workspace node_modules and reruns all postinstalls — including the
Electron binary fetch (~100MB) — because setup-node's 'cache: npm' only
caches the ~/.npm tarball cache.

Cache the installed tree itself with actions/cache (the SHA-pinned
v4.2.4 already used by e2e-desktop.yml), keyed on the exact lockfile
hash, and skip 'npm ci' on a hit:

- key includes runner.os + node26 + npm12 so a toolchain bump never
  reuses a stale tree
- NO restore-keys: a partial hit would leave a stale tree ('npm ci'
  skipped means nothing would repair it), so anything but an exact
  lockfile match reinstalls from scratch
- distinct keys for the discovery job (--ignore-scripts tree) and the
  check jobs (with-scripts tree + ~/.cache/electron), which differ in
  postinstall artifacts

Measured from run 31783969717: the npm-ci step is 30-45s per check job.
On warm cache this drops to a few seconds of restore, saving roughly
5-8 runner-minutes per PR run and ~1GB of registry traffic, and taking
~35s off every job on the merge-gate critical path.
…npm upgrade

Review findings on the caching commit:

- ~/.cache/electron was dead weight: with npm ci skipped on an exact
  cache hit, the download cache is never read (electron's unpacked
  binary lives in node_modules/electron/dist, inside the cached tree);
  it only inflated every saved archive by ~110MB.
- 'npm i -g npm@12' ran unconditionally in all 14 matrix jobs
  (~5-15s each); now a no-op when the bundled npm is already 12.x,
  which also keeps the installed major aligned with the npm12
  cache-key tag.

yaml + actionlint pass.
…g.yaml existence (NousResearch#86212)

profiles.create inherits the launch profile's provider+model when the
caller doesn't pin one — but the gate was 'config.yaml doesn't exist
yet'. Voice-section mirroring (NousResearch#85755) runs FIRST and legitimately
creates config.yaml (tts/stt), so inheritance silently skipped for
every non-clone profile since: the bot's editor showed 'Inherit
(launch profile)' while the profile actually had NO model section,
and the first message failed with 'No inference provider configured'
even though the main agent was authenticated and working (Bot Mode
tester report, screenshots).

Gate on what we actually care about: the profile's own raw config
lacking a complete model section (provider+default). Clones bring
their own section and stay untouched; explicit pins unchanged.

E2E: create receipt now model_inherited=true and the fresh profile's
config.yaml carries the launch profile's provider/model.
…earch#86227)

Three fixes for what profiles.describe & friends expose (tester
report with screenshot):

1. describe's toolsets used the RAW registry (get_all_toolsets) —
   leaking internal platform composites (hermes-discord, hermes-cron,
   feishu_drive, discord_admin, desktop_ui, ...) that are gated,
   platform-restricted, or deliberately hidden from users — and
   reported everything enabled whenever the profile had no pin.
   Now: the same filtered universe the `hermes tools` checklist
   offers (_get_effective_configurable_toolsets, platform-filtered),
   with enablement resolved via _get_platform_tools like the runtime.

2. skills.manage accepts optional `profile`: list/install scoped to
   that profile's skills dir via the home override, so editors can
   manage a bot's skills (incl. hub installs) from the main window.
   search/browse/inspect (the hub catalog) unchanged.

3. New mcp.catalog method: the bundled MCP catalog with per-profile
   installed/enabled state + required env keys, so capability UIs can
   offer the full menu and route un-setup entries through setup
   instead of silently listing dead servers.

E2E: describe now returns only user-facing toolsets with honest
enabled flags; mcp.catalog returns the 5-entry catalog; skills list
scopes to the named profile.
…ibility

A captured native-compaction checkpoint lives in the persisted
codex_reasoning_items sidecar, but the wire restructure that follows it
(prune_pre_checkpoint_items) ran unconditionally: the native gate only
decided whether context_management went into the request, and no signal
from it ever reached _chat_messages_to_responses_input.

So a single checkpoint kept deleting every pre-checkpoint item from all
later requests — after a mid-session swap out of the gpt-5.6 family,
after compression.enabled: false, after the rejection kill switch, and
after a session resume that reloads the sidecar from state.db. The model
receiving the opaque blob was no longer the one able to decode it, and
nothing was logged.

Thread a single native_compaction_eligible boolean, derived from the same
value that gates the context_management field, into the converter. When
ineligible: do not replay type: "compaction" items and do not prune. Safe
because native compaction never truncates Hermes' local history, so the
fallback still carries the full conversation.

All Responses call sites are covered: build_kwargs and convert_messages
derive the flag via _native_compaction_active, the auxiliary/compression
client is explicitly ineligible, and the converter defaults to False
(pre-feature wire) so future call sites are safe by construction.

Fixes NousResearch#85914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… enabled (NousResearch#86239)

Follow-up to NousResearch#86227: _DEFAULT_OFF_TOOLSETS entries (a2a, spotify,
discord, video, x_search, ...) and the region-specific yuanbao are
global opt-ins configured via `hermes tools`/Settings; showing them
unchecked in every per-profile capabilities editor is noise, and
showing a2a at all confuses users who never enabled agent-to-agent
serving (tester report). Enabled ones still show — hiding an ACTIVE
toolset would misrepresent the profile.
…arch#86243)

The /docs/skills hub page gains an embed mode for host apps that
iframe it as a skill PICKER (first consumer: Bot Mode's agent editor).
With ?embed=picker:
- docs chrome (navbar/footer/hero) is hidden
- every card gains '+ Add to this Agent', which posts
  {type: 'hermes-skill-pick', name, identifier, installCmd, source}
  to the parent window
The page never installs anything — the host validates event.origin
and performs the install through its own gateway (skills.manage).
Normal page rendering is untouched (no query param = no changes).
hermes-seaeye Bot and others added 25 commits August 14, 2026 22:16
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…ault (NousResearch#86414)

* fix(desktop): the main agent's model pick persists as the profile default

Reported: the default bot switches to the OpenAI API account instead of
the user's subscription, and doesn't retain the previous selection.

Root cause: the composer model picker always sent the switch as
--session scope, even for the PRIMARY profile's main agent. So the pick
never wrote config.yaml model.provider — and with model.provider unset,
resolve_provider('auto') falls through to a leftover OPENAI_API_KEY env
var and picks OpenAI/OpenRouter. The subscription the user selected was
only ever a per-session override that evaporated on the next session.

Fix: when the pick targets the primary profile's main agent
(touchesPrimary), send --global so it persists to config.yaml
(model.default + model.provider) via the existing model-switch persist
path. A SET model.provider already outranks the OPENAI_API_KEY env var
in resolve_provider (tier 2 vs tier 3), so the main agent now keeps the
chosen provider across restarts. Secondary chat tiles stay --session so
picking a model in one chat never rewrites the profile default (the
cross-session-contamination guard the old comment protected).

No change to resolve_provider's priority chain, so NousResearch#29285 (an explicit
env key beating a STALE oauth login) is untouched — we simply make the
user's explicit main-agent selection the config default it always
should have been.

* MoA presets stay session-scoped; update tests for primary-persist intent

Fix CI (ui shard 3of3): the primary main-agent pick now persists via
--global, but MoA (mixture-of-agents) presets must NOT — a transient
orchestration choice can't become the global gateway default. Exclude
provider==='moa' from the persist path (stays --session). Update the
primary-picker test to assert --global (the new intent) and keep the
MoA + secondary-tile tests asserting --session (the guards that prove
the narrowing). 19/19 green locally.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
test_run_prompt_submit_requeues_all_unstarted_notifications_with_real_threading
failed twice in one hour on CI slices for two UNRELATED PRs (86371,
86374) with `assert set() == {proc_batch_2, proc_batch_3}`. Root cause:
session.init/create tests earlier in the file start real per-session
notification poller daemon threads and never stop them. Those pollers
outlive their test and keep polling the PROCESS-GLOBAL
process_registry.completion_queue, stealing-and-requeuing the target
test's events mid-assertion so its bounded drain loop can starve.

Reproduced: with 30 leaked foreign-session pollers injected via a
sabotage conftest, the target test fails standalone ~1 in 3 runs with
the exact CI assertion; with the reap fixture active it passed 8/8
under the same sabotage.

Fix:
- tui_gateway/server.py: _start_notification_poller registers
  (stop_event, thread) in module-level _notification_pollers (pruned of
  dead threads on each spawn; threads get a stable
  tui-notif-poller-<sid> name for debugging).
- tests/test_tui_gateway_server.py: autouse fixture sets every
  registered live poller's stop event after each test and joins them
  under ONE shared 3s budget (the poller loop wakes at least every
  0.5s), so no poller survives into the next test. No per-thread
  timeout, no session-dict mutation — a first draft that mutated
  session state and joined per-thread hung the file; full-file runtime
  with this version is 15.2s vs 13.2s baseline.
…ousResearch#86473)

Adds the full MCP setup surface as profile-scoped gateway RPCs so a
desktop client (Bot Mode's bot editor, the core Capabilities tab) can
add/configure/test/authenticate/remove MCP servers for ANY profile, not
just the launch profile:

- mcp.servers.list (profile) -> configured servers (transport, auth,
  oauth_tokens_present, enabled, tool names; no secret values)
- mcp.servers.add (profile, name, config|preset, bearer_token?) -> reuses
  mcp_config._apply_mcp_preset / _save_mcp_server / _save_bearer_auth_token
- mcp.servers.set_api_key (profile, name, value, env_var?) -> http auth
  header template or stdio env ref, via save_env_value
- mcp.servers.test (profile, name) -> _probe_single_server + oauth state
- mcp.servers.remove (profile, name)
- mcp.servers.oauth.start/poll (profile, name[, session_id]) -> mirrors the
  PROVIDER oauth session/poll model (not the FastAPI dashboard flow): a
  background worker drives the same interactive machinery 'hermes mcp login'
  uses, capturing the browser redirect on a local loopback listener. Client
  opens auth_url via openExternal and polls until status=='approved'.

All handlers are profile-scoped via set_hermes_home_override in try/finally
(mirrors skills.manage). Shared helpers live in tui_gateway/mcp_rpc_helpers.py
and are aliased onto server.py's namespace so the rebound handler bodies
(HandlerRegistry.install) can resolve them — a plain def in methods_tools is
unreachable post-rebind. Reuses hermes_cli/mcp_config.py throughout; no config
logic duplicated; no raw yaml near config.yaml (config-read-guard safe).

Tests: tests/tui_gateway/test_mcp_profile_rpcs.py, 8 E2E against real temp
HERMES_HOME profiles asserting add/list/set_api_key/remove land in the RIGHT
profile's config.yaml and not the launch profile's. 8/8. Registration +
live mcp.servers.list verified in an imported gateway.

Co-authored-by: Teknium <teknium1@users.noreply.github.com>
…cess env

The browser-use CLI runs under its own Python (uv tool / uvx), which
can differ from Hermes's venv interpreter. PYTHONPATH/PYTHONHOME
inherited from the agent process point at Hermes's venv
site-packages, and a child interpreter honors them ahead of its own —
so the CLI imported compiled C-extensions (pydantic_core) built for
the wrong interpreter and crashed with ABI mismatch /
ModuleNotFoundError (issues 83427, 84841, 86006, 86104; hits the
desktop backend on py3.14 and any shell exporting PYTHONPATH).

Strip both vars in _base_subprocess_env() — the CLI manages its own
environment and never needs Hermes's import path.

Salvaged from PR 83471 by Benjamin (@n1majne3), the earliest of two
independent fixes (also PR 84022 by @jklance16, PYTHONPATH-only);
regression test covers both vars and preserves unrelated env.
Follow-up on the NousResearch#83854 salvage: prepend $HERMES_HOME/bin ahead of the
venv and user-local bin dirs, matching the managed-first Browser Use
CLI resolution policy — the worker resolves the same canonical binary
the agent process does.
…ly discarding them

Earlier releases accepted api_mode: openai on custom provider entries.
The canonical transport set is now {chat_completions, codex_responses,
anthropic_messages, bedrock_converse, codex_app_server}, and an
unrecognized value was silently ignored at both consumption sites
(_normalize_custom_provider_entry passes the raw string through and
agent_init's accepted-set check drops it; _parse_api_mode returns None),
falling through to hostname-based detection.

For hosts with a detection rule the provider silently switches
transports after an update. Observed live: a custom entry for
api.actual.inc with api_mode: openai (valid when written) flipped to
codex_responses via the hostname rule, and every reasoning-bearing
request to the relay's /v1/responses failed with a wrapped non-JSON
error while /v1/chat/completions worked throughout.

Fix: one shared alias map (_canonical_api_mode) consulted by both
sites. openai/openai_chat -> chat_completions, responses ->
codex_responses, anthropic/messages -> anthropic_messages, bedrock ->
bedrock_converse. Canonical names and unknown values pass through
unchanged, so invalid-config behavior is untouched.

Tests: alias map contract (every alias lands in _VALID_API_MODES),
normalizer canonicalization incl. the transport: key alias, and the
runtime gate accepting legacy spellings while still rejecting unknowns.
…probes

Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (NousResearch#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:

- the requests-based metadata/pricing probe
  (agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
  (hermes_cli/models.py::probe_api_models)

A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.

This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:

- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
  ssl_ca_cert before falling back to the env vars. Callers with no base_url
  keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
  passes it through open_credentialed_url, which gains an ssl_context seam on
  the cloned secure opener. Unmatched or public endpoints pass None and keep
  urllib's default policy.

Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
Upstream added stream=True to the /models metadata probe; widen the
fake_get signature so the captured verify assertion still runs.
…direct guard

ActualProfile.fetch_models() overrides ProviderProfile's default
implementation with its own Actual-specific base_url resolution
(ACTUAL_BASE_URL env var, hosted-vs-local normalization), but called raw
urllib.request.urlopen(req, timeout=timeout) directly instead of the base
class's open_credentialed_url(). Every other provider either uses the
base class default or forwards to it via super() and gets
SafeCredentialRedirectHandler for free — Actual is the only provider that
attaches a Bearer token to its own Request object and opens it with the
stdlib's default redirect handling, which forwards every header,
including Authorization, across a cross-origin redirect.

Actual's own feature surface makes the trigger realistic: ACTUAL_BASE_URL
is a first-class, documented way to point this provider at a self-hosted
or local-offline endpoint (see the local-loopback no-auth path already
handled elsewhere in this provider), so a misconfigured or compromised
endpoint 302-ing to another host leaks ACTUAL_API_KEY to it.

Fix: import and call the same open_credentialed_url() the base class
uses, keeping Actual's own URL-resolution logic unchanged.

Adds an end-to-end regression test using two real local HTTP servers (no
mocking of the security module itself) — one redirects, the other
records the Authorization header it receives — mirroring
test_urllib_security.py's own redirect tests. Also repoints the existing
fetch_models test's mock from urllib.request.urlopen to
hermes_cli.urllib_security.open_credentialed_url, since fetch_models no
longer calls the former. Mutation-verified: the new redirect test fails
on pre-fix code with the Authorization header observed at the redirect
target.
…gration) (NousResearch#86506)

delegation.max_iterations is the per-subagent tool-call budget. The old
default of 50 truncated substantial delegated work: leaf agents spend
~15-20 turns on reconnaissance before producing output, then ran out of
budget mid-task and returned 'completed but unfinished' summaries. 250
gives real delegated work room to finish.

Changes:
- config_defaults.py: delegation.max_iterations 50 -> 250; _config_version 35 -> 36
- tools/delegate_tool.py: DEFAULT_MAX_ITERATIONS fallback 50 -> 250 (kept in
  sync with the shipped default to prevent drift)
- config_migrations.py: _migrate_to_36 lifts configs still pinned at exactly
  the OLD default 50 -> 250 on update, so existing installs inherit the new
  headroom. Any other explicit value (deliberate override) is preserved;
  unset inherits 250 at read time.
- cli-config.yaml.example: doc the new default

The cap is per-child and children run concurrently (max_concurrent_children
default 3), so this raises worst-case fan-out cost; delegation.child_timeout_seconds
(default 0 = off) remains available as a wall-clock guardrail, and users can
still pin a lower max_iterations explicitly.

Verified: migration lifts 50->250, preserves a deliberate 120, leaves unset
untouched (3/3); DEFAULT_CONFIG reads version=36, max_iterations=250, fallback=250.
…turn

The statusbar gauge painted only what the backend reported as measured
occupancy, which a session has none of until a turn runs in this process.
Turning the gauge on mid-conversation, or resuming a chat, therefore showed
nothing until the next message.

Fetch session.context_breakdown as soon as the gauge is on screen instead of
when its popover opens. It is the same read-only estimate the popover already
used (chars/4 over the live prompt, tools and transcript — no provider call),
and it reports the measured figure once the backend has one. The popover
becomes presentational and reads the gauge s merged usage, so the bar and the
panel cannot disagree.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
…hrough bab9a85)

Clean merge - no conflicts. Includes the 6 commits deferred by PR #56's
partial sync: upstream resolved the model-cost-guard contradiction
itself in 16b54e2 (test-side fix, via their PR NousResearch#85970), confirming
the partial-sync call. Replay (git merge-tree) produced zero conflicts
and the merged tree passed the recorded semantic checks: drift grep for
'Hermes Desktop' in apps/desktop/{src,electron} is empty, and all brand
carve-out values are intact (upstream's apps/desktop/package.json edits
- UI-test sharding scripts, get-windows to optionalDependencies -
auto-merged around the branded fields).

CI files touched (1), benign: js-tests.yml adds exact-key node_modules
caching (SHA-pinned actions/cache, no restore-keys) and skips the
npm@12 global install when already on 12.x. No new actions, secrets,
triggers, or permissions. package-lock.json change mirrors the
get-windows optionalDependencies move - no version changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFEJ7TzKwG4tQujNy3CWjT
@uaixo uaixo added the ci-reviewed label Aug 15, 2026 — with Claude
@github-actions

github-actions Bot commented Aug 15, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 1fa00fc — Sync upstream hermes-agent main into NousAI-Assistant (154 c

⚠️ Warnings

OSV vulnerability scan · View job

5 known vulnerabilities found in pinned dependencies.

How to fix:

Review the findings in the Security tab. Update the affected dependencies if a patched version is available.


ℹ️ Info

CI-sensitive file review · View job

PR touches sensitive files, but the ci-reviewed label has been added, approving them.

Sensitive files changed:


debug info

CI timings

CI timings · View report · View job

Wall time 4m28s vs 4m46s (-6.3%). 18 job(s) slower, 22 faster, 2 unchanged.

  • JS & TS checks / ui-tui / check: -44.0s
  • Python tests / Run tests slice 4/12: -37.0s
  • JS & TS checks / apps/desktop / check:test:ui:shard-3of3: -36.0s
  • JS & TS checks / apps/desktop / check:test:desktop:all: -34.0s
  • JS & TS checks / List npm workspaces: -32.0s

@uaixo
uaixo marked this pull request as ready for review August 15, 2026 00:17
@uaixo
uaixo merged commit 5d1c7ea into NousAI-Assistant Aug 15, 2026
58 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.