Skip to content

merge: upstream parity sync — NousResearch/main → fork/main (2026-07-10, 386↑/1216↓) - #255

Merged
Kyzcreig merged 1224 commits into
mainfrom
sync/upstream-2026-07-10
Jul 10, 2026
Merged

Kyzcreig merged 1224 commits into
mainfrom
sync/upstream-2026-07-10

Conversation

@Kyzcreig

Copy link
Copy Markdown
Collaborator

Upstream parity sync — NousResearch/main → fork/main

Catch-up merge of upstream into the long-lived hard fork. Frozen upstream target caf557be.
Fork was 386 ahead / 1216 behind at merge-base 852c9b3cb.

Scale

  • 56 conflict files, 147 hunks (140 both-sides semantic + the mem0 arch-split) — the largest single ingest in the fork's history.
  • All 386 fork lead-commits preserved (LCM/compaction, relay pool lanes, mem0 self-host, restart recovery, denorm session list, etc.).

How it was resolved

Conflict resolution driven by codex under a written spec (docs/sync/RESOLUTION-SPEC-2026-07-10.md), then certified by Apollo (markers-clean, import-smoke, fork-feature survival). Every fleet-critical file resolved by reading both sides' history — no blind side-picks. Architectural ports + semantic reconciliations documented in docs/sync/review/.

STAGE-2 gates: 31 reds → 0

Four were real code regressions the merge introduced (fixed the code, preserving fork behavior):

  1. agent/agent_runtime_helpers.py — stray known_tool_ids.discard() from the set→Counter migration (NameError broke all 16 message-sequence-repair tests).
  2. cron/scheduler.py — one-shot dispatch claim was missing on the fork's tick() path (_process_one_job), only in run_one_job; a self-destructing one-shot could re-fire forever.
  3. tui_gateway/server.py — untrusted client source lost its _sanitize_client_source guard (hostile-label → persisted platform).
  4. gateway/slash_commands.py — /compress chat-only msgs projection clobbered by upstream's full-transcript filter (double-counted tool rows, broke both-axes feedback math).

The rest were stale tests updated to real upstream contracts the merge correctly adopted (FailoverReason.ssl_cert_verification, _notify_session_boundary platform arg, _notification_event_belongs_elsewhere sid arg, list_sessions_rich compact_rows, _SlashWorker profile_home, config-sync persist_override). No fork feature dropped; no test weakened. Full detail: docs/sync/review/stage2-reconciliation-fixes.md.

Result: 2267 passed across all touched subsystems (1 pre-existing order-pollution flake that passes isolated).

CI-gate traps pre-checked locally

  • secret_scan (gitleaks 8.18.4, CI-exact): 543 findings triaged → all upstream test fixtures / Unsloth doc placeholders; surgically allowlisted in .gitleaks.toml + negative-tested (a real injected key still trips).
  • contributor-check: all 271 merge-range authors already auto-skipped or in AUTHOR_MAP — 0 new entries needed.
  • config-migration dry-run: idempotent; fork toolsets messaging/send_message + moa/mixture_of_agents survive (resolve > 0). No dropped-toolset finding.
  • setup.py supply-chain: not touched by the merge / byte-identical to upstream — gate won't fire.

Merge topology

Landed as a real 2-parent merge commit — merge with --merge, NOT squash (squashing destroys the merge-base the next upstream sync needs).

Deploy

🔴 Merging this does NOT deploy to the fleet. Live rollout is a separate gated step (hermes update per host, host-1 ≠ the deploying gateway) — Ace's call, on his timing.

shannonsands and others added 30 commits July 7, 2026 17:17
…earch#60576)

Headless/hosted deploys run the dashboard server without COLORTERM in
the process environment, so chalk inside the PTY-spawned TUI child
downgraded every skin hex color to the xterm 256 palette — the default
skin's bronze banner border (#CD7F32) snapped to palette 173 (#D7875F,
salmon red) and the gold caduceus rendered red/yellow on fresh cloud
instances. Local launches never reproduced it because the operator's
interactive terminal leaks COLORTERM=truecolor into the server env.

xterm.js always renders 24-bit RGB, so the dashboard PTY child should
always advertise truecolor: backfill COLORTERM=truecolor in
_resolve_chat_argv via setdefault (an explicit operator value wins).

Verified with a clean-env PTY probe of the real TUI binary:
no COLORTERM -> 0 truecolor SGRs / 165 palette-256 (salmon 38;5;173);
with the backfill -> 166 truecolor SGRs, exact bronze 38;2;205;127;50.
…atus (NousResearch#60585)

The profile+gateway topology added in NousResearch#60537 sits entirely behind the
loopback/--insecure auth gate. But a hosted agent (Hermes Cloud) binds
non-loopback with OAuth, so should_require_auth is True, and NAS reads
/api/status over the network (fly-provider.ts getInstanceRuntimeStatus)
with no session token. On that gated path the whole topology block was
omitted, so the Portal could never render the profile list.

Split the topology readout by sensitivity:
- profile NAMES (profiles) + gateway_mode are low-sensitivity product
  surface and now ride the always-public status body, surviving the auth
  gate so NAS/the Portal can enumerate profiles.
- the per-gateway detail (gateways[], carrying host ports) is deployment
  recon and stays gated alongside hermes_home / config_path / env_path /
  gateway_pid / gateway_health_url.

The collector now runs unconditionally (still in the executor, off the
event loop). No new fields; only the gate placement changes.
…sResearch#60586)

The multiplex machinery already routes an inbound message to a profile via
SessionSource.profile (build_session_key namespacing + the per-turn
config/credential scope in SessionStore._resolve_profile_for_key). But the
relay path never populated it: _event_from_wire rebuilt the SessionSource
field-by-field and dropped any 'profile' the connector sent, so a
Team-Gateway (connector + relay) message could not be routed to a specific
profile the way the /p/<profile>/ HTTP prefix and per-credential polling
adapters already can.

Stamp source.profile from the wire payload in _event_from_wire. This is the
last missing link for NAS-driven per-profile routing over the relay in
multiplex mode; the connector populating the field ships separately
(gateway-gateway contract adds the optional wire field).

Back-compat: absent 'profile' → None → legacy agent:main namespace,
byte-identical to today for every single-profile gateway.
…p commands

Strict charset allowlist (alnum + - _, max 64) on the {name} path param of
the memory-provider config/setup endpoints. Prevents traversal-shaped names
from reaching find_provider_dir(), and setup now 404s when neither a
loadable provider nor a plugin manifest exists, so the command-running path
is only reachable for discoverable plugins. Adds regression tests.
…flag (NousResearch#60589)

The connector now depends on the single multiplexed gateway for per-profile
relay routing, so hosted deployments need to FORCE multiplexing on regardless
of the image's config.yaml. gateway.multiplex_profiles was config.yaml-only,
which a user could leave unset or flip off.

Add GATEWAY_MULTIPLEX_PROFILES as a standard operator override on top of the
existing config key — the same 'config.yaml is canonical, env is the operator
override' pattern the Telegram/Signal require_mention bridges use:

  env (recognized token) > config.yaml (top-level or nested gateway.*) > False

- gateway/config.py: _env_multiplex_profiles_override() resolves the env var
  tri-state — recognized truthy/falsy token → bool; unset/blank/unrecognized
  → None (fall through to config). Blank is deliberately None, not False, so a
  provisioned-but-unpopulated Fly secret ('') can't shadow a config.yaml opt-in
  (the empty-secret trap). Wired into GatewayConfig.from_dict so every consumer
  (run.py, session.py via self.config) sees the resolved value.
- hermes_cli/gateway.py: the named-profile-start guard
  (_guard_named_profile_under_multiplexer) reads config.yaml directly, so it
  gets the SAME env precedence — otherwise env-forced multiplex would leave the
  guard blind and someone could start a conflicting per-profile gateway that
  double-binds a bot token. Env-forced-on trips the guard even with no
  config.yaml key; env-forced-off disables it over a config opt-in.

Tests: full 3-tier precedence in test_config.py (incl. the discriminating
env-overrides-config cases + the empty/whitespace/unrecognized fall-through
trap + resolver tri-state), mutation-verified (flipping precedence fails
exactly the two env-wins tests); guard env cases in test_multiplex_lifecycle.py.

Force-on is safe on a single-profile instance: session keys stay byte-identical
(agent:main) and the _run_agent wrapper installs the per-turn secret scope, so
the fail-closed get_secret() path is satisfied.
…ashboard chat PTY resume

The chat PTY launch path landed on main after PR NousResearch#50558 and still called
_session_latest_descendant() with the old one-arg signature. Open the
requested profile's state DB (matching the REST endpoint) so profile-scoped
resume resolves descendants in the right database.
NousResearch#60643)

The April 2026 pin to WhiskeySockets/Baileys#01047deb existed only to
pick up the abprops bad-request fix (Baileys PR NousResearch#2473) before it was
released. That fix shipped in v7.0.0-rc11 (May 2026); our pinned commit
is now 48 commits behind rc13.

The git pin forced npm to clone the repo and compile Baileys from
TypeScript source on every fresh install (~3 min), which blew past the
dashboard pairing flow's timeout. Registry install takes ~3s.

Validation: all 9 bridge.js imports present in rc13, bridge.native.test.mjs
passes (13/13), live bridge boot renders pairing QR against real WA servers.
…desktop) (NousResearch#57225)

* feat(install): warn pip/Homebrew installs are unsupported (CLI, TUI, desktop)

pip and Homebrew are now Unsupported install methods per
website/docs/getting-started/platform-support.md. Surface a
warn-don't-block deprecation notice everywhere the install method is
already shown, pointing at the platform-support docs and noting these
installs will not receive further updates. NixOS (Tier 2) is untouched.

- hermes_cli/config.py: shared is_unsupported_install_method() /
  format_unsupported_install_warning() helpers so the wording and docs
  link stay consistent across every surface.
- hermes_cli/banner.py: generalize the existing pip-only banner
  warning to also cover Homebrew.
- hermes_cli/main.py: hermes update and hermes update --check print
  the warning before proceeding (still update; warn, don't block).
- tui_gateway/server.py: session.info gains install_warning.
- ui-tui: SessionPanel renders install_warning alongside the existing
  'N commits behind' notice.
- apps/desktop: SessionRuntimeInfo/GatewayEventPayload gain
  install_warning; applyRuntimeInfo + the live session.info event fire
  a snoozable warning toast via a new reportInstallMethodWarning(),
  mirroring the existing backend-contract-skew toast pattern. i18n
  strings added for en/zh/zh-hant/ja.
- Tests: updated pip banner assertions for the new wording, added a
  Homebrew banner test, and two tui_gateway session_info tests
  (install_warning present for pip, absent for git).

* fix(nix): make `hermes` in developement environment actually work

install modules as editable overlay with uv

* feat: print install method when running --version

* fix: correct detect install method when running from a subtree
…reporting

write_file() previously called _atomic_write() first and only ran the
JSON/YAML/TOML/Python syntax check afterward as an informational lint
delta -- a parse failure never set the top-level `error` key, so a
corrupt structured-data write still landed on disk (and file_tools.py's
files_modified gating, which keys off `error`, silently reported it as
a successful modification).

Move the in-process syntax check for JSON/YAML/TOML ahead of
_atomic_write() and refuse the write outright on a parse failure: no
temp file, no rename, nothing touches disk, and the result carries a
top-level `error` so callers correctly see it as unmodified.

Deliberately scoped to _FAIL_CLOSED_INPROC_EXTS (JSON/YAML/TOML), not
all of LINTERS_INPROC -- .py is excluded because this codebase's own
test fixtures (TestPatchReplacePostWriteVerification et al.) write
arbitrary non-Python text through *.py paths purely to exercise
write-mechanics; a hard block there broke 3 previously-passing tests
during development. Python keeps its pre-existing non-blocking
lint-delta report.

Adds tests/tools/test_write_file_syntax_gate.py: invalid JSON/YAML/YML/
TOML refused with nothing written (new file) and nothing modified
(existing file); valid JSON/YAML still written byte-for-byte; a
non-linted extension with garbage content is unaffected; invalid Python
is confirmed NOT hard-refused (still just reported).
…YAML isn't refused

safe_load() raises ComposerError on multi-document streams (k8s manifests)
and ConstructorError on application-defined tags (CloudFormation !Sub,
Ansible !vault) — both valid YAML syntax. Now that the linter's verdict is
a fail-closed write gate, those false positives would refuse legitimate
writes outright. Switch to yaml.parse() (scanner+parser only), which still
catches real syntax failures.
/update and other shutdown paths only waited on gateway session agents,
so active cron tool work was killed immediately in final-cleanup while
the scheduler could still mark the job successful (NousResearch#60432).
Cron jobs run through cron/scheduler.py's own ThreadPoolExecutor via a
standalone AIAgent (run_job/run_one_job), entirely outside
GatewayRunner._running_agents -- the dict _drain_active_agents() and
every other active-work check on that class reads. A gateway shutdown
(/update, /restart, and SIGUSR1 all funnel through the same stop())
could log active_at_start=0 and immediately kill tool subprocesses
while a cron job's terminal command was still running, with no wait
and no indication anything was interrupted.

Real-world impact (from the issue): a scheduled daily briefing cron
job was in flight during /update, its tool subprocess got killed
by the unconditional shutdown cleanup, and the job was never marked
failed -- it simply never completed or delivered, with no error
surfaced anywhere. A repro with a 30-minute `sleep` cron job in flight
during /update reproduced the same pattern: subprocess killed at
+0.22s of drain (active_at_start=0), the job's agent thread continued
in-process and produced a plausible-looking final response from the
truncated tool output, and the scheduler marked the run successful.

Root cause is layered, not a single line:

1. GatewayRunner._drain_active_agents() only waits on _running_agents.
   Cron work was invisible to it, so drain returned instantly whenever
   the only active work was a cron job.
2. Even with visibility, the shutdown's final tool-subprocess kill
   (process_registry.kill_all()) is a global, unconditional sweep with
   no per-job targeting -- a long-running cron job that outlives the
   drain timeout still gets its subprocess killed.
3. cron/scheduler.py had no way to detect that a job's tool subprocess
   was killed out from under it mid-run; the agent thread kept going
   and its eventual (often degraded but plausible-looking) response
   got reported as a normal successful completion.

Fix, three parts:

- cron/scheduler.py: expose get_running_job_ids() (thread-safe
  snapshot of the existing _running_job_ids set, already used to
  prevent double-dispatch) so the gateway can read cron's in-flight
  state without reaching into private module internals.

- gateway/run.py: GatewayRunner._active_cron_job_count() reads that
  snapshot. _drain_active_agents() now waits on
  (_running_agents OR active cron jobs), so a cron-only workload gets
  the same bounded wait chat sessions already get instead of an
  instant active_at_start=0. Shutdown drain logging gains
  cron_active_at_start/cron_active_now fields alongside the existing
  ones (unchanged, for compat).

- cron/scheduler.py: mark_running_jobs_interrupted(reason), called by
  gateway/run.py's _kill_tool_subprocesses() right after
  process_registry.kill_all(), marks every job still in
  _running_job_ids at that instant as failed/interrupted via the
  existing mark_job_run() -- and records the job IDs in
  _interrupted_job_ids BEFORE writing, so run_one_job()'s own
  eventual completion for the same run (racing in its own thread)
  checks that flag and skips its normal write instead of clobbering
  the interrupted status with a false "ok" produced from the
  now-truncated tool output. This does not attempt to correlate a
  killed PID to a specific job ID (process_registry tracks PIDs, not
  job IDs) -- any job still dispatched at the moment of a forced kill
  is treated as interrupted, matching the existing coarser precedent
  set by _interrupt_running_agents(), which interrupts every entry in
  _running_agents on a drain timeout without per-agent correlation
  either.

Deliberately out of scope (flagged in the issue as a separate,
lower-priority concern): startup-time reconciliation of cron runs that
started but never reached a terminal status.

Testing:

- tests/cron/test_shutdown_interrupt.py (12 tests): get_running_job_ids
  snapshot semantics, mark_running_jobs_interrupted marking/no-op/
  partial-failure behavior, and -- the core race guard -- run_one_job
  skipping its own last_status write (both the success path and the
  exception path) when the shutdown path already marked the run
  interrupted, with a control test proving ordinary un-interrupted
  completions are unaffected.

- tests/gateway/test_cron_active_work_drain.py (9 tests):
  _active_cron_job_count reading cron state and failing closed (0) if
  the cron module is unavailable; _drain_active_agents waiting for an
  in-flight cron job the same way it waits for chat sessions, timing
  out if the job outruns the window, and leaving existing chat-session
  drain behavior unchanged; a full runner.stop() integration test
  (drain-timeout path) proving mark_running_jobs_interrupted actually
  fires with the right job ID when a tool subprocess is force-killed,
  plus a no-op control when nothing cron-related is in flight.

- tests/gateway/test_shutdown_cache_cleanup.py: added
  _active_cron_job_count() to that file's hand-rolled _FakeGateway test
  double, which stop() now calls -- without it those 8 pre-existing
  tests AttributeError (caught by fail-then-pass below, not a
  production bug).

Fail-then-pass: reverted gateway/run.py + cron/scheduler.py, all 21
new tests fail (fixture/attribute errors -- the feature doesn't exist
yet); restored, all 21 pass.

Regression check: ran the full plausibly-affected surface --
tests/gateway/{test_gateway_shutdown,test_restart_drain,
test_restart_notification,test_restart_redelivery_dedup,
test_restart_resume_pending,test_restart_service_detection,
test_shutdown_cache_cleanup,test_stuck_loop,test_clean_shutdown_marker,
test_external_drain_control,test_session_state_cleanup,
test_update_command,test_update_streaming}.py plus tests/cron/ (944
tests) -- against a clean upstream/main checkout and against this
branch. Diffed the two FAILED lists: identical, 20 pre-existing
failures on both sides (Windows-locale/cp1252 file-encoding issues and
Unix-permission-bit assertions that don't apply on this Windows dev
box), zero new failures, zero fixed-by-accident. The 8
test_shutdown_cache_cleanup.py failures found mid-development were
from the _FakeGateway gap above, fixed in the same commit and
confirmed clean on the final rerun (diff against baseline: exit 0).

Fixes NousResearch#60432
Follow-up to the previous commit on NousResearch#60432. The status-write guard
(_consume_interrupted_flag, checked right before mark_job_run) closes
the false-success bookkeeping gap, but run_one_job delivers its result
BEFORE that check: delivery happens right after run_job() returns,
mark_job_run happens at the very end. A job whose tool subprocess was
killed mid-flight can still produce a plausible-looking final_response
from the truncated output, and that response would reach the user via
_deliver_result before the interrupted flag was ever consulted --
correct status in jobs.json, wrong message already sent.

Adds _is_interrupted(), a non-destructive peek at the same
_interrupted_job_ids set (_consume_interrupted_flag stays as the
consuming, authoritative check right before the status write -- this
needed a peek instead since the flag has to still be visible there).
Checked right after save_job_output, before the deliver_content
decision: if the run looked successful but was flagged interrupted,
force success=False with an explicit interruption message. This
routes delivery through the existing _summarize_cron_failure_for_delivery
path (the same one a real failure already uses) instead of the raw
final_response, so the user gets an honest "this run was interrupted"
instead of a truncated/misleading result.

Testing: 4 new tests in tests/cron/test_shutdown_interrupt.py --
_is_interrupted peek semantics (false/true/does-not-clear, as opposed
to the consuming _consume_interrupted_flag), and the delivery-gate
test itself, which mocks run_job to return a normal-looking success
with a "plausible final response" while the job is pre-marked
interrupted, and asserts _deliver_result receives the failure summary
("This run was interrupted.") instead, with the summarizer's error
argument confirmed to mention the interruption.

Fail-then-pass: reverted cron/scheduler.py only, the 4 new tests fail
(3 on the missing _is_interrupted attribute, 1 -- the delivery-gate
test -- on _summarize_cron_failure_for_delivery never being called,
i.e. the raw response would have gone out); restored, all 16 tests in
the file pass.

Regression: tests/cron/ (683 tests) + test_cron_active_work_drain.py +
test_gateway_shutdown.py + test_shutdown_cache_cleanup.py -- 11
pre-existing failures (Unix file-permission-bit and path-tilde
assertions that don't apply on this Windows dev box), matching the
same set already established as pre-existing in the prior commit's
regression check. Zero new failures.

Continues NousResearch#60432
…onto one drain surface

Keep NousResearch#60631's get_running_job_ids() snapshot + _active_cron_job_count()
(import-guarded for minimal test doubles) as the single read path, and
retarget NousResearch#60612's drain tests at it. Drops the redundant
cron_jobs_in_flight() helper so there is one surface, not two.
Guard _finalize_session's db.end_session() call against gateway-owned
sessions (telegram, bluebubbles, discord, etc.).  The TUI is a viewer
for these sessions, not the lifecycle owner.  Unconditionally ending
them in state.db creates a Groundhog Day routing loop: the gateway's
NousResearch#54878 self-heal detects the stale entry, recovers to the parent
session, context compression splits back to the reaped child, and the
cycle repeats on every inbound message — causing complete conversational
context amnesia.

Fixes NousResearch#60609
…hardcoded list

The salvaged guard used a hand-maintained frozenset of 14 platform names —
several of which (line, wechat, facebook, imessage, googlechat) aren't
actual Hermes Platform values, while real ones (whatsapp_cloud, feishu,
wecom, dingtalk, qqbot, yuanbao, plugin platforms like irc) were missing.
Resolve the source through gateway.config.Platform instead (built-ins +
registered plugin platforms via _missing_), with an explicit exclusion set
for self-owned/local sources. Adds tests for the guard and both reap paths.
…S-free) (NousResearch#60730)

For air-gapped / self-hosted-IdP deploys with NO Nous Portal, let the gateway
obtain its caller-identity bearer from a generic OAuth2 client_credentials grant
against the operator's own IdP (e.g. Microsoft Entra ID) instead of only
resolve_nous_access_token(). The connector's OIDC tenant resolver reads a claim
(default tid) off that token as the tenant.

- gateway/relay: new canonical _resolve_relay_identity_token() — client_credentials
  when gateway.idp.token_url (or GATEWAY_RELAY_IDP_* env) is set, else Nous Portal
  (unchanged default). Wired into self_provision_relay().
- hermes_cli/gateway_enroll: _resolve_identity_token() delegates to the canonical
  resolver so the enroll CLI and the runtime self-provision path share ONE impl.

Config via gateway.idp.{token_url,client_id,client_secret,scope} in config.yaml
(env override GATEWAY_RELAY_IDP_*). No behaviour change when unset.

Tests: tests/gateway/relay/test_identity_token_resolver.py (6 — mode selection,
request shape, config/env precedence, fail-closed). Relay suite 162 pass.

Validated via the cross-repo gateway<->connector live E2E (provision, managed
self-provision, inbound round-trip, /link) against a connector running the OIDC
tenant resolver with zero NAS config.
_save_provider_state() sets auth_store['active_provider'] as a side effect.
The Z.AI endpoint probe runs from credential-pool env seeding for any user
with a Z.AI key in env — persisting the probe cache must not silently make
zai the active provider. Use _store_provider_state(set_active=False).

Follow-up to PR NousResearch#41201 salvage.
Review findings (hermes-pr-review Phase 2, 3-angle):
- _save_auth_store() does real filesystem I/O (mkdir, O_EXCL create, fsync,
  atomic replace) and can raise on disk-full/permissions/lock-timeout. The
  persist ran bare in the success path, so a persist failure aborted
  _resolve_zai_base_url() after detection had already succeeded. Wrap the
  persist in try/except: log a warning and still return the detected URL
  (worst case: next start re-probes).
- Readability: stage the payload in a local detected_endpoint instead of
  writing through the stale pre-lock 'state' dict, which is no longer what
  gets persisted.
OutThisLife and others added 18 commits July 10, 2026 02:13
Windows drops webContents zoom on minimize/restore, so the UI snapped back
to 100% while Settings still read 125%. Zoom was only reasserted on the main
window's did-finish-load, never on show/restore and never for session windows.

Reassert the persisted level on show/restore + first load, wired once in
wireCommonWindowHandlers so the main window and secondary session windows
share it. The pet overlay opts out (zoom:false): it sizes its own OS window
to fit the sprite in unzoomed CSS px and has its own Alt+wheel scale, so
inheriting the global zoom would render the mascot larger than its window and
crop it (and it shares the renderer origin's zoom localStorage key).

Salvages NousResearch#61245; keeps its pure-helper tests and adds a scope assertion.

Co-authored-by: HexLab98 <liruixinch@outlook.com>
…245-ui-zoom

fix(desktop): re-apply UI zoom on show/restore, scoped to chat windows (supersedes NousResearch#61245)
Dashboard Chat is an xterm mirror of a TUI inside the gateway, so
server-side clipboard.paste never sees the browser clipboard. Upload
pasted/dropped images to the profile's images/ dir (same place
clipboard.paste / image.attach use), then drive /image over the PTY.

Uses a dedicated /api/chat/image-upload endpoint (magic-byte check,
25MB cap, profile scope) instead of relative managed-files uploads that
400 on local dashboards without a locked root. Ctrl/Cmd+Shift+V also
tries clipboard.read() for images before falling back to text, since
preventDefault on that chord suppresses the DOM paste event.

Salvages NousResearch#57912 (client composition + /image PTY drive) and folds in
NousResearch#48563's upload endpoint + drop path.

Co-authored-by: bird <6666242+bird@users.noreply.github.com>
Co-authored-by: tt-a1i <53142663+tt-a1i@users.noreply.github.com>
…912-dash-paste

fix(web): paste/drop images into dashboard Chat via HERMES_HOME/images
Old packaged Desktop apps re-entering bootstrap against an existing
~/.hermes/hermes-agent were still passing the baked-in --commit pin, which
detached the managed checkout back to the app stamp (e.g. 0.15.1) after
hermes update had already moved it forward.

Skip the packaged commit pin when activeRoot already has git metadata;
keep branch args and fresh-install commit pinning unchanged. Port of
NousResearch#59902 onto the post-ts-ify bootstrap-runner.ts.

Co-authored-by: helix4u <4317663+helix4u@users.noreply.github.com>
…902-bootstrap-repin

fix(desktop): prevent bootstrap stale commit repin on existing checkouts
Switching connection mode (local / cloud agent / remote) no longer
full-window-reloads into the cold-boot CONNECTING screen. The primary
backend is torn down in place (no renderer reload); the shell + Settings
stay up while session lists are wiped so sidebar skeletons retrigger, then
the socket re-dials and config/sessions refresh. Cold-boot CONNECTING
latches off after the first successful boot; the intentional teardown
suppresses the backend-exit toast. Dev affordance: a "Preview soft switch"
button under Gateway diagnostics (Electron has no ?query= entry).

Gateway settings UI brought in line with the rest of Settings:
- Mode cards use the shared selectableCardClass on an equal-height
  auto-rows-fr grid, stacking 1→3 (never an orphaned 2+1); titles wrap
  instead of truncating.
- Remote gateway's auth detail moves into a ? tooltip in the title; drop
  the redundant "connects to the one you choose" from the cloud card.
- textStrong buttons force px-0 so the underline sits flush with the label.
- Tooltip chip uses box-decoration-break: clone so the background hugs each
  wrapped line (bg only on the text), capped at max-w-64.

Fully i18n'd (en + zh; ja/zh-hant inherit via defineLocale).
…402-desktop-cloud-mode

feat(desktop): Hermes Cloud connection mode (salvage of NousResearch#55402)
…teway-switch-ux

feat(desktop): soft gateway switch + gateway-settings polish
Radix's hoverable-content grace area can leave tips stuck over Electron drag regions; disable it and make tip content pointer-events-none so open state tracks the trigger only.
…p-stuck

fix(desktop): stop Tip from sticking open and blocking clicks
Catch-up merge of upstream (frozen target caf557b) into the long-lived hard fork.
Fork was 386 ahead / 1216 behind at merge-base 852c9b3. 56 conflict files,
147 hunks (140 both-sides semantic + mem0 arch-split). All 386 fork lead-commits
preserved; import-smoke + marker-clean verified before commit.

Architectural ports & semantic reconciliations (full detail in docs/sync/review/):
- plugins/memory/mem0: take-fork wholesale; removed upstream _backend/_setup refactor
  shards + their tests (fork __init__ is the production superset). Upstream setup-wizard
  refactor noted for deliberate follow-up, not reintroduced as orphan modules.
- agent/auxiliary_client _resolve_task_provider_model: reconciled api_mode — first-class/
  custom/auto keep upstream api_mode=None (downstream transport detection); plugin
  providers keep the fork eager ProviderProfile.api_mode read.
- agent/chat_completion_helpers: keep-both — fork relay/fallback route announce + audit
  sink + x-hermes-lane pool headers preserved; added upstream unavailable-fallback skip.
- tools/delegate_tool: fork inherit_context/boomerang kept; code_execution stays exempt
  from subagent toolset stripping; took upstream ACP transport-field removal.
- gateway/run.py, gateway/slash_commands.py: interleave — fork restart-loop/recovery gates,
  compaction telemetry, branch-thread, usage cards preserved; upstream async session-store
  calls used where hunk-local.
- gateway/session.py: keep-both — fork durable reasoning/model identity + answerable
  lookup; upstream model_override + rewrite_transcript success-return (clear undo only
  after successful rewrite).
- hermes_state.py: keep-both — fork gated effective_last_active denorm + backfill; upstream
  MAX_FTS5_QUERY_CHARS, handoff index, search_query, compact_rows (denorm fast-path bypassed
  when upstream-only search/compact options requested).
- cron/scheduler.py: keep-both — fork per-job reasoning/fallback + loud alert; upstream
  dispatch-claiming, credential-exfil guard, TG DM-topic probe, deferred teardown.
- tui_gateway/server.py: reconciled fork client source attribution with upstream
  platform_override; _make_agent accepts both, disabled reasoning reports "none".
- run_agent.py: interleaved upstream intrinsic-persistence markers with fork interrupt-close
  finish-reason re-persist; row ids attached to persisted dicts (resume discriminator no
  longer depends on long-lived id() sets).
- agent/model_metadata.py: keep-both — fork two-tier token breakdown + request telemetry;
  upstream async context lookup + local-probe cache + bounded tool-token cache.
- Desktop: upstream Electron TS migration/package shape; fork app-side media/project/composer
  behavior preserved.

Fork-critical features re-verified present post-merge: relay x-hermes-lane headers,
messaging/moa toolsets + send_message/mixture_of_agents tools, code_execution exemption.

Resolution by codex under docs/sync/RESOLUTION-SPEC-2026-07-10.md; certified (markers,
import-smoke, feature-survival) by Apollo.

POST-MERGE RECONCILIATION FIXES (STAGE-2 gates: 31 reds → 0; detail in
docs/sync/review/stage2-reconciliation-fixes.md):
- agent/agent_runtime_helpers.py: removed a stray set-era `known_tool_ids.discard()` the
  set→Counter merge left behind (NameError broke all 16 message-sequence-repair tests).
- cron/scheduler.py: added the upstream one-shot dispatch claim to _process_one_job (the
  fork's tick path), not just run_one_job — a self-destructing one-shot could otherwise
  re-fire forever on the built-in ticker.
- tui_gateway/server.py: restored the fork's _sanitize_client_source guard on the untrusted
  client `source` in _make_agent (the merge trusted it verbatim → hostile-label injection).
- gateway/slash_commands.py: restored the fork's chat-only /compress `msgs` projection (the
  merge took upstream's full-transcript filter → double-counted tool rows, broke both-axes
  compress-feedback math).
- tests updated to real upstream contracts the merge adopted: FailoverReason.ssl_cert_verification;
  _notify_session_boundary platform arg; _notification_event_belongs_elsewhere sid arg;
  list_sessions_rich compact_rows; _SlashWorker profile_home; config-sync persist_override;
  _resolve_session_agent_runtime override stubs. No fork feature dropped, no test weakened.

STAGE-2 pytest gates green (2267 passed; 1 pre-existing order-pollution flake that passes
isolated). CI-gate traps + config-migration dry-run follow before deploy (deploy is a
separate gated step, Ace's call).
@greptile-apps

greptile-apps Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

Too many files changed for review. (1409 files found, 100 file limit)

Kyzcreig added 7 commits July 10, 2026 04:42
…t-up-to-date

#253 added TestCompactRows::test_compact_projection_tracks_schema (compact rows must
carry every schema column). The parity merge's fork denorm column effective_last_active
is deliberately stripped from list rows (consumers use the computed last_active; the raw
value is internal to the recency-ordering CTEs, and the fork denorm-reland oracle tests
assert its absence). Reconciled by adding effective_last_active to the test's sanctioned
exclusion set alongside system_prompt — preserving the fork's clean-list-row contract
without weakening #253's schema-tracking guard.
TypeScript CI on the parity PR caught three desktop merge-resolution gaps where the
merge kept the fork side of a conflicted file but consumers/tests reference upstream
symbols the fork lacks:

- apps/desktop/src/lib/media.ts: re-add upstream's downloadGatewayMediaFile export +
  its readDesktopFileDataUrl import (consumed by artifacts/index.tsx + markdown-text.tsx).
- use-composer-actions.test.ts: the merge kept the fork's readComposerImagePreview impl
  but unioned upstream's attachmentPreviewDataUrl test block (a function the fork doesn't
  export). Removed the orphaned describe + its unused $connection/attachmentPreviewDataUrl
  imports; the routing behavior is covered by desktop-fs.test.ts.
- use-prompt-actions/index.test.tsx: add the upstream-added required PromptActionsOptions
  field openMemoryGraph (real, used by the impl + desktop-controller) to the test options.

Verified each symbol against upstream caf557b. No behavior change to shipping code beyond
restoring the dropped upstream export.
Five more merge-reconciliation gaps the full Python CI suite caught (all passed on
fork/main; regressions the merge introduced by taking one side of a conflicted seam):

- agent/conversation_loop.py preflight compression: the merge kept upstream's
  'Pre-API pressure check' block, which calls the compressor's REMOVED
  should_defer_preflight_to_real_usage ratchet (fork replaced it with skew
  CALIBRATION). The getattr-lambda fallback silently returned False → compressed
  when it should defer. Switched the gate to the fork's should_compress_calibrated.
- agent/conversation_loop.py MoA aggregator-cost block (upstream-only): gated its
  last_aggregator_slot read on isinstance(dict), so a non-MoA/MagicMock client can't
  route a mock provider/model into is_notional_anthropic_provider (TypeError).
- gateway/run.py auto-resume: reconciled upstream's global restart-loop guard
  (restart_loop_guard.check_and_record, coarse: skips ALL auto-resume) with the fork's
  PRIMARY per-session F2 replay-mark→SUSPEND breaker. The global guard fired first and
  returned 0 before any session resumed, so _resumed_this_boot never populated and F2
  could never suspend the looping session. Now the global guard still records every boot
  but only short-circuits when the per-session breaker is disabled; F2 owns the break.
- gateway/slash_commands.py: converted 10 raw self.session_store.X() calls in async
  slash-command methods (redo/compress/merge/branch) to await async_session_store.X(),
  satisfying the upstream async-session-store AST guard (no sync SQLite on the event loop).
- tests/gateway/test_compression_session_id_persistence.py: the fork's hand-rolled AST
  walker missed the session_entry.session_id assignments after upstream restructured
  _handle_message_with_agent; rewrote it with a robust block-recursion (and it now
  recognizes the async await ...async_session_store._save()). Verified non-vacuous: it
  finds both assignments and confirms each is followed by _save().
…ssion + test drift

SECURITY FIX (real regression):
- tools/approval.py _run_approval_gate: the upstream gate-refactor read the raw
  env_var_enabled('HERMES_CRON_SESSION') to detect cron context. A cron-SPAWNED
  subagent runs in a bare ThreadPoolExecutor worker that does NOT inherit the
  os.environ marker — its cron identity lives only in the re-bound ContextVar
  (_bind_child_cron_session → set_cron_session). So a cron child AUTO-APPROVED
  dangerous commands (curl|bash), defeating cron_mode: deny. Switched to the
  context-aware _is_cron_session() (ContextVar-first, os.environ fallback) and
  removed the now-dead duplicate cron-gate block the merge left unreachable after
  check_dangerous_command's return _run_approval_gate(...).

TEST reconciliations to real upstream contracts the merge adopted:
- test_run_agent_tool_exec.py: (a) 13 pre-tool-block patches retargeted from the
  fork's deprecated get_pre_tool_call_block_message shim to upstream's centralized
  resolve_pre_tool_block (which every dispatch path now uses); (b) MCP tool names
  updated from the fork's single-underscore mcp_x to upstream's mcp__server__tool
  convention (MCP_TOOL_NAME_PREFIX='mcp__'); (c) malformed-JSON-args test updated
  to upstream's STRICTER behavior — malformed args are now rejected with an error
  tool result instead of the fork's lenient coerce-to-{} + run (BEHAVIOR CHANGE,
  documented inline: safer — never execute a tool with fabricated empty args).
- test_web_server_sessiondb_eventloop.py: fake _DB now accepts SessionDB(read_only=)
  + list_sessions_rich(compact_rows=) and DEFAULT_DB_PATH.exists() is stubbed True —
  upstream added a read-only open + fresh-install exists() guard + compact_rows.

All targeted suites green: test_run_agent_tool_exec (61), test_cron_subagent_session (5),
test_approval (303), test_web_server_sessiondb_eventloop (9).
Fixes the last 14 test failures surfaced by the full CI matrix after the
upstream parity merge. Each is a fork/upstream reconciliation the conflict
resolution took one side of; classified against the fork/main baseline
(pass-on-fork => my regression, fix code; fail-on-fork => stale test, update).

Real regressions fixed in code (all passed on fork/main, broke in the merge):
- gateway hygiene announce stranded in the no-op else branch: a *rotated*
  hygiene compaction silently sent no in-chat "Context compacted" announce
  on the live fleet. Restored the fork's "announce on any successful
  compaction (rotate OR in-place)" structure.
- run_agent flush: the persist-user-message override lost the platform-message-id
  path (upstream's inline override handled content/timestamp only), breaking
  backfill-on-reconnect dedup (has_platform_message_id). Re-threaded _row_platform_id.
- telegram adapter _handle_polling_network_error read instance-only retry
  attrs; getattr-fallback so the bare-adapter recovery path works (NousResearch#59614).
- length-continuation + truncated-tool-call retry caps drifted to 4 while the
  rest of the retry subsystem (codex/invalid-tool/invalid-json/empty-content)
  stayed at 3 and a comment still said 3 — restored the fork's coherent 3.
- /usage: upstream Context:/Compressions: header lines rendered above the
  context breakdown, violating the fork's "breakdown FIRST" contract. Removed.
- browser_connect: a log string embedded the --remote-debugging-port literal
  that trips the keychain-guard launcher scanner (false positive); reworded.

Stale tests updated to the merged (upstream) contract:
- cron approve/deny approval tests patched env_var_enabled; the security fix
  routes cron detection through _is_cron_session() (ContextVar) now.
- compaction legacy test: small-ctx floor now applies on update_model too.
- todo hydrate: requires a paired assistant todo tool call (GHSA-5g4g-6jrg-mw3g).
- persist voice-input: override applies to the DB row only, never the live
  dict (NousResearch#48677); assert the live content is preserved.
- telegram restart-parity guard: dataflow-resolve the _start_polling_resilient
  wrapper's forwarded drop_pending_updates param (keeps the guard's teeth on
  the real recovery ladders).

Full affected-subsystem sweep green; no new failures introduced.
…ice 3)

The merge took upstream's always-TEMPFAIL-75 service-restart exit code, but
Ace's fleet uses the fork's deployment model: exit 0 under systemd
(Restart=always relaunches on clean exit; 75 would trip stepped-restart
backoff) and 75 only on darwin/launchd (KeepAlive.SuccessfulExit=false needs
non-zero). Branch on INVOCATION_ID + platform. Skipped on macOS; the
exit==0 path runs on Linux CI (test_gateway_stop_systemd_service_restart_exits_cleanly)
and the 75 path is covered by the paired launchd test.
@Kyzcreig
Kyzcreig merged commit 98b40ec into main Jul 10, 2026
40 checks passed
@Kyzcreig
Kyzcreig deleted the sync/upstream-2026-07-10 branch July 10, 2026 13:42
@Kyzcreig
Kyzcreig restored the sync/upstream-2026-07-10 branch September 21, 2026 10:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.