Repository navigation
merge: upstream parity sync — NousResearch/main → fork/main (2026-07-10, 386↑/1216↓) - #255
Merged
Merged
Conversation
…earch#60576) Headless/hosted deploys run the dashboard server without COLORTERM in the process environment, so chalk inside the PTY-spawned TUI child downgraded every skin hex color to the xterm 256 palette — the default skin's bronze banner border (#CD7F32) snapped to palette 173 (#D7875F, salmon red) and the gold caduceus rendered red/yellow on fresh cloud instances. Local launches never reproduced it because the operator's interactive terminal leaks COLORTERM=truecolor into the server env. xterm.js always renders 24-bit RGB, so the dashboard PTY child should always advertise truecolor: backfill COLORTERM=truecolor in _resolve_chat_argv via setdefault (an explicit operator value wins). Verified with a clean-env PTY probe of the real TUI binary: no COLORTERM -> 0 truecolor SGRs / 165 palette-256 (salmon 38;5;173); with the backfill -> 166 truecolor SGRs, exact bronze 38;2;205;127;50.
…atus (NousResearch#60585) The profile+gateway topology added in NousResearch#60537 sits entirely behind the loopback/--insecure auth gate. But a hosted agent (Hermes Cloud) binds non-loopback with OAuth, so should_require_auth is True, and NAS reads /api/status over the network (fly-provider.ts getInstanceRuntimeStatus) with no session token. On that gated path the whole topology block was omitted, so the Portal could never render the profile list. Split the topology readout by sensitivity: - profile NAMES (profiles) + gateway_mode are low-sensitivity product surface and now ride the always-public status body, surviving the auth gate so NAS/the Portal can enumerate profiles. - the per-gateway detail (gateways[], carrying host ports) is deployment recon and stays gated alongside hermes_home / config_path / env_path / gateway_pid / gateway_health_url. The collector now runs unconditionally (still in the executor, off the event loop). No new fields; only the gate placement changes.
…sResearch#60586) The multiplex machinery already routes an inbound message to a profile via SessionSource.profile (build_session_key namespacing + the per-turn config/credential scope in SessionStore._resolve_profile_for_key). But the relay path never populated it: _event_from_wire rebuilt the SessionSource field-by-field and dropped any 'profile' the connector sent, so a Team-Gateway (connector + relay) message could not be routed to a specific profile the way the /p/<profile>/ HTTP prefix and per-credential polling adapters already can. Stamp source.profile from the wire payload in _event_from_wire. This is the last missing link for NAS-driven per-profile routing over the relay in multiplex mode; the connector populating the field ships separately (gateway-gateway contract adds the optional wire field). Back-compat: absent 'profile' → None → legacy agent:main namespace, byte-identical to today for every single-profile gateway.
…p commands
Strict charset allowlist (alnum + - _, max 64) on the {name} path param of
the memory-provider config/setup endpoints. Prevents traversal-shaped names
from reaching find_provider_dir(), and setup now 404s when neither a
loadable provider nor a plugin manifest exists, so the command-running path
is only reachable for discoverable plugins. Adds regression tests.
…flag (NousResearch#60589) The connector now depends on the single multiplexed gateway for per-profile relay routing, so hosted deployments need to FORCE multiplexing on regardless of the image's config.yaml. gateway.multiplex_profiles was config.yaml-only, which a user could leave unset or flip off. Add GATEWAY_MULTIPLEX_PROFILES as a standard operator override on top of the existing config key — the same 'config.yaml is canonical, env is the operator override' pattern the Telegram/Signal require_mention bridges use: env (recognized token) > config.yaml (top-level or nested gateway.*) > False - gateway/config.py: _env_multiplex_profiles_override() resolves the env var tri-state — recognized truthy/falsy token → bool; unset/blank/unrecognized → None (fall through to config). Blank is deliberately None, not False, so a provisioned-but-unpopulated Fly secret ('') can't shadow a config.yaml opt-in (the empty-secret trap). Wired into GatewayConfig.from_dict so every consumer (run.py, session.py via self.config) sees the resolved value. - hermes_cli/gateway.py: the named-profile-start guard (_guard_named_profile_under_multiplexer) reads config.yaml directly, so it gets the SAME env precedence — otherwise env-forced multiplex would leave the guard blind and someone could start a conflicting per-profile gateway that double-binds a bot token. Env-forced-on trips the guard even with no config.yaml key; env-forced-off disables it over a config opt-in. Tests: full 3-tier precedence in test_config.py (incl. the discriminating env-overrides-config cases + the empty/whitespace/unrecognized fall-through trap + resolver tri-state), mutation-verified (flipping precedence fails exactly the two env-wins tests); guard env cases in test_multiplex_lifecycle.py. Force-on is safe on a single-profile instance: session keys stay byte-identical (agent:main) and the _run_agent wrapper installs the per-turn secret scope, so the fail-closed get_secret() path is satisfied.
…ashboard chat PTY resume The chat PTY launch path landed on main after PR NousResearch#50558 and still called _session_latest_descendant() with the old one-arg signature. Open the requested profile's state DB (matching the REST endpoint) so profile-scoped resume resolves descendants in the right database.
NousResearch#60643) The April 2026 pin to WhiskeySockets/Baileys#01047deb existed only to pick up the abprops bad-request fix (Baileys PR NousResearch#2473) before it was released. That fix shipped in v7.0.0-rc11 (May 2026); our pinned commit is now 48 commits behind rc13. The git pin forced npm to clone the repo and compile Baileys from TypeScript source on every fresh install (~3 min), which blew past the dashboard pairing flow's timeout. Registry install takes ~3s. Validation: all 9 bridge.js imports present in rc13, bridge.native.test.mjs passes (13/13), live bridge boot renders pairing QR against real WA servers.
…desktop) (NousResearch#57225) * feat(install): warn pip/Homebrew installs are unsupported (CLI, TUI, desktop) pip and Homebrew are now Unsupported install methods per website/docs/getting-started/platform-support.md. Surface a warn-don't-block deprecation notice everywhere the install method is already shown, pointing at the platform-support docs and noting these installs will not receive further updates. NixOS (Tier 2) is untouched. - hermes_cli/config.py: shared is_unsupported_install_method() / format_unsupported_install_warning() helpers so the wording and docs link stay consistent across every surface. - hermes_cli/banner.py: generalize the existing pip-only banner warning to also cover Homebrew. - hermes_cli/main.py: hermes update and hermes update --check print the warning before proceeding (still update; warn, don't block). - tui_gateway/server.py: session.info gains install_warning. - ui-tui: SessionPanel renders install_warning alongside the existing 'N commits behind' notice. - apps/desktop: SessionRuntimeInfo/GatewayEventPayload gain install_warning; applyRuntimeInfo + the live session.info event fire a snoozable warning toast via a new reportInstallMethodWarning(), mirroring the existing backend-contract-skew toast pattern. i18n strings added for en/zh/zh-hant/ja. - Tests: updated pip banner assertions for the new wording, added a Homebrew banner test, and two tui_gateway session_info tests (install_warning present for pip, absent for git). * fix(nix): make `hermes` in developement environment actually work install modules as editable overlay with uv * feat: print install method when running --version * fix: correct detect install method when running from a subtree
…reporting write_file() previously called _atomic_write() first and only ran the JSON/YAML/TOML/Python syntax check afterward as an informational lint delta -- a parse failure never set the top-level `error` key, so a corrupt structured-data write still landed on disk (and file_tools.py's files_modified gating, which keys off `error`, silently reported it as a successful modification). Move the in-process syntax check for JSON/YAML/TOML ahead of _atomic_write() and refuse the write outright on a parse failure: no temp file, no rename, nothing touches disk, and the result carries a top-level `error` so callers correctly see it as unmodified. Deliberately scoped to _FAIL_CLOSED_INPROC_EXTS (JSON/YAML/TOML), not all of LINTERS_INPROC -- .py is excluded because this codebase's own test fixtures (TestPatchReplacePostWriteVerification et al.) write arbitrary non-Python text through *.py paths purely to exercise write-mechanics; a hard block there broke 3 previously-passing tests during development. Python keeps its pre-existing non-blocking lint-delta report. Adds tests/tools/test_write_file_syntax_gate.py: invalid JSON/YAML/YML/ TOML refused with nothing written (new file) and nothing modified (existing file); valid JSON/YAML still written byte-for-byte; a non-linted extension with garbage content is unaffected; invalid Python is confirmed NOT hard-refused (still just reported).
…YAML isn't refused safe_load() raises ComposerError on multi-document streams (k8s manifests) and ConstructorError on application-defined tags (CloudFormation !Sub, Ansible !vault) — both valid YAML syntax. Now that the linter's verdict is a fail-closed write gate, those false positives would refuse legitimate writes outright. Switch to yaml.parse() (scanner+parser only), which still catches real syntax failures.
/update and other shutdown paths only waited on gateway session agents, so active cron tool work was killed immediately in final-cleanup while the scheduler could still mark the job successful (NousResearch#60432).
Cron jobs run through cron/scheduler.py's own ThreadPoolExecutor via a
standalone AIAgent (run_job/run_one_job), entirely outside
GatewayRunner._running_agents -- the dict _drain_active_agents() and
every other active-work check on that class reads. A gateway shutdown
(/update, /restart, and SIGUSR1 all funnel through the same stop())
could log active_at_start=0 and immediately kill tool subprocesses
while a cron job's terminal command was still running, with no wait
and no indication anything was interrupted.
Real-world impact (from the issue): a scheduled daily briefing cron
job was in flight during /update, its tool subprocess got killed
by the unconditional shutdown cleanup, and the job was never marked
failed -- it simply never completed or delivered, with no error
surfaced anywhere. A repro with a 30-minute `sleep` cron job in flight
during /update reproduced the same pattern: subprocess killed at
+0.22s of drain (active_at_start=0), the job's agent thread continued
in-process and produced a plausible-looking final response from the
truncated tool output, and the scheduler marked the run successful.
Root cause is layered, not a single line:
1. GatewayRunner._drain_active_agents() only waits on _running_agents.
Cron work was invisible to it, so drain returned instantly whenever
the only active work was a cron job.
2. Even with visibility, the shutdown's final tool-subprocess kill
(process_registry.kill_all()) is a global, unconditional sweep with
no per-job targeting -- a long-running cron job that outlives the
drain timeout still gets its subprocess killed.
3. cron/scheduler.py had no way to detect that a job's tool subprocess
was killed out from under it mid-run; the agent thread kept going
and its eventual (often degraded but plausible-looking) response
got reported as a normal successful completion.
Fix, three parts:
- cron/scheduler.py: expose get_running_job_ids() (thread-safe
snapshot of the existing _running_job_ids set, already used to
prevent double-dispatch) so the gateway can read cron's in-flight
state without reaching into private module internals.
- gateway/run.py: GatewayRunner._active_cron_job_count() reads that
snapshot. _drain_active_agents() now waits on
(_running_agents OR active cron jobs), so a cron-only workload gets
the same bounded wait chat sessions already get instead of an
instant active_at_start=0. Shutdown drain logging gains
cron_active_at_start/cron_active_now fields alongside the existing
ones (unchanged, for compat).
- cron/scheduler.py: mark_running_jobs_interrupted(reason), called by
gateway/run.py's _kill_tool_subprocesses() right after
process_registry.kill_all(), marks every job still in
_running_job_ids at that instant as failed/interrupted via the
existing mark_job_run() -- and records the job IDs in
_interrupted_job_ids BEFORE writing, so run_one_job()'s own
eventual completion for the same run (racing in its own thread)
checks that flag and skips its normal write instead of clobbering
the interrupted status with a false "ok" produced from the
now-truncated tool output. This does not attempt to correlate a
killed PID to a specific job ID (process_registry tracks PIDs, not
job IDs) -- any job still dispatched at the moment of a forced kill
is treated as interrupted, matching the existing coarser precedent
set by _interrupt_running_agents(), which interrupts every entry in
_running_agents on a drain timeout without per-agent correlation
either.
Deliberately out of scope (flagged in the issue as a separate,
lower-priority concern): startup-time reconciliation of cron runs that
started but never reached a terminal status.
Testing:
- tests/cron/test_shutdown_interrupt.py (12 tests): get_running_job_ids
snapshot semantics, mark_running_jobs_interrupted marking/no-op/
partial-failure behavior, and -- the core race guard -- run_one_job
skipping its own last_status write (both the success path and the
exception path) when the shutdown path already marked the run
interrupted, with a control test proving ordinary un-interrupted
completions are unaffected.
- tests/gateway/test_cron_active_work_drain.py (9 tests):
_active_cron_job_count reading cron state and failing closed (0) if
the cron module is unavailable; _drain_active_agents waiting for an
in-flight cron job the same way it waits for chat sessions, timing
out if the job outruns the window, and leaving existing chat-session
drain behavior unchanged; a full runner.stop() integration test
(drain-timeout path) proving mark_running_jobs_interrupted actually
fires with the right job ID when a tool subprocess is force-killed,
plus a no-op control when nothing cron-related is in flight.
- tests/gateway/test_shutdown_cache_cleanup.py: added
_active_cron_job_count() to that file's hand-rolled _FakeGateway test
double, which stop() now calls -- without it those 8 pre-existing
tests AttributeError (caught by fail-then-pass below, not a
production bug).
Fail-then-pass: reverted gateway/run.py + cron/scheduler.py, all 21
new tests fail (fixture/attribute errors -- the feature doesn't exist
yet); restored, all 21 pass.
Regression check: ran the full plausibly-affected surface --
tests/gateway/{test_gateway_shutdown,test_restart_drain,
test_restart_notification,test_restart_redelivery_dedup,
test_restart_resume_pending,test_restart_service_detection,
test_shutdown_cache_cleanup,test_stuck_loop,test_clean_shutdown_marker,
test_external_drain_control,test_session_state_cleanup,
test_update_command,test_update_streaming}.py plus tests/cron/ (944
tests) -- against a clean upstream/main checkout and against this
branch. Diffed the two FAILED lists: identical, 20 pre-existing
failures on both sides (Windows-locale/cp1252 file-encoding issues and
Unix-permission-bit assertions that don't apply on this Windows dev
box), zero new failures, zero fixed-by-accident. The 8
test_shutdown_cache_cleanup.py failures found mid-development were
from the _FakeGateway gap above, fixed in the same commit and
confirmed clean on the final rerun (diff against baseline: exit 0).
Fixes NousResearch#60432
Follow-up to the previous commit on NousResearch#60432. The status-write guard (_consume_interrupted_flag, checked right before mark_job_run) closes the false-success bookkeeping gap, but run_one_job delivers its result BEFORE that check: delivery happens right after run_job() returns, mark_job_run happens at the very end. A job whose tool subprocess was killed mid-flight can still produce a plausible-looking final_response from the truncated output, and that response would reach the user via _deliver_result before the interrupted flag was ever consulted -- correct status in jobs.json, wrong message already sent. Adds _is_interrupted(), a non-destructive peek at the same _interrupted_job_ids set (_consume_interrupted_flag stays as the consuming, authoritative check right before the status write -- this needed a peek instead since the flag has to still be visible there). Checked right after save_job_output, before the deliver_content decision: if the run looked successful but was flagged interrupted, force success=False with an explicit interruption message. This routes delivery through the existing _summarize_cron_failure_for_delivery path (the same one a real failure already uses) instead of the raw final_response, so the user gets an honest "this run was interrupted" instead of a truncated/misleading result. Testing: 4 new tests in tests/cron/test_shutdown_interrupt.py -- _is_interrupted peek semantics (false/true/does-not-clear, as opposed to the consuming _consume_interrupted_flag), and the delivery-gate test itself, which mocks run_job to return a normal-looking success with a "plausible final response" while the job is pre-marked interrupted, and asserts _deliver_result receives the failure summary ("This run was interrupted.") instead, with the summarizer's error argument confirmed to mention the interruption. Fail-then-pass: reverted cron/scheduler.py only, the 4 new tests fail (3 on the missing _is_interrupted attribute, 1 -- the delivery-gate test -- on _summarize_cron_failure_for_delivery never being called, i.e. the raw response would have gone out); restored, all 16 tests in the file pass. Regression: tests/cron/ (683 tests) + test_cron_active_work_drain.py + test_gateway_shutdown.py + test_shutdown_cache_cleanup.py -- 11 pre-existing failures (Unix file-permission-bit and path-tilde assertions that don't apply on this Windows dev box), matching the same set already established as pre-existing in the prior commit's regression check. Zero new failures. Continues NousResearch#60432
…onto one drain surface Keep NousResearch#60631's get_running_job_ids() snapshot + _active_cron_job_count() (import-guarded for minimal test doubles) as the single read path, and retarget NousResearch#60612's drain tests at it. Drops the redundant cron_jobs_in_flight() helper so there is one surface, not two.
Guard _finalize_session's db.end_session() call against gateway-owned sessions (telegram, bluebubbles, discord, etc.). The TUI is a viewer for these sessions, not the lifecycle owner. Unconditionally ending them in state.db creates a Groundhog Day routing loop: the gateway's NousResearch#54878 self-heal detects the stale entry, recovers to the parent session, context compression splits back to the reaped child, and the cycle repeats on every inbound message — causing complete conversational context amnesia. Fixes NousResearch#60609
…hardcoded list The salvaged guard used a hand-maintained frozenset of 14 platform names — several of which (line, wechat, facebook, imessage, googlechat) aren't actual Hermes Platform values, while real ones (whatsapp_cloud, feishu, wecom, dingtalk, qqbot, yuanbao, plugin platforms like irc) were missing. Resolve the source through gateway.config.Platform instead (built-ins + registered plugin platforms via _missing_), with an explicit exclusion set for self-owned/local sources. Adds tests for the guard and both reap paths.
…S-free) (NousResearch#60730) For air-gapped / self-hosted-IdP deploys with NO Nous Portal, let the gateway obtain its caller-identity bearer from a generic OAuth2 client_credentials grant against the operator's own IdP (e.g. Microsoft Entra ID) instead of only resolve_nous_access_token(). The connector's OIDC tenant resolver reads a claim (default tid) off that token as the tenant. - gateway/relay: new canonical _resolve_relay_identity_token() — client_credentials when gateway.idp.token_url (or GATEWAY_RELAY_IDP_* env) is set, else Nous Portal (unchanged default). Wired into self_provision_relay(). - hermes_cli/gateway_enroll: _resolve_identity_token() delegates to the canonical resolver so the enroll CLI and the runtime self-provision path share ONE impl. Config via gateway.idp.{token_url,client_id,client_secret,scope} in config.yaml (env override GATEWAY_RELAY_IDP_*). No behaviour change when unset. Tests: tests/gateway/relay/test_identity_token_resolver.py (6 — mode selection, request shape, config/env precedence, fail-closed). Relay suite 162 pass. Validated via the cross-repo gateway<->connector live E2E (provision, managed self-provision, inbound round-trip, /link) against a connector running the OIDC tenant resolver with zero NAS config.
(cherry picked from commit 6ed8849)
_save_provider_state() sets auth_store['active_provider'] as a side effect. The Z.AI endpoint probe runs from credential-pool env seeding for any user with a Z.AI key in env — persisting the probe cache must not silently make zai the active provider. Use _store_provider_state(set_active=False). Follow-up to PR NousResearch#41201 salvage.
Review findings (hermes-pr-review Phase 2, 3-angle): - _save_auth_store() does real filesystem I/O (mkdir, O_EXCL create, fsync, atomic replace) and can raise on disk-full/permissions/lock-timeout. The persist ran bare in the success path, so a persist failure aborted _resolve_zai_base_url() after detection had already succeeded. Wrap the persist in try/except: log a warning and still return the detected URL (worst case: next start re-probes). - Readability: stage the payload in a local detected_endpoint instead of writing through the stale pre-lock 'state' dict, which is no longer what gets persisted.
Windows drops webContents zoom on minimize/restore, so the UI snapped back to 100% while Settings still read 125%. Zoom was only reasserted on the main window's did-finish-load, never on show/restore and never for session windows. Reassert the persisted level on show/restore + first load, wired once in wireCommonWindowHandlers so the main window and secondary session windows share it. The pet overlay opts out (zoom:false): it sizes its own OS window to fit the sprite in unzoomed CSS px and has its own Alt+wheel scale, so inheriting the global zoom would render the mascot larger than its window and crop it (and it shares the renderer origin's zoom localStorage key). Salvages NousResearch#61245; keeps its pure-helper tests and adds a scope assertion. Co-authored-by: HexLab98 <liruixinch@outlook.com>
…245-ui-zoom fix(desktop): re-apply UI zoom on show/restore, scoped to chat windows (supersedes NousResearch#61245)
Dashboard Chat is an xterm mirror of a TUI inside the gateway, so server-side clipboard.paste never sees the browser clipboard. Upload pasted/dropped images to the profile's images/ dir (same place clipboard.paste / image.attach use), then drive /image over the PTY. Uses a dedicated /api/chat/image-upload endpoint (magic-byte check, 25MB cap, profile scope) instead of relative managed-files uploads that 400 on local dashboards without a locked root. Ctrl/Cmd+Shift+V also tries clipboard.read() for images before falling back to text, since preventDefault on that chord suppresses the DOM paste event. Salvages NousResearch#57912 (client composition + /image PTY drive) and folds in NousResearch#48563's upload endpoint + drop path. Co-authored-by: bird <6666242+bird@users.noreply.github.com> Co-authored-by: tt-a1i <53142663+tt-a1i@users.noreply.github.com>
…912-dash-paste fix(web): paste/drop images into dashboard Chat via HERMES_HOME/images
Old packaged Desktop apps re-entering bootstrap against an existing ~/.hermes/hermes-agent were still passing the baked-in --commit pin, which detached the managed checkout back to the app stamp (e.g. 0.15.1) after hermes update had already moved it forward. Skip the packaged commit pin when activeRoot already has git metadata; keep branch args and fresh-install commit pinning unchanged. Port of NousResearch#59902 onto the post-ts-ify bootstrap-runner.ts. Co-authored-by: helix4u <4317663+helix4u@users.noreply.github.com>
…902-bootstrap-repin fix(desktop): prevent bootstrap stale commit repin on existing checkouts
Switching connection mode (local / cloud agent / remote) no longer full-window-reloads into the cold-boot CONNECTING screen. The primary backend is torn down in place (no renderer reload); the shell + Settings stay up while session lists are wiped so sidebar skeletons retrigger, then the socket re-dials and config/sessions refresh. Cold-boot CONNECTING latches off after the first successful boot; the intentional teardown suppresses the backend-exit toast. Dev affordance: a "Preview soft switch" button under Gateway diagnostics (Electron has no ?query= entry). Gateway settings UI brought in line with the rest of Settings: - Mode cards use the shared selectableCardClass on an equal-height auto-rows-fr grid, stacking 1→3 (never an orphaned 2+1); titles wrap instead of truncating. - Remote gateway's auth detail moves into a ? tooltip in the title; drop the redundant "connects to the one you choose" from the cloud card. - textStrong buttons force px-0 so the underline sits flush with the label. - Tooltip chip uses box-decoration-break: clone so the background hugs each wrapped line (bg only on the text), capped at max-w-64. Fully i18n'd (en + zh; ja/zh-hant inherit via defineLocale).
…402-desktop-cloud-mode feat(desktop): Hermes Cloud connection mode (salvage of NousResearch#55402)
…teway-switch-ux feat(desktop): soft gateway switch + gateway-settings polish
Radix's hoverable-content grace area can leave tips stuck over Electron drag regions; disable it and make tip content pointer-events-none so open state tracks the trigger only.
…p-stuck fix(desktop): stop Tip from sticking open and blocking clicks
Catch-up merge of upstream (frozen target caf557b) into the long-lived hard fork. Fork was 386 ahead / 1216 behind at merge-base 852c9b3. 56 conflict files, 147 hunks (140 both-sides semantic + mem0 arch-split). All 386 fork lead-commits preserved; import-smoke + marker-clean verified before commit. Architectural ports & semantic reconciliations (full detail in docs/sync/review/): - plugins/memory/mem0: take-fork wholesale; removed upstream _backend/_setup refactor shards + their tests (fork __init__ is the production superset). Upstream setup-wizard refactor noted for deliberate follow-up, not reintroduced as orphan modules. - agent/auxiliary_client _resolve_task_provider_model: reconciled api_mode — first-class/ custom/auto keep upstream api_mode=None (downstream transport detection); plugin providers keep the fork eager ProviderProfile.api_mode read. - agent/chat_completion_helpers: keep-both — fork relay/fallback route announce + audit sink + x-hermes-lane pool headers preserved; added upstream unavailable-fallback skip. - tools/delegate_tool: fork inherit_context/boomerang kept; code_execution stays exempt from subagent toolset stripping; took upstream ACP transport-field removal. - gateway/run.py, gateway/slash_commands.py: interleave — fork restart-loop/recovery gates, compaction telemetry, branch-thread, usage cards preserved; upstream async session-store calls used where hunk-local. - gateway/session.py: keep-both — fork durable reasoning/model identity + answerable lookup; upstream model_override + rewrite_transcript success-return (clear undo only after successful rewrite). - hermes_state.py: keep-both — fork gated effective_last_active denorm + backfill; upstream MAX_FTS5_QUERY_CHARS, handoff index, search_query, compact_rows (denorm fast-path bypassed when upstream-only search/compact options requested). - cron/scheduler.py: keep-both — fork per-job reasoning/fallback + loud alert; upstream dispatch-claiming, credential-exfil guard, TG DM-topic probe, deferred teardown. - tui_gateway/server.py: reconciled fork client source attribution with upstream platform_override; _make_agent accepts both, disabled reasoning reports "none". - run_agent.py: interleaved upstream intrinsic-persistence markers with fork interrupt-close finish-reason re-persist; row ids attached to persisted dicts (resume discriminator no longer depends on long-lived id() sets). - agent/model_metadata.py: keep-both — fork two-tier token breakdown + request telemetry; upstream async context lookup + local-probe cache + bounded tool-token cache. - Desktop: upstream Electron TS migration/package shape; fork app-side media/project/composer behavior preserved. Fork-critical features re-verified present post-merge: relay x-hermes-lane headers, messaging/moa toolsets + send_message/mixture_of_agents tools, code_execution exemption. Resolution by codex under docs/sync/RESOLUTION-SPEC-2026-07-10.md; certified (markers, import-smoke, feature-survival) by Apollo. POST-MERGE RECONCILIATION FIXES (STAGE-2 gates: 31 reds → 0; detail in docs/sync/review/stage2-reconciliation-fixes.md): - agent/agent_runtime_helpers.py: removed a stray set-era `known_tool_ids.discard()` the set→Counter merge left behind (NameError broke all 16 message-sequence-repair tests). - cron/scheduler.py: added the upstream one-shot dispatch claim to _process_one_job (the fork's tick path), not just run_one_job — a self-destructing one-shot could otherwise re-fire forever on the built-in ticker. - tui_gateway/server.py: restored the fork's _sanitize_client_source guard on the untrusted client `source` in _make_agent (the merge trusted it verbatim → hostile-label injection). - gateway/slash_commands.py: restored the fork's chat-only /compress `msgs` projection (the merge took upstream's full-transcript filter → double-counted tool rows, broke both-axes compress-feedback math). - tests updated to real upstream contracts the merge adopted: FailoverReason.ssl_cert_verification; _notify_session_boundary platform arg; _notification_event_belongs_elsewhere sid arg; list_sessions_rich compact_rows; _SlashWorker profile_home; config-sync persist_override; _resolve_session_agent_runtime override stubs. No fork feature dropped, no test weakened. STAGE-2 pytest gates green (2267 passed; 1 pre-existing order-pollution flake that passes isolated). CI-gate traps + config-migration dry-run follow before deploy (deploy is a separate gated step, Ace's call).
|
Too many files changed for review. ( |
…t-up-to-date #253 added TestCompactRows::test_compact_projection_tracks_schema (compact rows must carry every schema column). The parity merge's fork denorm column effective_last_active is deliberately stripped from list rows (consumers use the computed last_active; the raw value is internal to the recency-ordering CTEs, and the fork denorm-reland oracle tests assert its absence). Reconciled by adding effective_last_active to the test's sanctioned exclusion set alongside system_prompt — preserving the fork's clean-list-row contract without weakening #253's schema-tracking guard.
TypeScript CI on the parity PR caught three desktop merge-resolution gaps where the merge kept the fork side of a conflicted file but consumers/tests reference upstream symbols the fork lacks: - apps/desktop/src/lib/media.ts: re-add upstream's downloadGatewayMediaFile export + its readDesktopFileDataUrl import (consumed by artifacts/index.tsx + markdown-text.tsx). - use-composer-actions.test.ts: the merge kept the fork's readComposerImagePreview impl but unioned upstream's attachmentPreviewDataUrl test block (a function the fork doesn't export). Removed the orphaned describe + its unused $connection/attachmentPreviewDataUrl imports; the routing behavior is covered by desktop-fs.test.ts. - use-prompt-actions/index.test.tsx: add the upstream-added required PromptActionsOptions field openMemoryGraph (real, used by the impl + desktop-controller) to the test options. Verified each symbol against upstream caf557b. No behavior change to shipping code beyond restoring the dropped upstream export.
Five more merge-reconciliation gaps the full Python CI suite caught (all passed on fork/main; regressions the merge introduced by taking one side of a conflicted seam): - agent/conversation_loop.py preflight compression: the merge kept upstream's 'Pre-API pressure check' block, which calls the compressor's REMOVED should_defer_preflight_to_real_usage ratchet (fork replaced it with skew CALIBRATION). The getattr-lambda fallback silently returned False → compressed when it should defer. Switched the gate to the fork's should_compress_calibrated. - agent/conversation_loop.py MoA aggregator-cost block (upstream-only): gated its last_aggregator_slot read on isinstance(dict), so a non-MoA/MagicMock client can't route a mock provider/model into is_notional_anthropic_provider (TypeError). - gateway/run.py auto-resume: reconciled upstream's global restart-loop guard (restart_loop_guard.check_and_record, coarse: skips ALL auto-resume) with the fork's PRIMARY per-session F2 replay-mark→SUSPEND breaker. The global guard fired first and returned 0 before any session resumed, so _resumed_this_boot never populated and F2 could never suspend the looping session. Now the global guard still records every boot but only short-circuits when the per-session breaker is disabled; F2 owns the break. - gateway/slash_commands.py: converted 10 raw self.session_store.X() calls in async slash-command methods (redo/compress/merge/branch) to await async_session_store.X(), satisfying the upstream async-session-store AST guard (no sync SQLite on the event loop). - tests/gateway/test_compression_session_id_persistence.py: the fork's hand-rolled AST walker missed the session_entry.session_id assignments after upstream restructured _handle_message_with_agent; rewrote it with a robust block-recursion (and it now recognizes the async await ...async_session_store._save()). Verified non-vacuous: it finds both assignments and confirms each is followed by _save().
…ssion + test drift
SECURITY FIX (real regression):
- tools/approval.py _run_approval_gate: the upstream gate-refactor read the raw
env_var_enabled('HERMES_CRON_SESSION') to detect cron context. A cron-SPAWNED
subagent runs in a bare ThreadPoolExecutor worker that does NOT inherit the
os.environ marker — its cron identity lives only in the re-bound ContextVar
(_bind_child_cron_session → set_cron_session). So a cron child AUTO-APPROVED
dangerous commands (curl|bash), defeating cron_mode: deny. Switched to the
context-aware _is_cron_session() (ContextVar-first, os.environ fallback) and
removed the now-dead duplicate cron-gate block the merge left unreachable after
check_dangerous_command's return _run_approval_gate(...).
TEST reconciliations to real upstream contracts the merge adopted:
- test_run_agent_tool_exec.py: (a) 13 pre-tool-block patches retargeted from the
fork's deprecated get_pre_tool_call_block_message shim to upstream's centralized
resolve_pre_tool_block (which every dispatch path now uses); (b) MCP tool names
updated from the fork's single-underscore mcp_x to upstream's mcp__server__tool
convention (MCP_TOOL_NAME_PREFIX='mcp__'); (c) malformed-JSON-args test updated
to upstream's STRICTER behavior — malformed args are now rejected with an error
tool result instead of the fork's lenient coerce-to-{} + run (BEHAVIOR CHANGE,
documented inline: safer — never execute a tool with fabricated empty args).
- test_web_server_sessiondb_eventloop.py: fake _DB now accepts SessionDB(read_only=)
+ list_sessions_rich(compact_rows=) and DEFAULT_DB_PATH.exists() is stubbed True —
upstream added a read-only open + fresh-install exists() guard + compact_rows.
All targeted suites green: test_run_agent_tool_exec (61), test_cron_subagent_session (5),
test_approval (303), test_web_server_sessiondb_eventloop (9).
Fixes the last 14 test failures surfaced by the full CI matrix after the upstream parity merge. Each is a fork/upstream reconciliation the conflict resolution took one side of; classified against the fork/main baseline (pass-on-fork => my regression, fix code; fail-on-fork => stale test, update). Real regressions fixed in code (all passed on fork/main, broke in the merge): - gateway hygiene announce stranded in the no-op else branch: a *rotated* hygiene compaction silently sent no in-chat "Context compacted" announce on the live fleet. Restored the fork's "announce on any successful compaction (rotate OR in-place)" structure. - run_agent flush: the persist-user-message override lost the platform-message-id path (upstream's inline override handled content/timestamp only), breaking backfill-on-reconnect dedup (has_platform_message_id). Re-threaded _row_platform_id. - telegram adapter _handle_polling_network_error read instance-only retry attrs; getattr-fallback so the bare-adapter recovery path works (NousResearch#59614). - length-continuation + truncated-tool-call retry caps drifted to 4 while the rest of the retry subsystem (codex/invalid-tool/invalid-json/empty-content) stayed at 3 and a comment still said 3 — restored the fork's coherent 3. - /usage: upstream Context:/Compressions: header lines rendered above the context breakdown, violating the fork's "breakdown FIRST" contract. Removed. - browser_connect: a log string embedded the --remote-debugging-port literal that trips the keychain-guard launcher scanner (false positive); reworded. Stale tests updated to the merged (upstream) contract: - cron approve/deny approval tests patched env_var_enabled; the security fix routes cron detection through _is_cron_session() (ContextVar) now. - compaction legacy test: small-ctx floor now applies on update_model too. - todo hydrate: requires a paired assistant todo tool call (GHSA-5g4g-6jrg-mw3g). - persist voice-input: override applies to the DB row only, never the live dict (NousResearch#48677); assert the live content is preserved. - telegram restart-parity guard: dataflow-resolve the _start_polling_resilient wrapper's forwarded drop_pending_updates param (keeps the guard's teeth on the real recovery ladders). Full affected-subsystem sweep green; no new failures introduced.
…ice 3) The merge took upstream's always-TEMPFAIL-75 service-restart exit code, but Ace's fleet uses the fork's deployment model: exit 0 under systemd (Restart=always relaunches on clean exit; 75 would trip stepped-restart backoff) and 75 only on darwin/launchd (KeepAlive.SuccessfulExit=false needs non-zero). Branch on INVOCATION_ID + platform. Skipped on macOS; the exit==0 path runs on Linux CI (test_gateway_stop_systemd_service_restart_exits_cleanly) and the 75 path is covered by the paired launchd test.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream parity sync — NousResearch/main → fork/main
Catch-up merge of upstream into the long-lived hard fork. Frozen upstream target
caf557be.Fork was 386 ahead / 1216 behind at merge-base
852c9b3cb.Scale
How it was resolved
Conflict resolution driven by codex under a written spec (
docs/sync/RESOLUTION-SPEC-2026-07-10.md), then certified by Apollo (markers-clean, import-smoke, fork-feature survival). Every fleet-critical file resolved by reading both sides' history — no blind side-picks. Architectural ports + semantic reconciliations documented indocs/sync/review/.STAGE-2 gates: 31 reds → 0
Four were real code regressions the merge introduced (fixed the code, preserving fork behavior):
agent/agent_runtime_helpers.py— strayknown_tool_ids.discard()from the set→Counter migration (NameError broke all 16 message-sequence-repair tests).cron/scheduler.py— one-shot dispatch claim was missing on the fork'stick()path (_process_one_job), only inrun_one_job; a self-destructing one-shot could re-fire forever.tui_gateway/server.py— untrusted clientsourcelost its_sanitize_client_sourceguard (hostile-label → persisted platform).gateway/slash_commands.py—/compresschat-onlymsgsprojection clobbered by upstream's full-transcript filter (double-counted tool rows, broke both-axes feedback math).The rest were stale tests updated to real upstream contracts the merge correctly adopted (
FailoverReason.ssl_cert_verification,_notify_session_boundaryplatform arg,_notification_event_belongs_elsewheresid arg,list_sessions_richcompact_rows,_SlashWorkerprofile_home, config-sync persist_override). No fork feature dropped; no test weakened. Full detail:docs/sync/review/stage2-reconciliation-fixes.md.Result: 2267 passed across all touched subsystems (1 pre-existing order-pollution flake that passes isolated).
CI-gate traps pre-checked locally
.gitleaks.toml+ negative-tested (a real injected key still trips).messaging/send_message+moa/mixture_of_agentssurvive (resolve > 0). No dropped-toolset finding.Merge topology
Landed as a real 2-parent merge commit — merge with
--merge, NOT squash (squashing destroys the merge-base the next upstream sync needs).Deploy
🔴 Merging this does NOT deploy to the fleet. Live rollout is a separate gated step (
hermes updateper host, host-1 ≠ the deploying gateway) — Ace's call, on his timing.