chore: reconcile fork fixes with current upstream - #68
Conversation
Review-response sweep found one straggler: model-picker.tsx's visibleDownloads filter still used raw toLowerCase().includes(), so a hyphen-style query wouldn't match a space-separated download target — the same bug class this PR fixes, in the same picker. The catalog menu's equivalent filter was already converted.
Keep command-palette literal filter semantics unchanged; opt model surfaces into the shared fold. Preserve original highlighter contracts and trim new regressions to two invariants. Native Electron catalog before/after and visibility dialog verified; campaign suites remain queued.
`session-unread-tile.test.ts` intermittently fails CI with `Test timed out in 15000ms` at the first case. It has hit at least three independent branches, `main` included, so it is not tied to any one change. The cost is module reconstruction. `beforeEach`/`afterEach` both call `vi.resetModules()`, so each of the three cases re-imports six modules, including `@/lib/chat-runtime` and the pane-tree store. Locally that is ~3.3s on median but the tail reaches 11.8s (3.6x the median) on a warm machine; a loaded CI runner pushes that past the 15s budget. One observed CI run passed at 14614ms, 386ms under the limit, which is the same test sitting on the wrong side of the same boundary. Drop `resetModules` and undo the per-test state explicitly instead: - collect the `registry.register` disposers and run them in `afterEach`, matching what `session-states.test.ts` already does; - reset `$layoutTree` and `$activeTreeGroup` at the top of `setup()`. `declareDefaultTree` only seeds the layout when it is empty, so without this the second case would adopt the first case's tree. The test keeps its teeth: reverting the `$focusedStoredSessionId` change from a5b5043 still fails exactly the same two cases as before this patch. Slowest of 30 consecutive local runs goes from 11.84s to 3.70s, with the median roughly unchanged.
The previous commit bumped pyproject.toml and uv.lock but missed the
LAZY_DEPS exact pin for platform.discord, so
test_pyproject_pins_match_lazy_deps_pins and
test_every_lazy_deps_exact_pin_matches_uv_lock fail with
{'brotlicffi': {'platform.discord': {'lazy_pin': '1.2.0.1', 'uv_lock': ['1.2.0.2']}}}.
Update the third registration site to keep all three in lockstep.
…ming DecodingError
Remount scope-owned credential state on Applies-to changes and bind onboarding requests to their initiating route. Cancel polling and invalidate late results when setup closes or reopens, without undoing writes already sent. Co-authored-by: By JTT <29462570+jordan-thirkle@users.noreply.github.com>
…longer stalled; tell the model results land between turns Three orchestrator failures traced through the Sep 7 gpt-6-astra campaign sessions: 1. delegation.independent_completions (new, default false). NousResearch#104299 made every ungrouped task its own completion message, so a 15-task call woke the orchestrator up to 15 times; one chain received 132 notices and answered 130 of them with "already incorporated". A multi-task call now returns as ONE consolidated message unless the flag is on; `group` is inert until then. 2. Queued units were killed before they started. Units of one call share a pool slot but the executor was still sized by slots, so with 15 units live a new unit queued behind a full pool; the stale monitor's clock ran from dispatch, interrupted it at 450 s, and the child exited `interrupted 0.02s` when its thread finally came up (13 such lanes in one session). The executor now grows to the number of live units and the stall clock arms when the runner actually starts. 3. The tool text said "do not wait or poll — just continue" without saying that completions are delivered only BETWEEN turns. A model that never ends its turn (one 203-minute turn, 717 API calls) never received 40 finished results. Tool description, dispatch note and completion header now say to finish independent work, give a one-line status, and end the turn.
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Salvage NousResearch#101887 after native Actions 34097643131 reproduced WinError 32 using a ready Electron app with its cwd inside the live release. Reuse the existing install-scoped process cleanup before promotion and wait after forced termination, preserving rollback. Co-authored-by: fangliquan <fangliquan@qq.com>
Document actual groups.promote/groups.demote parameters and required old-writer fencing before confirmation. Demotion is a controlled rejoin step, not an atomic promote-then-demote handover or log reconciliation. Clarify replica coverage, confirmation meaning, and lineage readback. Corrected redo of NousResearch#104342; its nonexistent groups.peer methods and unsafe handover ordering are not carried forward. Fixes NousResearch#104309 Refs NousResearch#104904 Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
…ousResearch#104150) _paginate_full_list wrapped the paginated list call in try/except TypeError to detect the mcp 1.x calling convention. The same except also caught TypeErrors raised INSIDE the modern list call — e.g. a server response decode failure — and retried with the legacy cursor= keyword, replacing the real error with a misleading 'unexpected keyword argument cursor' and making genuine MCP pagination failures undiagnosable. Probe list_method's signature instead (_list_method_accepts_params): the legacy cursor= fallback fires only when the method genuinely doesn't accept the mcp 2.0 params= keyword (or takes **kwargs), so a TypeError from inside the list call propagates to the caller. Regression tests: the decode TypeError surfaces and the legacy retry doesn't run; a genuinely 1.x-shaped method keeps using the cursor fallback.
…s resolve ClawHub's detail endpoint now answers a slug claimed by multiple owners with 409 AMBIGUOUS_SKILL_SLUG; the bare GET in _skill_detail returned None for every such slug, so 'skills install clawhub/@owner/slug' (and the owner/skills/slug URL form) failed at fetch time even though the requester already knew the owner (NousResearch#104117). - _skill_detail forwards expected_owner as the ?owner= query param on the detail GET (params already flows through _get_json's **kwargs). - _parse_identifier also accepts the clawhub/@owner/slug combination: the @ surfaces only after the clawhub/ prefix is stripped, so the had_at check now re-runs on the stripped form. GitHub-style owner/repo/skill paths stay rejected.
Adapt the earliest routing repair in scroasdale PR NousResearch#44268 to the current media helpers, retaining metadata in URL fallbacks and refusing successful text-only receipts for failed local uploads. Also informed by jasondschoeman-pixel issue NousResearch#104357 and ericmaddox PR NousResearch#104760. Co-authored-by: scroasdale <67333169+scroasdale@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
… revocation (NousResearch#102602) * fix(relay): re-dial once with a fresh token before treating a 4401 as revocation A 4401 close after a successful handshake was read unconditionally as the connector having revoked this gateway's per-gateway secret (opt-out), so the transport latched auth_revoked, the adapter went relay_disabled, and all messaging stopped until a manual restart. But the connector sends the same plain 4401 'unauthorized' for an EXPIRED upgrade token (make_upgrade_token TTL is 300s). Scale-to-zero makes that routine: the instance is suspended while a re-dial is in flight, the token was minted before the freeze, the dial completes on resume with a token past its TTL, the connector refuses it with 4401, and the gateway misreads an expired token as a revoked credential. Production incident 2026-09-02 - the connector DB secret was never revoked. Fix, in gateway/relay/ws_transport.py: - The first post-handshake 4401 is provisional. The reader schedules ONE immediate re-dial (_redial_with_fresh_token) that bypasses the reconnect backoff; _dial_and_start mints a fresh token on every call. The backoff supervisor design is untouched. - Only a 4401 against that fresh token (either refused at the upgrade or closed after a descriptor on that connection, tracked by dial generation) latches auth_revoked. Terminal behaviour is otherwise unchanged. - If the fresh dial fails for a non-auth reason, hand off to the normal backoff supervisor. - Read the Close frame reason (_close_reason_of, sibling of _close_code_of). A 4401 whose reason is exactly 'expired' never latches revocation and takes the normal reconnect path - forward-compatible hook for the connector change landing separately. - disconnect() cancels the one-shot retry task alongside the supervisor. Tests (tests/gateway/relay/test_ws_transport.py, real websockets server): - 4401 once, next dial accepted -> reconnected, auth_revoked False, exactly 2 dials. - 4401 on every dial -> auth_revoked True after exactly one retry (2 dials), no supervisor, no further dials. - 4401 reason 'expired' (once and repeated) -> never latched, reconnects via the normal supervisor. - Existing 7d-B tests keep passing (4401 before any handshake stays retryable; the revoking-every-dial stub still latches). * fix(relay): tolerate a partially built transport in disconnect() teardown Three teardown tests construct WebSocketRelayTransport via object.__new__ without __init__, so the new _auth_retry attribute was absent and disconnect() raised AttributeError. Read it with getattr like the other optional teardown handles. * fix(relay): one live dialer — the fresh-token retry is a flag on the next dial, not a second dialer Review (round 1) found a race: a reader that dies with a provisional 4401 while the backoff supervisor is already mid-dial started _redial_with_fresh_token as a SECOND concurrent dialer; both installed sockets/readers and the supervisor could overwrite the accepted retry socket (3 dials, wrong socket). Now a provisional 4401 sets _auth_retry_pending; _dial_and_start consumes it and stamps _auth_retry_generation on whichever dialer performs the next dial. The reader arms a dialer only when none is live (_dialer_running), and an upgrade-time 4401 on that generation latches revocation from either dialer (_latch_if_fresh_token_refused). The reader also iterates its captured ws handle, not self._ws. Regression test reproduces the race with a deterministic fake connect; it fails with the _dialer_running guard removed. * fix(relay): a dial whose reader died mid-hello is a failed dial; the retry marker survives network failures Review round 2: - BLOCKER: with one-live-dialer, a reader that dies while its own dialer is still inside _dial_and_start (hello in flight) arms nothing — that is the dialer's job — but the dialer then returned 'connected', leaving no socket, no reader, no dialer. _dial_and_start now raises ConnectionError when the reader it installed has already finished, so both dialers take their normal failure path (retry -> supervisor; supervisor -> backoff). - MAJOR: _auth_retry_pending was consumed before websockets.connect, so a connect-time network failure un-marked the retry and the NEXT dial's real fresh-token 4401 read as another first strike (revocation never latched). The marker is now consumed only when the token reaches an auth outcome: upgrade accepted (stamp the generation) or upgrade 4401'd (judge it). Two regression tests, each mutation-checked red against its own guard.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
…rk-upstream-sync-20260908
64b2e7b to
cb08558
Compare
૮ >ﻌ< ა ci reviewran on 1d69242 — fix: repair CI event and provider registry regressions
|
| Package | Before | After |
|---|---|---|
| apps/desktop | 0.17.0 |
0.17.1 |
How to fix:
Add the ci-reviewed label after verifying the version changes are expected.
⚠️ Warnings
CI timings · View report · View job
Wall time 14m7s vs 7m52s (+79.4%). 10 job(s) slower, 20 faster, 1 unchanged.
- JS & TS checks / JS & TS checks shard 2/5: +197.0s
- Python tests / Run tests slice 1/8: +95.0s
- Python tests / Run tests slice 4/8: -76.0s
- Python tests / Run tests slice 8/8: +68.0s
- JS & TS checks / JS & TS checks shard 5/5: -63.0s
OSV vulnerability scan · View job
28 known vulnerabilities found in pinned dependencies.
- CVE-2026-67213 in website/package-lock.json
- CVE-2026-82417 in scripts/whatsapp-bridge/package-lock.json
- CVE-2026-82417 in website/package-lock.json
- CVE-2026-75931 in package-lock.json
- CVE-2026-75931 in website/package-lock.json
- CVE-2026-83610 in package-lock.json
- CVE-2026-83610 in package-lock.json
- CVE-2026-71554 in uv.lock
- CVE-2026-73088 in package-lock.json
- CVE-2026-73088 in website/package-lock.json
- GHSA-8423-8fgw-73vq in uv.lock
- CVE-2026-70608 in package-lock.json
- CVE-2026-73089 in package-lock.json
- CVE-2026-73089 in website/package-lock.json
- CVE-2026-75975 in package-lock.json
- CVE-2026-75975 in website/package-lock.json
- CVE-2026-75899 in package-lock.json
- CVE-2026-75899 in website/package-lock.json
- CVE-2026-84371 in package-lock.json
- CVE-2026-56876 in package-lock.json
How to fix:
Review the findings in the Security tab. Update the affected dependencies if a patched version is available.
ℹ️ Info
CI-sensitive file review · View job
PR touches sensitive files, but the ci-reviewed label has been added, approving them.
Sensitive files changed:
commented
Sep 8, 2026
💡 Codex Reviewhermes-agent/agent/turn_truncation.py Line 53 in 64b2e7b When repeated AGENTS.md reference: AGENTS.md:L47-L48 This test reads AGENTS.md reference: AGENTS.md:L390-L395 On Linux CI this patch makes the interpreter claim it is Darwin, and the rest of this new file repeatedly patches AGENTS.md reference: AGENTS.md:L351-L358 hermes-agent/agent/api_request_hooks.py Line 53 in 64b2e7b
AGENTS.md reference: AGENTS.md:L79-L82 ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
left a comment
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1d69242d63
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| return isinstance(tool_calls, list) and any( | ||
| cls._get_tool_call_id_static(tc) == tool_call_id | ||
| and cls._get_tool_call_name_static(tc) in {"todo", "todo_list"} | ||
| cls._get_tool_call_id_static(tc) == tool_call_id and cls._get_tool_call_name_static(tc) == "todo" |
There was a problem hiding this comment.
Accept todo_list when restoring plans
The registered wire tool is todo_list, so normal persisted tool responses no longer satisfy this check and _hydrate_todo_store restores an empty plan when a session resumes. The same narrowing in acp_adapter/server.py also suppresses native plan updates during ACP history replay; accept the canonical todo_list name in both sibling paths while retaining todo for legacy history.
AGENTS.md reference: AGENTS.md:L47-L48
Useful? React with 👍 / 👎.
| return () => { | ||
| cancelled = true | ||
| clearTimeout(retryTimer) |
There was a problem hiding this comment.
Preserve queue failure accounting across busy renders
When a foreground queue drain calls the real submit pipeline, submit.ts synchronously changes busy to true before awaiting the gateway, which reruns this effect and executes this cleanup. That sets cancelled before a rejection reaches onFail, so the failure count and retry timer are skipped; when submit resets busy to false, the queue can immediately fire again with its count still at zero, bypassing the advertised backoff and attempt cap indefinitely.
Useful? React with 👍 / 👎.
| name: web-perplexity | ||
| version: 1.0.0 | ||
| description: "Perplexity Search API web search and query-relevant page snippets. Requires PERPLEXITY_API_KEY — get one at https://www.perplexity.ai/account/api." |
There was a problem hiding this comment.
Move the Perplexity provider out of the core tree
This manifest bundles and auto-loads a paid third-party SaaS integration in plugins/, making Hermes responsible for maintaining Perplexity's vendor API. Repository policy explicitly requires paid-service connectors to ship as standalone user or pip plugins rather than in this tree.
AGENTS.md reference: plugins/AGENTS.md:L23-L28
Useful? React with 👍 / 👎.
| def test_invalid_pid_inputs_do_not_crash(self): | ||
| assert _subprocess_compat.pid_is_hermes(-1) is False | ||
| assert _subprocess_compat.pid_is_hermes(0) is False | ||
| assert _subprocess_compat.pid_is_hermes("not-a-pid") is False | ||
| assert _subprocess_compat.pid_is_hermes(True) is False | ||
| with mock.patch.object(_subprocess_compat, "IS_WINDOWS", True): |
There was a problem hiding this comment.
Run PID ownership probes only on Windows
These Windows process-identity tests now run on non-Windows hosts by patching the module-level IS_WINDOWS flag, so they exercise the Linux interpreter and mocked helpers rather than the Windows subprocess behavior they are meant to protect. Restore the windows_only markers so the tests run in the actual Windows lane instead of providing false cross-platform confidence.
AGENTS.md reference: AGENTS.md:L351-L358
Useful? React with 👍 / 👎.
| file_operations holds Git-Bash-style paths there). ``HERMES_NATIVE_FILE_READ=0`` | ||
| turns the fast path off.""" |
There was a problem hiding this comment.
Keep the native-read override in config.yaml
Existing users with terminal.native_file_read: false now have that setting silently ignored because _native_read_enabled only checks HERMES_NATIVE_FILE_READ; on POSIX local environments their file operations unexpectedly switch back to the native path. This is behavioral configuration and must remain in config.yaml rather than requiring a process environment override.
AGENTS.md reference: AGENTS.md:L79-L82
Useful? React with 👍 / 👎.
| # A reconciliation PR may merge newer upstream history into the fork. | ||
| # First-parent keeps imported upstream authors out of the PR attribution set. | ||
| NEW_EMAILS=$(git log ${MERGE_BASE}..HEAD --first-parent --format='%ae' --no-merges | sort -u) |
There was a problem hiding this comment.
Do not hide merged contributor branches from attribution
When a PR contains a merge commit whose second parent is a contributor branch not already on main, --first-parent --no-merges ignores both the merge and every authored commit behind it, so the check reports no unmapped email and permits the contribution to land without credit. Exclude commits already reachable from origin/main instead of discarding all non-first-parent history, and apply the same correction to scripts/audit_pr_attribution.py.
AGENTS.md reference: AGENTS.md:L71-L72
Useful? React with 👍 / 👎.
| session.owner_task_id = to_owner | ||
| session.task_id = to_task_id | ||
| session.session_key = to_session_key | ||
| session.handoff_note = note |
There was a problem hiding this comment.
Persist handed-off process ownership before returning
A successful handoff mutates only the in-memory ProcessSession; it never rewrites processes.json. If the gateway restarts while a host-local or systemd-scoped handed-off process is still running, checkpoint recovery restores the old sa-* owner and child session route, so its eventual completion is suppressed as subagent-owned instead of reaching the parent that was told it now owns the process. Persist the updated ownership and routing fields as part of the successful transfer.
AGENTS.md reference: tools/AGENTS.md:L92-L98
Useful? React with 👍 / 👎.
| duration = run.elapsed() | ||
| entry = _build_result_entry(child, result, task_index, duration, schema) | ||
| run.append_sibling_write_reminder(entry) | ||
| run.account_background_processes(entry) |
There was a problem hiding this comment.
Account for child processes on failure paths
account_background_processes is invoked only after a child returns successfully. When await_child instead returns a timeout/error entry, or the surrounding path raises, _run_single_child returns without adding orphaned_processes or unread_completions and then cleanup terminates those processes, so the parent receives none of the runtime truth this feature promises and cannot relaunch needed work. Run the accounting step for failure entries before every cleanup path as well.
AGENTS.md reference: tools/AGENTS.md:L92-L96
Useful? React with 👍 / 👎.
| queued = {target: receipt for target, receipt in | ||
| job.get("_bot_chat_delivery_receipts", {}).items() | ||
| if receipt["status"] in ("queued", "claimed")} or None |
There was a problem hiding this comment.
Reconcile settled Bot Chat receipts after cron returns
For a live Bot Chat owner, the cron worker records a queued receipt and returns before the owner processes it, but no production path revisits that receipt afterward; the consumer only updates the mailbox JSON. Consequently last_delivery_queued and last_status=delivery_queued remain permanently stale even after the receipt becomes settled or failed (especially visible for one-shot or infrequent jobs), so hermes cron list continues claiming completion is unverified. Add a reconciliation path that reads terminal receipts and updates the persisted job state.
Useful? React with 👍 / 👎.
| from hermes_cli.shared_session_attach import configure_tui_attachment | ||
| try: | ||
| configure_tui_attachment(env, resume_session_id) |
There was a problem hiding this comment.
Implement the owner side before enabling auto-attach
For a resume targeting a session with a live owner, this newly wired path cannot succeed against any runtime built from this commit: a repo-wide search finds no production writer for metadata.shared_runtime_url and no /api/session-attach route, with both appearing only in this client and mocked tests. configure_tui_attachment therefore raises for every current Desktop/TUI owner and the launcher exits instead of attaching, leaving the cooperative fan-out flow dead on its real resolution chain. Implement the authenticated server advertisement and handshake before invoking the client path.
AGENTS.md reference: AGENTS.md:L275-L276
Useful? React with 👍 / 👎.
This prerequisite sync PR rebases the fork-owned fixes from merged #59 onto the current NousResearch/hermes-agent main tree.
git diff --checkpasses.Merge this before rebuilding the dependent stacked PRs.