Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions FORK_NOTES.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ This document explains the fork-specific changes on `main` that diverge from ups

| ID | Target file | What it does | Why we need it | Upstream status |
|---|---|---|---|---|
| **P-055** | `agent/memory_provider.py`, `hermes_cli/web_server.py`, `tools/memory_tool.py`, `plugins/memory/openviking/__init__.py`, `plugins/memory/hindsight/__init__.py`, focused tests | Adds an optional read-only runtime-status hook and `GET /api/memory/providers/{name}/status`; OpenViking and Hindsight aggregate health, readiness, component/model/queue/memory or bank/runtime metrics without returning secrets. Provider config saves accept `activate:false`, while omitted `activate` keeps the existing save-and-activate contract. Direct profile-scoped OpenViking endpoints now participate in connection resolution and availability checks. Built-in `MEMORY.md` configuration now publishes and enforces a 1–8000 character range, with the existing 2200-character default. | The managed Desktop runtime has a profile-isolated `HERMES_HOME`; its memory UI could only activate providers and never persisted connection fields, so OpenViking remained `needs_config` and neither backend had in-app observability. Its built-in memory editor also needs one bounded config contract shared with Core. | Generic provider observability/configuration surface; suitable for upstream after UI review. |
| **P-054** | `model_tools.py`, `tools/environments/base.py`, `tools/environments/local.py`, `tools/environments/modal_utils.py`, `tools/terminal_tool.py`, `tui_gateway/cli_delegation.py`, `tui_gateway/server.py`, Claude Code/Codex skills and focused tests | Extends P-047 to foreground terminal delegations: the agent dispatcher preserves `tool_call_id` through registry dispatch, pipe-backed environments expose a fail-open live-output callback keyed by that id, and the gateway coalesces foreground and background chunks into `delegation.cli.output` at ≤2 Hz. Plain-text output becomes redacted `raw` progress events; completion aggregates model/session/workdir and Token usage from Claude/Codex JSON, with a defensive parser for Codex's human-readable header and `tokens used` footer. Terminal results now include the shell's resolved final cwd, dynamic shell expressions are no longer reported as a literal `$`, and bundled skills request stream-json/`--json` in foreground mode. | A foreground Claude Code/Codex call previously emitted only started/completed, leaving the Desktop spinner silent for minutes; if the completed event was missed the row stayed running forever, while cwd and Token fields remained empty (notably when an older Codex skill omitted `--json`). | Generic terminal/gateway observability and suitable for upstream; event names and classifier remain CN Desktop-driven. Related: P-047. |
| **P-053** | `scripts/update_thirdparty.py`, `tests/scripts/test_update_thirdparty.py` | Adds a maintenance script that checks GitHub releases for **ripgrep** and **rtk**, updates *every* pinned occurrence across `scripts/install.sh`, `scripts/install.ps1`, `tools/rtk_provision.py`, and `.github/workflows/tests.yml` (which pins ripgrep twice — the `test` and `e2e` jobs), recomputes the CI SHA256 for ripgrep, and supports opt-in Chinese mirror fallback via `--china-mirror` / `--mirror` / `HERMES_THIRDPARTY_MIRROR`. | Hermes-CN pins these external binaries in multiple file formats (Bash, PowerShell, Python, YAML), and direct GitHub access is often slow or blocked in mainland China; a single updater keeps versions consistent and provides mirror fallback. | CN-specific maintenance tool; could be upstreamed after introducing `_ripgrep_common.py` / `_rtk_common.py` |
| **P-025** | `hermes_cli/web_server.py` | `/api/providers/oauth` now (1) serves from a 20s per-profile in-process TTL cache, (2) runs each provider's status check concurrently via `asyncio.to_thread` (OFF the FastAPI event loop) instead of serially inline, and (3) busts the cache on every connect/disconnect (disconnect clear paths, PKCE submit, device-code/loopback poll→`approved`). Adds a `refresh=true` escape hatch. | The desktop Models page enumerated every OAuth provider's status serially on every open AND every window refocus; some checks touch the network/subprocess, and because the handler is `async` they blocked the event loop that also serves the chat gateway WebSocket — so 模型页 took seconds to open and could stutter live chat. | Should be upstreamed (generic responsiveness fix) |
Expand Down
1 change: 1 addition & 0 deletions FORK_NOTES.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@

| ID | 目标文件 | 做了什么 | 为什么需要 | 上游状态 |
|---|---|---|---|---|
| **P-055** | `agent/memory_provider.py`、`hermes_cli/web_server.py`、`tools/memory_tool.py`、`plugins/memory/openviking/__init__.py`、`plugins/memory/hindsight/__init__.py`、聚焦测试 | 新增可选只读运行状态 hook 与 `GET /api/memory/providers/{name}/status`;OpenViking/Hindsight 分别聚合健康、就绪、组件、模型、队列、记忆或 bank/runtime 指标,且不返回秘密。Provider 配置保存支持 `activate:false`,省略时仍保持原有“保存并启用”兼容行为。直接写入档案 `memory.openviking` 的 endpoint 也会参与连接解析与可用性判断。内置 `MEMORY.md` 配置继续默认 2200 字符,并统一发布和执行 1–8000 字符范围。 | Desktop managed runtime 使用隔离的 `HERMES_HOME`;旧记忆页只会激活、不会保存连接字段,导致 OpenViking 一直是 `needs_config`,两个后端也都无法在 Hermes 内监控;内置记忆编辑器也需要与 Core 共享同一套容量边界。 | 通用 provider 可观测与配置接口,待 UI 评审后适合上游。 |
| **P-054** | `model_tools.py`、`tools/environments/base.py`、`tools/environments/local.py`、`tools/environments/modal_utils.py`、`tools/terminal_tool.py`、`tui_gateway/cli_delegation.py`、`tui_gateway/server.py`、Claude Code/Codex 技能及聚焦测试 | 在 P-047 上补齐前台 terminal 委派:agent 调度层把 `tool_call_id` 透传到工具注册表,pipe 型环境再按该 id 提供失败开放的实时输出回调;gateway 将前后台 chunk 都以 ≤2 Hz 合并为 `delegation.cli.output`。普通文本输出归一化为脱敏 `raw` 进度事件;终态从 Claude/Codex JSON 汇总 model/session/workdir 与 Token,并兼容解析 Codex 人类可读头部和 `tokens used` 尾部。terminal 结果补充 shell 执行后的真实 cwd,动态 shell 表达式不再被错误展示为字面量 `$`,内置技能在前台也优先请求 stream-json/`--json`。 | 前台 Claude Code/Codex 过去只有 started/completed,Desktop 会静默转圈数分钟;completed 事件一旦丢失就永久显示运行中,而且旧 Codex 技能未加 `--json` 时目录与 Token 始终为空。 | 通用 terminal/gateway 可观测性可上游;事件名与分类器仍由 CN Desktop 驱动。关联 P-047。 |
| **P-053** | `scripts/update_thirdparty.py`、`tests/scripts/test_update_thirdparty.py` | 新增维护脚本,检查 **ripgrep** 与 **rtk** 的 GitHub 最新 release,自动同步 `scripts/install.sh`、`scripts/install.ps1`、`tools/rtk_provision.py`、`.github/workflows/tests.yml` 中**每一处**版本钉子(tests.yml 里 `test` 与 `e2e` 两个 job 各钉一次 ripgrep),并重新计算 ripgrep 在 CI 工作流里的 SHA256;支持 `--china-mirror` / `--mirror` / `HERMES_THIRDPARTY_MIRROR` 显式开启的国内镜像回退。 | Hermes-CN 把这些外部二进制版本钉在多种文件格式(Bash、PowerShell、Python、YAML)中,且国内访问 GitHub 经常慢或不通;统一 updater 可保持多文件版本一致,并提供镜像回退。 | CN 专属维护工具;待上游引入 `_ripgrep_common.py` / `_rtk_common.py` 后可考虑 upstream |
| **P-025** | `hermes_cli/web_server.py` | `/api/providers/oauth` 现在:(1) 命中 20s 的按 profile 进程内 TTL 缓存;(2) 用 `asyncio.to_thread` 并发跑各 provider 的状态检查(移出 FastAPI 事件循环),不再串行内联;(3) 在每次连接/断开时失效缓存(断开的两条清理路径、PKCE submit、设备码/loopback 轮询到 `approved`)。另加 `refresh=true` 逃生阀。 | 桌面端模型页每次打开、以及每次窗口重新聚焦都会串行枚举所有 OAuth provider 的状态;部分检查会联网/起子进程,而该 handler 是 `async`,于是阻塞了同时服务聊天网关 WebSocket 的事件循环——模型页要等好几秒,还会拖累实时会话。 | 建议上游(通用响应性修复) |
Expand Down
11 changes: 11 additions & 0 deletions agent/memory_provider.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@
on_pre_compress(messages) -> str — extract before context compression
on_memory_write(action, target, content, metadata=None) — mirror built-in memory writes
on_delegation(task, result, **kwargs) — parent-side observation of subagent work
get_runtime_status() -> dict | None — read-only backend health and metrics
backup_paths() -> list[str] — extra on-disk paths to include in `hermes backup`
"""

Expand Down Expand Up @@ -260,6 +261,16 @@ def get_config_schema(self) -> List[Dict[str, Any]]:
"""
return []

def get_runtime_status(self) -> Optional[Dict[str, Any]]:
"""Return a read-only backend status snapshot for the dashboard.

Providers that support monitoring return common connection fields plus
a provider-specific ``details`` payload. The dashboard executes this
hook off the event loop. Returning ``None`` keeps existing providers
compatible and explicitly marks runtime monitoring as unsupported.
"""
return None

def save_config(self, values: Dict[str, Any], hermes_home: str) -> None:
"""Write non-secret config to the provider's native location.

Expand Down
33 changes: 27 additions & 6 deletions agent/models_dev.py
Original file line number Diff line number Diff line change
Expand Up @@ -568,7 +568,12 @@ class ModelCapabilities:

supports_tools: bool = True
supports_vision: bool = False
supports_pdf: bool = False
supports_audio: bool = False
supports_video: bool = False
supports_reasoning: bool = False
supports_reasoning_control: bool = False
open_weights: bool = False
context_window: int = 200000
max_output_tokens: int = 8192
model_family: str = ""
Expand Down Expand Up @@ -624,12 +629,14 @@ def get_model_capabilities(
cache/snapshot only (non-blocking) for hot paths like ``/api/model/info``.

Extracts from model entry fields:
- reasoning (bool) → supports_reasoning
- tool_call (bool) → supports_tools
- attachment (bool) → supports_vision
- limit.context (int) → context_window
- limit.output (int) → max_output_tokens
- family (str) → model_family
- reasoning → supports_reasoning
- reasoning_options → supports_reasoning_control
- tool_call → supports_tools
- modalities.input → vision/PDF/audio/video support
- open_weights → open_weights
- limit.context → context_window
- limit.output → max_output_tokens
- family → model_family
"""
models = _get_provider_models(provider, allow_network=allow_network)
if models is None:
Expand All @@ -653,7 +660,16 @@ def get_model_capabilities(
supports_vision = "image" in input_mods
else:
supports_vision = bool(entry.get("attachment", False))
input_modality_set = set(input_mods) if isinstance(input_mods, list) else set()
supports_pdf = "pdf" in input_modality_set
supports_audio = "audio" in input_modality_set
supports_video = "video" in input_modality_set
supports_reasoning = bool(entry.get("reasoning", False))
reasoning_options = entry.get("reasoning_options")
supports_reasoning_control = (
isinstance(reasoning_options, list) and len(reasoning_options) > 0
)
open_weights = bool(entry.get("open_weights", False))

# Extract limits
limit = entry.get("limit", {})
Expand All @@ -671,7 +687,12 @@ def get_model_capabilities(
return ModelCapabilities(
supports_tools=supports_tools,
supports_vision=supports_vision,
supports_pdf=supports_pdf,
supports_audio=supports_audio,
supports_video=supports_video,
supports_reasoning=supports_reasoning,
supports_reasoning_control=supports_reasoning_control,
open_weights=open_weights,
context_window=context_window,
max_output_tokens=max_output_tokens,
model_family=model_family,
Expand Down
2 changes: 1 addition & 1 deletion agent/models_dev_snapshot.json

Large diffs are not rendered by default.

39 changes: 31 additions & 8 deletions hermes_cli/inventory.py
Original file line number Diff line number Diff line change
Expand Up @@ -147,10 +147,10 @@ def build_models_payload(
show $/Mtok columns and gate paid models on free accounts —
mirroring the ``hermes model`` CLI picker. Adds network calls
(pricing fetch + Nous tier check); only set for interactive pickers.
- ``capabilities``: add a per-row ``capabilities`` map
``{model: {fast, reasoning}}`` so pickers can gate the model-options
controls (fast toggle / reasoning) to what each model actually
supports, instead of offering knobs the backend would reject.
- ``capabilities``: add a per-row ``capabilities`` map sourced from the
bundled models.dev snapshot. Besides the legacy ``fast``/``reasoning``
gates, each known model includes tools, vision, reasoning, context-window,
output-token, and family metadata for richer desktop picker tags.
- ``force_fresh_nous_tier``: bypass the short Nous free-tier cache when
selecting Portal-recommended Nous models and applying tier gating. Keep
this false for UI picker opens; explicit auth/model flows can opt in
Expand Down Expand Up @@ -264,13 +264,15 @@ def build_models_payload(


def _apply_capabilities(rows: list[dict]) -> None:
"""Attach a ``{model: {fast, reasoning}}`` map to each provider row.
"""Attach models.dev-backed capability metadata to each provider row.

`fast` mirrors ``model_supports_fast_mode`` (the same gate the runtime
enforces). `reasoning` comes from the models.dev catalog when known and
defaults to True otherwise — the effort dial is broadly accepted and a
no-op on models that ignore it, whereas hiding it from a capable-but-
uncatalogued model is the worse failure.
uncatalogued model is the worse failure. The richer fields are emitted only
when models.dev knows the model, allowing older/static desktop catalogs to
remain a fallback for custom or newly-added model IDs.
"""
from hermes_cli.models import model_supports_fast_mode

Expand All @@ -281,22 +283,43 @@ def _apply_capabilities(rows: list[dict]) -> None:

for row in rows:
slug = row.get("slug") or ""
caps: dict[str, dict[str, bool]] = {}
caps: dict[str, dict[str, object]] = {}

for model in row.get("models") or []:
reasoning = True
meta = None
if get_model_capabilities is not None and slug:
try:
meta = get_model_capabilities(slug, model)
if meta is None and str(slug).lower().startswith("custom:"):
canonical_slug = str(slug).split(":", 1)[1].strip().lower()
meta = get_model_capabilities(canonical_slug, model)
if meta is not None:
reasoning = bool(meta.supports_reasoning)
except Exception:
reasoning = True

caps[model] = {
model_caps: dict[str, object] = {
"fast": bool(model_supports_fast_mode(model)),
"reasoning": reasoning,
}
if meta is not None:
model_caps.update({
"supports_tools": bool(meta.supports_tools),
"supports_vision": bool(meta.supports_vision),
"supports_pdf": bool(meta.supports_pdf),
"supports_audio": bool(meta.supports_audio),
"supports_video": bool(meta.supports_video),
"supports_reasoning": bool(meta.supports_reasoning),
"supports_reasoning_control": bool(
meta.supports_reasoning_control
),
"open_weights": bool(meta.open_weights),
"context_window": int(meta.context_window),
"max_output_tokens": int(meta.max_output_tokens),
"model_family": str(meta.model_family or ""),
})
caps[model] = model_caps

row["capabilities"] = caps

Expand Down
Loading
Loading