Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
8757a8f
fix: align AI-recent notes with WebUI prefill hook
AJV20 May 28, 2026
571bb10
Merge remote-tracking branch 'origin/master' into webui-context-prefi…
AJV20 May 28, 2026
5f42e87
fix: skip stale repair for compression parents
ai-ag2026 May 28, 2026
f879fd6
fix: add dry-run discoverability safe repair
ai-ag2026 May 28, 2026
2ee2491
fix: defer streaming KaTeX for pending equations
ai-ag2026 May 28, 2026
9190ab4
Fix empty partial activity tail recency
May 28, 2026
ce59e7c
fix: defer stale stream repair for active workers
ai-ag2026 May 28, 2026
9e54039
fix(profiles): write API key to .env instead of config.yaml on profil…
gavinssr May 28, 2026
821d4a7
test: keep redaction fixture visible in session index
ai-ag2026 May 28, 2026
d77e8f0
test: update _write_endpoint_to_config tests for api_key→.env migration
gavinssr May 28, 2026
10573ab
Fix session media image rendering
May 28, 2026
1b5e6f6
fix: mirror WebUI prefill env for AI-recent notes
AJV20 May 28, 2026
9e69db9
fix: show cron sessions in project filter
AJV20 May 28, 2026
3469a2f
fix: avoid interruption marker for completed journal runs
ai-ag2026 May 28, 2026
e4ef50a
test: cover provider-neutral notes sources
AJV20 May 28, 2026
cbd3704
fix: preserve literal prefill script paths
AJV20 May 28, 2026
04e0f90
test: force master in git workspace fixtures
AJV20 May 28, 2026
8e6ed66
fix: clarify gateway chat auth errors
AJV20 May 28, 2026
790fc70
test: keep bare git fixtures on master
AJV20 May 28, 2026
923b719
fix: surface gateway auth errors in browser
AJV20 May 28, 2026
dc5b4b1
Merge PR #3037
May 28, 2026
007ba46
Merge PR #3048
May 28, 2026
11ea6c3
Merge PR #3060
May 28, 2026
83f8080
Merge PR #3053
May 28, 2026
c642c1e
Merge PR #3069
May 28, 2026
921b94a
Merge PR #3046
May 28, 2026
4412aea
Merge PR #3059
May 28, 2026
1c89c7d
Merge PR #3064
May 28, 2026
a3fc305
Merge PR #3077
May 28, 2026
371f77c
stage-batch36: stamp v0.51.154 / Release DZ
May 28, 2026
0a2dabc
stage-batch36: tighten #3064 MEDIA: token gate to non-user-role messages
May 28, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,26 @@

## [Unreleased]

## [v0.51.154] — 2026-05-28 — Release DZ (stage-batch36 — 9-PR medium-risk cleanup: cron project chip + KaTeX streaming + recovery + .env keys + discoverability repair + media MEDIA tokens + gateway 401 + notes prefill + cron filter)

### Added

- Session discoverability audit now has a default-dry-run `--repair-safe` routine for deterministic cleanup: stale persisted WebUI-as-CLI flags can be cleared from sidecars/index entries, and messageful WebUI rows present only in `state.db` can be materialized into sidecars/index entries when `--apply --backup-dir <dir>` is explicitly provided.

### Changed

- The third-party notes drawer's "Recently used by AI" list now follows the provider-neutral WebUI-specific `HERMES_WEBUI_PREFILL_MESSAGES_SCRIPT` / `webui_prefill_messages_script` hook when configured, including argv-style hooks such as `[python3, /path/to/recall.py]` and command strings such as `python3 /path/to/recall.py`, before falling back to the legacy generic `prefill_messages_script`. Configured third-party notes sources such as Joplin, Obsidian, Notion, and llm-wiki remain visible even before runtime tool inventory hydrates.

### Fixed

- Streaming KaTeX render passes now skip parser-owned equation placeholders that may still be receiving text, preventing long equations from being marked rendered before the final parser flush completes. (#2976)
- Cron sessions assigned to the dedicated Cron Jobs project now remain hidden from the default sidebar while still appearing when that project chip is selected.
- Compression parent sessions are no longer repaired as stale interrupted turns when a continuation already exists, preventing false "Response interrupted" markers and hidden continuation rows after auto-compression session rotation. (Refs #2361)
- Empty partial activity rows preserved from cancelled turns no longer define sidebar recency, anchor the initial paginated message window, or get restored after newer completed turns. Long sessions with old activity-only partials after recent replies now stay grouped by their latest real message and open on the recent readable transcript. (#3057)
- Local `MEDIA:` image tokens in chat history now include the current session id and can render exact image paths already present in that session transcript, so agent-generated artifacts outside the active workspace no longer show as broken thumbnails while arbitrary local paths remain blocked.
- Gateway-backed browser chat now turns Gateway API Server 401s into a specific `gateway_auth_error` explaining that `HERMES_WEBUI_GATEWAY_API_KEY` must match `API_SERVER_KEY`, instead of surfacing the Gateway's generic "Invalid API key" body as if the model provider key failed. The browser error renderer recognizes this event type as "Gateway authentication failed" instead of falling back to a generic "Error" heading. `/api/health/agent` also reports redacted gateway-chat configuration status (`enabled`, backend, base URL configured, API key configured) as an operator diagnostic payload; it is not currently rendered as a user-facing health banner.
- New profiles with an API key supplied at create time now write the key to the profile's `.env` under the correct provider-specific variable (e.g. `KIMI_API_KEY`, `DEEPSEEK_API_KEY`) at mode 0o600, instead of writing it to `config.yaml` where Hermes Agent never reads it.

## [v0.51.153] — 2026-05-28 — Release DY (stage-batch35 — 11-PR low-risk cleanup: title-language + clarify SSE + upload filename + discoverability + SSE reconnect + gateway image + docker docs)

### Changed
Expand Down
9 changes: 8 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -180,7 +180,14 @@ HERMES_WEBUI_GATEWAY_API_KEY=... \
`api_server`, or `api-server` enable the bridge. Generic truthy values such as
`1` or `true` are ignored so existing deployments do not change execution
ownership accidentally. If `HERMES_WEBUI_GATEWAY_API_KEY` is omitted, WebUI falls
back to `API_SERVER_KEY` when present.
back to `API_SERVER_KEY` when present. When Gateway returns HTTP 401, WebUI
reports a `gateway_auth_error` that points at this WebUI↔Gateway key mismatch
rather than showing the Gateway's generic provider-style "Invalid API key" body.
`/api/health/agent` also includes a redacted `gateway_chat` block so operators can
see whether gateway mode, base URL, and API-key presence are configured without
exposing the key value. That `gateway_chat` field is an operator diagnostic
payload only; it is not currently rendered as a user-facing health banner in the
browser UI.

The bridge is best used by operators who already run Hermes Gateway/API Server
locally and want browser-originated chat to use the same runtime/tool path as
Expand Down
45 changes: 38 additions & 7 deletions api/gateway_chat.py
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,40 @@ def _gateway_api_key(environ: dict[str, str] | None = None) -> str:
).strip()


def gateway_chat_config_status(config_data=None, environ: dict[str, str] | None = None) -> dict:
"""Return redacted Gateway-backed chat configuration status."""
mode = webui_chat_backend_mode(config_data, environ)
base_url = _gateway_base_url(config_data, environ)
return {
"enabled": mode == "gateway",
"backend": mode,
"base_url_configured": bool(base_url),
"api_key_configured": bool(_gateway_api_key(environ)),
}


def _gateway_http_error_event(exc: urllib.error.HTTPError, err_body: str, *, api_key_configured: bool) -> dict:
safe = _redact_text(err_body or str(exc))[:500]
if exc.code == 401:
return {
"label": "Gateway authentication failed",
"type": "gateway_auth_error",
"message": "Gateway rejected the WebUI API key (HTTP 401).",
"hint": (
"Set HERMES_WEBUI_GATEWAY_API_KEY to the same value as the Hermes Gateway "
"API_SERVER_KEY, or disable HERMES_WEBUI_CHAT_BACKEND=gateway."
if not api_key_configured
else "Check that HERMES_WEBUI_GATEWAY_API_KEY matches the Hermes Gateway API_SERVER_KEY."
),
}
return {
"label": "Gateway request failed",
"type": "gateway_http_error",
"message": f"Gateway returned HTTP {exc.code}.",
"hint": safe or "Check the configured Gateway API server.",
}


def _gateway_sse_delta(payload: dict) -> str:
"""Extract assistant text from an OpenAI-compatible streaming chunk."""
try:
Expand Down Expand Up @@ -296,13 +330,10 @@ def put_gateway_event(event, data):
err_body = exc.read(2048).decode("utf-8", errors="replace")
except Exception:
err_body = ""
safe = _redact_text(err_body or str(exc))[:500]
put_gateway_event("apperror", {
"label": "Gateway request failed",
"type": "gateway_http_error",
"message": f"Gateway returned HTTP {exc.code}.",
"hint": safe or "Check the configured Gateway API server.",
})
put_gateway_event(
"apperror",
_gateway_http_error_event(exc, err_body, api_key_configured=bool(_gateway_api_key())),
)
except Exception as exc:
safe = _redact_text(str(exc))[:500]
put_gateway_event("apperror", {
Expand Down
192 changes: 181 additions & 11 deletions api/models.py
Original file line number Diff line number Diff line change
Expand Up @@ -369,12 +369,35 @@ def _message_timestamp(message):
return None


def _is_empty_partial_activity_message(message):
"""Return True for cancelled/recovered activity rows with no reply text."""
if not isinstance(message, dict):
return False
if message.get('role') != 'assistant' or not message.get('_partial'):
return False
content = message.get('content', '')
if isinstance(content, str):
return not content.strip()
if isinstance(content, list):
for part in content:
if isinstance(part, dict):
if part.get('type') == 'text' and str(part.get('text') or part.get('content') or '').strip():
return False
continue
if str(part or '').strip():
return False
return True
return not str(content or '').strip()


def _last_message_timestamp(messages):
if not isinstance(messages, list):
return None
for message in reversed(messages):
if isinstance(message, dict) and message.get('role') == 'tool':
continue
if _is_empty_partial_activity_message(message):
continue
ts = _message_timestamp(message)
if ts:
return ts
Expand Down Expand Up @@ -1175,6 +1198,19 @@ def _run_journal_has_visible_output(session, stream_id: str | None) -> bool:
return False


def _run_journal_terminal_state(session, stream_id: str | None) -> str | None:
if not stream_id:
return None
try:
from api.run_journal import latest_run_summary
summary = latest_run_summary(session.session_id, stream_id)
except Exception:
return None
if not summary.get('terminal'):
return None
return str(summary.get('terminal_state') or '') or None


def _journal_is_still_arriving(session, stream_id: str | None) -> bool:
"""Return True for journals that may become visible on a later read.

Expand Down Expand Up @@ -1689,13 +1725,35 @@ def _apply_core_sync_or_error_marker(
_pending_text = " ".join(str(session.pending_user_message or "").split())
_already_checkpointed = False
if _pending_text and session.messages:
_last_msg = session.messages[-1]
if isinstance(_last_msg, dict) and _last_msg.get('role') == 'user':
_last_text = " ".join(str(_last_msg.get('content') or "").split())
_already_checkpointed = _last_text == _pending_text
for _last_msg in reversed(session.messages):
if isinstance(_last_msg, dict) and _last_msg.get('role') == 'user':
_last_text = " ".join(str(_last_msg.get('content') or "").split())
_already_checkpointed = _last_text == _pending_text
break
_recovered_ts = int(time.time())
if isinstance(session.pending_started_at, (int, float)) and session.pending_started_at > 0:
_recovered_ts = int(session.pending_started_at)
_stream_id = stream_id_for_recheck or session.active_stream_id
_pending_started_at = session.pending_started_at
if _run_journal_terminal_state(session, _stream_id) == 'completed':
if not _already_checkpointed:
_append_recovered_pending_turn(session, timestamp=_recovered_ts)
_append_journaled_partial_output(
session,
_stream_id,
dedupe_existing=True,
)
session.active_stream_id = None
session.pending_user_message = None
session.pending_attachments = []
session.pending_started_at = None
session.save(touch_updated_at=touch_updated_at)
logger.info(
"Session %s: cleared stale pending state for completed stream %s without error marker",
sid,
_stream_id,
)
return True
if not _already_checkpointed:
_append_recovered_pending_turn(session, timestamp=_recovered_ts)
else:
Expand All @@ -1709,10 +1767,8 @@ def _apply_core_sync_or_error_marker(
_append_recovered_turn_to_context(session, recovered)
recovered_output = _append_journaled_partial_output(
session,
stream_id_for_recheck or session.active_stream_id,
_stream_id,
)
_stream_id = stream_id_for_recheck or session.active_stream_id
_pending_started_at = session.pending_started_at
session.active_stream_id = None
session.pending_user_message = None
session.pending_attachments = []
Expand Down Expand Up @@ -1850,6 +1906,72 @@ def _apply_core_sync_or_error_marker(
_REPAIR_STALE_PENDING_GRACE_SECONDS = 30


def _has_compression_continuation(session) -> bool:
"""Return True when ``session`` is an archived compression parent.

Context compression rotates the live WebUI session id: the old sidecar is
preserved for lineage while the new child owns the running/completed turn.
Stale-pending repair must not append an interruption marker to that old
parent just because its stream bookkeeping disappeared after the rotation.
"""
sid = getattr(session, 'session_id', None)
if not sid:
return False

def _row_is_continuation(row) -> bool:
if not isinstance(row, dict):
return False
child_sid = row.get('session_id')
if not child_sid or child_sid == sid:
return False
if row.get('parent_session_id') != sid:
return False
# Any child row is enough evidence that this pending state belongs to a
# compression lineage, not a dead standalone turn. The child may itself
# temporarily carry a bad pre_compression_snapshot flag from older code;
# do not filter it out here or the guard misses the exact regression.
return True

try:
with LOCK:
for child in SESSIONS.values():
if getattr(child, 'session_id', None) == sid:
continue
if getattr(child, 'parent_session_id', None) == sid:
return True
except Exception:
pass

try:
if SESSION_INDEX_FILE.exists():
entries = json.loads(SESSION_INDEX_FILE.read_text(encoding='utf-8'))
if isinstance(entries, list) and any(_row_is_continuation(e) for e in entries):
return True
except Exception:
logger.debug("Failed to inspect session index for compression continuation", exc_info=True)

# Index rows can lag behind rapid compression/save races. Fall back to a
# shallow JSON metadata scan; session files write parent_session_id before
# the messages array, so this avoids loading multi-MB transcripts.
try:
needle = f'"parent_session_id": "{sid}"'
for path in SESSION_DIR.glob('*.json'):
if path.name.startswith('_') or path.stem == sid:
continue
try:
head = path.read_text(encoding='utf-8', errors='ignore')[:4096]
except TypeError:
head = path.read_text(encoding='utf-8')[:4096]
except OSError:
continue
if needle in head:
return True
except Exception:
logger.debug("Failed to scan session files for compression continuation", exc_info=True)

return False


def _repair_stale_pending(session) -> bool:
"""Recover a sidecar stuck with messages=[] and stale pending state.

Expand All @@ -1872,6 +1994,18 @@ def _repair_stale_pending(session) -> bool:
or not _seen_stream_id
or _seen_stream_id in _active_stream_ids()):
return False
if getattr(session, 'pre_compression_snapshot', False):
logger.debug(
"_repair_stale_pending: skipping pre-compression snapshot %s",
getattr(session, 'session_id', '?'),
)
return False
if _has_compression_continuation(session):
logger.debug(
"_repair_stale_pending: skipping compression parent %s with continuation",
getattr(session, 'session_id', '?'),
)
return False

# Grace-period guard: bail if the turn is too fresh to be a real crash.
# Falsy pending_started_at (None, 0, missing) means "old enough" — preserve
Expand Down Expand Up @@ -2228,6 +2362,38 @@ def _is_intentionally_background_sidebar_session(session: dict) -> bool:
return source == 'cron' or sid.startswith('cron_')


def _include_project_hidden_background_sidebar_sessions(
candidates: list[dict],
visible: list[dict],
) -> list[dict]:
"""Keep project-assigned background sessions addressable by project chips.

Cron sessions stay hidden from the default sidebar, but if they have a
project assignment they must still be present in the client cache so the
dedicated project chip can reveal them (#3019).
"""
visible_ids = {
str(session.get('session_id'))
for session in visible
if session.get('session_id')
}
out = list(visible)
for session in candidates:
sid = str(session.get('session_id') or '')
if not sid or sid in visible_ids:
continue
if not _is_intentionally_background_sidebar_session(session):
continue
if not session.get('project_id'):
continue
if _sidebar_message_count(session) <= 0:
continue
row = dict(session)
row['default_hidden'] = True
out.append(row)
return out


def _preserve_messageful_sidebar_discoverability(
candidates: list[dict],
visible: list[dict],
Expand Down Expand Up @@ -2503,8 +2669,10 @@ def all_sessions(diag=None):
and not s.get('worktree_path')
)]
result = _prefer_fuller_snapshots_for_sidebar(result)
visible_result = [s for s in result if not _hide_from_default_sidebar(s)]
result = _preserve_messageful_sidebar_discoverability(result, visible_result)
sidebar_candidates = result
visible_result = [s for s in sidebar_candidates if not _hide_from_default_sidebar(s)]
result = _preserve_messageful_sidebar_discoverability(sidebar_candidates, visible_result)
result = _include_project_hidden_background_sidebar_sessions(sidebar_candidates, result)
_strip_sidebar_internal_flags(result)
# Backfill: sessions created before Sprint 22 have no profile tag.
# Attribute them to 'default' so the client profile filter works correctly.
Expand Down Expand Up @@ -2542,8 +2710,10 @@ def all_sessions(diag=None):
and not getattr(s, 'worktree_path', None)
)]
result = _prefer_fuller_snapshots_for_sidebar(result)
visible_result = [s for s in result if not _hide_from_default_sidebar(s)]
result = _preserve_messageful_sidebar_discoverability(result, visible_result)
sidebar_candidates = result
visible_result = [s for s in sidebar_candidates if not _hide_from_default_sidebar(s)]
result = _preserve_messageful_sidebar_discoverability(sidebar_candidates, visible_result)
result = _include_project_hidden_background_sidebar_sessions(sidebar_candidates, result)
_strip_sidebar_internal_flags(result)
for s in result:
if not s.get('profile'):
Expand Down
Loading
Loading