Skip to content

fix(desktop): stop live-sync duplicating every message on send - #64564

Closed
Kyzcreig wants to merge 517 commits into
NousResearch:mainfrom
ANG-Ventures:fix/desktop-optimistic-stamp-dup
Closed

Kyzcreig wants to merge 517 commits into
NousResearch:mainfrom
ANG-Ventures:fix/desktop-optimistic-stamp-dup

Conversation

@Kyzcreig

Copy link
Copy Markdown

Symptom

In the desktop app (remote mode), every message in a session renders twice — reproducible by opening a new session and sending a single message. The duplication is display-only: the session DB stores each message exactly once (verified: N rows, N distinct contents). A reload shows the transcript correctly.

Root cause (two defects on the optimistic-stamp seam)

The desktop renders a sent turn optimistically with client-minted ids (user-<ts> / assistant-<ts>) and relies on the message.complete frame carrying the committed DB row ids to stamp those optimistic rows to their persisted integer ids. The live session-sync poll (session.changes) then recognizes the committed rows by id and skips them. Two defects broke that contract, so the poll re-appended every committed row as a duplicate:

  1. Server (tui_gateway/server.py) — the message.complete payload never included the committed ids ({text, usage, status, +reasoning/warning/rendered} only). So extractCommittedMessageIds() returned [], markTurnComplete() early-returned, and no stamping happened.

  2. Client (use-session-changes.ts) — even with ids present, the stamp predicate only matched user-prefixed or pending rows. But completeAssistantMessage() clears the streamed assistant row's pending flag before markTurnComplete() runs, so the completed assistant row (id assistant-<ts>, pending:false) was skipped → kept its optimistic id → duplicated by the poll. The existing unit test masked this by passing pending:true, a state that doesn't occur at runtime.

Fix

  • Server: add _completed_turn_committed_ids(history) — extracts [user_id, assistant_id] for the finished turn from session["history"] via the _db_persisted_row_id each message carries (stamped in run_agent._flush_messages_to_session_db) — and ship it as payload["message_ids"]. Tool rows are excluded (they're not optimistic client rows; the poll handles them by their own ids), matching the client's two-optimistic-row contract.
  • Client: recognize any optimistic id (user-/assistant- prefix) or a still-pending row in both stampOptimisticTranscriptRows and optimisticTranscriptIds.

Tests (fail on the naive code, pass on the fix)

  • tests/tui_gateway/test_completed_turn_committed_ids.py (6): order, tool-loop turn (returns [user, final-assistant-text]), empty-tool-calls-row exclusion, missing/non-int id skipping, invalid input.
  • use-session-changes.test.ts (2 new): an already-completed (pending:false) assistant row still stamps; end-to-end no-duplicate when the poll returns the own committed rows.

Verified RED→GREEN on both halves; existing suites green (session_changes 7, finalize_persist 11, desktop chat-hooks + message-stream 38). Server helper additionally verified against real session data (returns [user_id, assistant_id] in order from an actual completed desktop turn).

Kyzcreig and others added 30 commits July 1, 2026 07:35
…nfig override) (#158)

* feat(desktop): parseDesktopBackendRoot pinned-grammar config reader (P1.1, AC-4)

No-YAML-dep targeted scalar read of desktop.backend_root with indentation-anchored
grammar: nested-under-desktop only, comment strip, quote unwrap, ~/\$HOME expand,
first-wins; rejects tab-indent/flow-style/multi-doc -> null (fail-safe). 18/18.

* feat(desktop): backend runs from runtime tree by default (P1.1-1.5, AC-1/3/4/5)

resolveBackendRoot 3-tier precedence (config override > runtime > dev) wired into
createActiveBackend; parseDesktopBackendRoot no-YAML-dep config read; ready-file emits
project_root (AC-3 effect gate, fail-closed); shared fallback surface (AC-5). No live
behavior change on installs without a runtime tree. 313/314 desktop mechanics green.

* fix(desktop): reason-accurate backend-fallback log (override-unusable may still land on runtime tree)

* fix(desktop): honor leading --- YAML document-start marker in backend_root config reader

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…n guard) (#159)

* test(telegram): lock restart/reconnect message-loss parity (regression guard)

Discord had a drain-window message-loss bug (PR #157) because it is a *push*
transport. Telegram is a *pull* transport (getUpdates long-poll) with a
server-side queue, so downtime is backfilled for free — but only because of
three independent properties a refactor could each silently break. Lock them
as behavior contracts so the loss bug can never appear on Telegram.

INV-1  reconnect(is_reconnect=True) preserves the queue (drop_pending_updates=
       False); cold boot drops it. (positive + negative control)
INV-2  every start_polling recovery ladder preserves the queue unconditionally
       (AST source-contract; bootstrap-vs-recovery discriminated by enclosing
       function, not line number; MISSING kwarg = explicit FAIL).
INV-3a disconnect() drains in order updater.stop -> app.stop -> app.shutdown.
INV-3b PTB-provenance guard: Application.stop() in the installed PTB actually
       calls update_queue.join(); fails loudly (naming version) on a PTB bump.

Test-only + doc note; no production code change. All 6 tests mutation-proven to
gate (see PR body). Ground-truthed against pinned PTB 22.7.

SPEC: ~/.hermes/plans/2026-07-01_telegram-parity-regression-SPEC.md
Spike: ~/.hermes/plans/2026-07-01_telegram-parity-SPIKE.md

* test(telegram): harden mock guard + PTB provenance walk (Greptile follow-ups)

Two Greptile correctness nits (Confidence 4/5, 'safe to merge'):
- _ensure_telegram_mock: hasattr(__file__) is True for a MagicMock too, so a
  prior test's mock in sys.modules was misread as 'real telegram present'.
  Now distinguishes a genuine module (real str __file__, not a MagicMock).
- PTB provenance test: outer walk didn't stop at the first Application class
  and could match a nested/inner class. Iterate module-level statements only
  and break at the first match.

Both mutation-re-verified to still gate; 6 passed; ruff clean.

---------

Co-authored-by: Apollo <apollo@ang.ventures>
…pe B) (#160)

Telegram is loss-safe (#159) but had a KNOWN-OPEN residual: a hard kill
(SIGKILL/OOM/power) after the agent answered an update but before PTB confirmed
its offset makes Telegram re-deliver the answered update on next boot -> the
agent answers it twice. This guard suppresses a proven-duplicate re-delivery,
exactly once, and fails OPEN in every uncertain case (never drops a genuine
message -- suppressing on uncertainty would convert a rare dup into a loss).

Two Telegram id spaces, NEVER interchanged (SPEC section 0):
- update_id  (platform_update_id): envelope seq, used ONLY to SCOPE the guard
  to genuine candidate re-deliveries (update_id <= persisted HWM).
- message_id (stamped platform_message_id): used ONLY for the answerability
  lookup (was this message already answered?).

Design:
- gateway/telegram_redelivery.py: TelegramHwmTracker (in-memory HWM, coalesced
  <=30s + shutdown flush, atomic 0600, fail-open read), in_redelivery_scope
  (HWM-primary, 120s time-cap fallback only when no HWM), decide_redelivery
  (the fail-open decision table), RedeliverySuppressionCounter (content-free
  PHASE=tg_redelivery_suppressed observability).
- gateway/run.py: _is_telegram_boot_redelivered_duplicate wired into the
  general message path after session resolution, before the agent run; HWM
  flush on graceful stop; INV-6 companion-row persist for aggregated turns.
- telegram/adapter.py: _enqueue_text_event accumulates ALL constituent
  message_ids of a buffered multi-update turn.
- INV-6: aggregated turns stamp NON-first constituents as observed=1 companion
  rows (reuses shipped observed column + exact-match primitive, no schema
  change, cache-safe) so a re-delivery of any constituent is answerable.

Config: telegram.redelivery_guard (default on). Deploys on next gateway restart.

Tests (26, mutation-proven, two tiers):
- Tier-1 unit: decision table + HWM tracker + fail-open reads.
- Tier-2 LIVE (real SessionStore): the B-1 update_id/message_id conflation
  mutation (feed update_id -> miss -> the silent no-op), AC-8 aggregate
  companion rows make the middle constituent answerable, AC-11 side-effects.
Mutations proven RED: fail-closed-on-unanswerable, drop out-of-scope
short-circuit, B-1 conflation, stamp-only-first-constituent.

SPEC: ~/.hermes/plans/2026-07-01_telegram-redelivery-guard-SPEC.md (v1.3,
3 Opus review passes: BLOCK->AWC->AWC, all folded).

Co-authored-by: Apollo <apollo@ang.ventures>
…ully (>100% basis artifact) (#161)

The in-turn A-floor observability marker printed 'kept_tail 75378 / pre 74447 = 101.3%' — a
confusing >100% that reads as 'kept more than existed'. raw_tail_tokens is a documented UPPER BOUND
estimated on the RAW (pre-sanitize) suffix, while pre_tokens is counted on the sanitized/in-context
basis; the two are DIFFERENT bases, so the raw bound can legitimately exceed pre. Compaction itself is
unaffected — only the displayed fraction was misleading.

New _fmt_gross_frac() caps the displayed ratio at 100% and marks it as a bound with a basis note
('≥100% (raw-tail bound N ≥ pre M; raw vs sanitized basis)') when raw exceeds pre; normal within-bound
cases still render as a plain percent. The threshold COMPARISON is unchanged (a bound over the 10%
gross-error line still correctly triggers the two-line degrade) — this is display-only.

3 unit tests incl. the real captured 75378/74447 case; 116 compaction tests green, ruff clean.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…file HERMES_HOME (#162)

The author-to-shared create path (#122) resolved the shared tree as
HERMES_HOME/'skills-shared'. A specialist-profile gateway runs with
HERMES_HOME=<root>/profiles/<name>, so it resolved profiles/<name>/skills-shared
— which doesn't exist (the shared tree is top-level only). Result:
_valid_shared_groups() returned [] and EVERY create from a non-default profile
hard-errored ('category X is not a shared group … (none found)'). Only the
default profile worked.

Fix: _shared_skills_root() derives the top-level root with the same profile
rule as hermes_constants.get_default_hermes_root (parent dir named 'profiles'
-> grandparent), applied to the module's HERMES_HOME so the established test
seam (patching the module global) keeps working. Covers standard
(~/.hermes/profiles/<p>) and Docker (/opt/data/profiles/<p>) layouts; a
non-profile home resolves exactly as before (vanilla upstream unchanged).
Fail-open on any resolution error to the old HERMES_HOME-relative path.

Tests: new TestProfileHomeSharedRoot suite — profile-home resolution, groups
found, create lands in the top-level tree (and NOT under the profile), Docker
layout, a RED-proof documenting the old expression's miss, vanilla/fleet
layouts unchanged (INV-1/2), and fail-open. 35/35 module, 222/222 blast radius.
…ead code) + hoist compaction_stats imports (#163)

The #137 reaper comment justified ignoring the idle heuristic by claiming
_touch_activity is dead code with a fake monotonically-growing idle signal.
Verified against live code: _touch_activity (run_agent.py) IS wired at ~20
call sites (every API call, tool call, stream delta) and two watchdogs
depend on it (cron inactivity kill, gateway poll-loop eviction). The
reaper's task-liveness design is still correct — idleness is the wrong
question for leaked slots (agent object may be GC'd; a quiet live turn can
look idle) — but the stated rationale would mislead the next person
touching restart lifecycle. Comment-only change in gateway/run.py.

Also hoist function-local 'import hashlib' / 'from collections import
Counter' (x4) in agent/compaction_stats.py to module level — _row_signature
runs per-row in the A-floor fallback path.

Tests: 108 compaction-stats tests green under worktree PYTHONPATH.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(pricing): add claude-fable-5 snapshot rate ($10/$50, premium tier)

Real claude-app/claude-fable-5 turns landed unpriced (cost_usd NULL) —
flagged by the tokens-reprice-sweep on 2026-07-01. Add the official-docs
snapshot entry at its premium tier ($10/M in, $50/M out, cache read
$1.00, cache write $12.50; matches the OpenRouter catalog rate) so the
engine prices it and the nightly sweep back-heals the NULL rows.

Test: behavior contract through estimate_usage_cost across all notional
providers + bare anthropic.

* test: also assert empty/None provider (vendor-inference fallback path)
Replace duplicate private partition helpers with one consume-once _signature_partition helper and reuse it for hygiene and in-turn compaction stats.

Verified: python -m pytest tests/agent/test_inturn_afloor_partition.py tests/agent/test_compaction_stats_reconcile.py -q -o addopts="" (97 passed); python -m pytest tests/agent/test_compaction_*.py tests/agent/test_inturn_compaction_stats_replay.py tests/agent/test_inturn_afloor_partition.py tests/run_agent/test_infinite_compaction_loop.py tests/run_agent/test_in_place_compaction.py tests/tui_gateway/test_compaction_status.py tests/context_engine/test_compaction_render_e2e.py -q -o addopts="" (259 passed); python -m ruff check agent/compaction_stats.py tests/agent/test_inturn_afloor_partition.py tests/agent/test_compaction_stats_reconcile.py (passed); git diff --check (passed).

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…del_ids merge (#166)

Plugin provider catalogs namespace curated entries with the provider's own
slug ("my-proxy/claude-x") while the provider's live /models endpoint
returns bare ids ("claude-x"). The curated+live merge compared spellings
with a naive lowercased exact match, so every model appeared TWICE per
provider in every /model picker (desktop Models dialog, TUI, gateway) —
adapters display the bare tail, so both rows render identically.

Fix: canonicalize both sides on the provider's OWN "<slug>/" prefix only,
matched on the "/" boundary. Deliberately NOT a blanket namespace strip —
vendor-namespaced aggregator entries (openai/gpt-5 vs azure/gpt-5) that
merely share a trailing name must not collapse.

Same identifier-representation bug class as #135, different site.

RED->GREEN: 4 behavior tests fail on the old merge loop, pass on the fix;
plus over-collapse, live-only-discovery, live-first-order, and boundary
guards. 200 adjacent picker/provider tests green.
A pasted/clipboard/dropped/picked composer image is written to the
CLIENT's local Electron userData dir (saveImageBuffer / saveClipboardImage)
even when the app runs against a REMOTE gateway. The preview readback,
however, went through readDesktopFileDataUrl(), which in remote mode routes
to the backend's GET /api/fs/read-data-url. The backend is a different
machine and has no such file, so it returns 404 "File not found" — surfaced
to the user as the "Image preview failed" toast on every paste.

The attachment itself was unaffected (submit uploads the bytes via
image.attach_bytes, which already reads the local path through the Electron
bridge) — only the thumbnail preview was broken.

Fix: add readComposerImagePreview(), which reads the preview from the LOCAL
Electron bridge first and only falls back to the remote-aware read when the
local bridge is unavailable or can't read the path (e.g. a path dragged from
the remote file tree). This mirrors what the submit-time upload already does
and keeps the local-mode behavior identical.

Tests: 3 new cases in desktop-fs.test.ts assert remote-mode preview reads
the local bridge and never touches the backend FS REST, plus the
remote-file-tree fallback path.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…tail (#168)

* feat(lcm): honor compression.target_ratio via a token-budgeted fresh tail

LCM previously ignored compression.target_ratio entirely and kept a fixed
32-message fresh tail (~6.5K tokens on live traffic) regardless of context
size. This maps the fleet-standard knob to a token budget:

  budget = target_ratio × threshold_tokens
         (= target_ratio × threshold × context_length)

capped at lcm.fresh_tail_max_tokens (default 60K, guards 1M-window models)
and 0.9 × threshold_tokens (convergence clamp — compaction can never no-op
loop even at ratio=1.0).

At compress() entry the dynamic count K is computed ONCE (frozen-K): walk
backward accumulating count_message_tokens until the budget is spent,
floored at the legacy fresh_tail_count (the tail can only widen, never
shrink below today's behavior). All in-pass cut sites use len(list) - K —
sound because the leaf-loop span is removal-only (the only list-growing
transform, stub insertion via _sanitize_active_context_messages, runs
strictly after the last cut; an AST source-contract test with planted-
violation negative controls guards this invariant against future edits).
Out-of-pass readers (preflight, _raw_backlog_messages) compute fresh.

Also:
- protect_last_n now tracks the dynamic K after each compaction so the
  compaction-stats consumer (find_inturn_kept_cut slack window) stays
  centered; pre-first-compaction value unchanged.
- rotate_active_session preserves max(static, last-dynamic-K) rows so the
  lifecycle frontier can never advance past the wider in-memory tail.
- fail-open fallback compressor now reads the configured target_ratio
  (was a hardcoded 0.20).
- lcm_status exposes the resolved budget + last dynamic count for live
  verification.
- disable switch: lcm.fresh_tail_token_budget_enabled: false (or env
  LCM_FRESH_TAIL_TOKEN_BUDGET_ENABLED=0), fail-safe ON for garbage values.

Degenerate inputs (no context window, flag off, invalid ratio, estimator
exception) reproduce the legacy fixed-count cut exactly.

Ingest/replay/dedup paths untouched (PR #107 territory): zero edits to
_reconcile_ingest_cursor_from_store, the scaffold-overlap guard, or the
min-evidence bar at the replay-skip site.

Verification: 45 new tests (config plumbing, budget resolution, chokepoint
helper, frozen-K boundary consistency, AST guard + planted violations,
rotate floor incl. window-grow, fail-open ratio); full context_engine +
compaction-stats suite 348 passed; offline A/B shakedown on a mixed-density
corpus at 200K: ON-arm tail 98 msgs / 36,946 tokens vs budget 37,500,
legacy arm 32 msgs, 3 fires per arm, cadence_ratio 1.01 (gate ≤ 2.0).

Spec: ~/.hermes/plans/2026-07-01_lcm-target-ratio-token-budget-tail-SPEC.md
(v0.4, three Opus review passes: BLOCK → AWC → AWC, all blockers folded)

* fix(lcm): Greptile review — env-path (0,1] clamp for LCM_TARGET_RATIO + AugAssign list-growth coverage in the frozen-K AST guard

- LCM_TARGET_RATIO env override now applies the same (0,1] range guard as
  the config-file path (out-of-range → default 0.20), with tests.
- The frozen-K AST guard also flags augmented-assignment list growth
  (working_messages += [...]) inside the leaf loop — scoped to message-list
  vars / list-literal RHS so integer counters don't false-positive — with a
  planted += negative control.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Tests-workflow wall time is 95% raw pytest (install is ~10s w/ warm uv
cache). Slices are independent parallel jobs and the per-file subprocess
model is import/IO-bound, not CPU-bound, so both knobs have headroom:

- ci.yml: slice_count 8 -> 10
- tests.yml: HERMES_TEST_WORKERS=12 (was cpu_count*2 = 8 on 4-vCPU)
- run_tests.sh: pass HERMES_TEST_WORKERS through the hermetic env -i
  (without this the env var silently never reached the runner)

Expected: ~5min -> ~3min wall. This PR's own CI run is the measurement.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
#172)

* fix(gateway): reconcile /compress feedback across both transcript axes

The gateway /compress path compresses a chat-only projection of the stored
transcript and, when the rewrite happens, drops every non-chat row (tool
results / system / contentless turns) — usually the bulk of a tool-heavy
session. But every user-facing number was computed on whichever axis was
convenient, so a real 453K→59K compaction reported itself as:

    🗜️ No changes from compression: 179 messages
    Chat size: ~31,406 tokens (unchanged; ...)
    Full request size: 453,542 → ~32,036 tokens ...

— a self-contradicting message (observed live 2026-07-02).

Fixes, per the honesty spec:

- F1: summarize_manual_compression grows an opt-in enhanced mode
  (non_chat_count/tokens, transcript_rewritten, full_before_count) that
  reconciles BOTH axes: CASE A 'Compacted stored transcript: N → M'
  (chat compact, tool rows dropped), CASE B both axes shrank, CASE C
  'No changes: transcript preserved (N messages: X chat + Y tool/system)'.
  Classic positional callers (CLI, TUI) are byte-identical.
- F2: the after-size basis is now the rows the next request actually
  carries — the compressed list only when the store was rewritten, the ORIGINAL
  history otherwise (no more fabricated shrink on a no-op), and the
  no-rewrite full-request line says 'unchanged' instead of printing
  estimator noise as a delta.
- F4: last_prompt_tokens is only zeroed when the transcript was actually
  rewritten; a true no-op keeps the provider-measured figure.
- F3: the fixed overhead (resident agent's real system prompt + tool
  schemas) is resolved ONCE up front and used for BOTH the before and
  after full-request estimates — the before was previously measured with
  the memory-only temp agent's empty overhead while claiming to include
  system + tools.
- New locale key gateway.compress.full_request_unchanged (all 16 catalogs).

Tests: 7 new unit tests for the helper's cases, 2 new gateway E2E-style
regressions (tool-heavy CASE A + CASE C incl. token-reset gating), all
existing compress/i18n/CLI tests green (70 passed).

* fix(review): fold Greptile P2s — all-or-nothing enhanced gate, CASE C token-line invariant, skip unused after-estimate

- enhanced gate now requires non_chat_tokens too (no silent '~0 tokens
  reclaimed')
- CASE C forces the unchanged token wording in the helper itself instead of
  trusting callers to pass equal before/after
- gateway no-rewrite path skips the after-estimate entirely (the reply says
  'unchanged' and never reads it; estimating over a tool-heavy history is
  not free)
- 2 regression tests

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
#173)

* feat(gateway): granular CompactionStats breakdown for manual /compress

Follow-up to #172. The honest two-line form fixed the contradiction, but
the rich per-bucket breakdown the auto-compaction announce renders
(Messages / Context / Removed-from-live-context with the tool-result
sub-split / Model / recovery pointer) was never wired into manual
/compress — despite compaction_stats.py's docstring claiming it as a
consumer. Ace asked for the detailed version.

- After a real rewrite that changed the transcript, /compress now builds
  build_hygiene_stats over (history, chat-projection, compressed) and
  renders _format_granular_announce — the exact same reconciling renderer
  and validate() contract the hygiene announce uses. Example:

    🗜️ Compressed: 7 → 3 stored messages
       Messages:  7 → 3   (kept 3 recent chat)
       Context:   ~564 → ~49 tokens   (freed ~515, 91% smaller)
       Removed from live context (5 messages):
         • 2 tool-result messages  →  ~502 reclaimed   (raw tool output)
         • 1 other messages  →  ~19 reclaimed   (system + tool-call turns, cleared)
         • 2 folded messages   →  ~22 reclaimed   (older chat condensed into 1 summary)
       Model: claude-app/claude-opus-4-8
    ↩ previous transcript preserved: sess-1 (searchable via session_search)
    Full request size: 453,542 → ~49 tokens (...)

- Store-correct recovery pointer: lcm.db (lcm_grep/lcm_expand) for the LCM
  engine, session_search + old session id for a rotated built-in store.
- Degrades to the #172 two-line enhanced form on ANY stats build/validate
  failure (a reconcile bug can never break /compress or ship wrong math),
  and on no-op passes (len unchanged) where there is nothing to break down.
- 2 new gateway regressions: full granular render (buckets, sub-split,
  model line, recovery pointer) + forced-failure degradation.

* fix(review): classify built-in SUMMARY_PREFIX rows as summary; warn on granular build failure

Greptile P2s on #173:
- _is_summary_message only matched the LCM summary regex, so a built-in
  compressor summary row ([CONTEXT COMPACTION — REFERENCE ONLY] ...)
  classified as 'kept chat' and the Messages line miscounted (kept N
  instead of kept N-1 + 1 summary). Now the built-in marker is detected
  too; regression assert added (kept 2 recent chat + 1 summary).
- Unexpected exceptions in the granular stats build now log at WARNING
  (parity with the validate() failure path); DEBUG hid real regressions.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…nscript basis (#174)

* fix(gateway): label manual /compress granular block by its stored-transcript basis

The granular CompactionStats renderer (#173) was written for the auto-
compaction announce, whose input IS the live message list — so 'Context:' /
'Removed from live context' / 'kept in context' accurately describe wire
savings there. Manual /compress reuses the same renderer but feeds it the
STORED transcript (session_store.load_transcript), which under LCM has already
had its bulk compacted OFF the wire in earlier passes. The reduction is storage
reclaimed, not request-size — so the block overstated wire savings (Ace's live
case: '~689K → ~37K, freed 651K, 95% smaller' against a real 303K request).

Make _format_granular_announce basis-aware via a keyword-only  arg
(default 'live' → byte-identical auto-announce output). Manual /compress passes
basis='stored', rendering 'Stored transcript:' / 'Removed from stored
transcript' / 'reclaimed' / 'kept in transcript'. The separate provider-measured
'Full request size:' line stays the wire truth beside it, so the user sees both
storage-reclaimed AND real request-size cut.

Numbers unchanged (build_hygiene_stats / CompactionStats / validate untouched) —
only labels. Adds TestGranularBasisLabel (default==live, stored relabels without
changing any number, renderer never emits the wire line) + an LCM-engine E2E
through the real _handle_compress_command reproducing Ace's exact scenario.

* test: assert kept-in-transcript unconditionally when Replacement cost present (Greptile P2 #174)

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
… aren't all "tui" (#175)

* feat(tui_gateway): client-identity source so desktop/dashboard/mobile aren't all "tui"

The desktop app, the in-browser dashboard chat, and the Ink stdio TUI all drive
the same tui_gateway JSON-RPC server, which hardcoded platform="tui" on every
agent it built. So every desktop-app turn recorded platform="tui" in blackbox
(turns.platform) — the source dimension on the token/cost dashboard couldn't
tell the desktop client from the terminal, and a future mobile app would fold
in too.

Let a client self-declare a `source` on session.create/resume; thread it
through _make_agent -> AIAgent(platform=...) (which already feeds the blackbox
on_session_end hook), and store it on the session dict so the deferred/lazy
build, /new reset, resume, and branch paths all preserve it. `source` is
untrusted client input feeding a persisted DB dimension, so it is sanitized to
a conservative slug (`^[a-z][a-z0-9_-]{0,23}$`) with a `tui` fallback — the
label stays open-ended (new clients name themselves; no central enum to keep in
sync, per the existing "fixed set of platform names goes stale" note).

Desktop app: stamp source:'desktop' once, centrally, in the requestGateway
wrapper for the session-originating methods (create/resume) — never overriding
an explicit caller source (e.g. the dashboard sidecar's 'tool').

No schema/blackbox change (the platform pipe already existed); stdio Ink and
internal rebuilds keep the historical "tui" default so nothing regresses.
Historical desktop turns are indistinguishable from real TUI turns, so the
clean split starts at deploy (no backfill possible).

Tests: tui_gateway/test_client_source_platform.py (sanitizer table + _make_agent
platform derivation); desktop use-gateway-request.test.tsx (injection contract:
stamps create/resume, skips other methods, never overrides explicit source).

* test: accept _make_agent(source=…) in remaining narrow test fakes

The new source kwarg on _make_agent tripped three narrow monkeypatch fakes in
tests/test_tui_gateway_server.py (session_id/session_db-only signatures). Widen
them to **_kw so they tolerate the new kwarg, matching the sibling fakes that
already splat. No behavior change.

* fix(tui_gateway): reject non-str source (Greptile P2)

_sanitize_client_source used str(raw or ''), so a JSON-RPC boolean true would
coerce to the slug 'true' instead of falling back to 'tui'. Guard with
isinstance(raw, str) up front; add True/0/[] cases to the sanitizer table.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…ead, one number story (#176)

The manual /compress granular block led with the stored-transcript estimate
(Stored transcript: ~328K -> ~43K, chars/4 over the archive) while the footer
showed the REAL provider-measured wire count (e.g. 236K). Two numbers, two
populations, two tokenizers -> they disagreed by 100K+ every time and confused
the operator (and the agent explaining it).

The real provider-measured before-count already exists at /compress time
(session_entry.last_prompt_tokens). This makes the granular block lead with it.

Change (stored basis only, when a real measured count is available):
- Context line becomes the WIRE story: measured before (exact, no ~) ->
  next-request estimate (~), freed/pct computed on wire numbers, tagged
  "before measured, after next-request estimate".
- Stored-transcript totals are demoted into the "Removed from stored transcript"
  header as an explicit "token-est reclaimed from archive" parenthetical, so
  they can never be read as request-size savings.
- The trailing "Full request size:" line is skipped when the block already
  carried the wire story (no double-reporting).

Fallbacks unchanged and covered:
- No real measured count (post-restart) -> prior stored-basis rendering.
- Auto-compaction announce (basis="live") -> byte-identical; wire kwargs are
  ignored on the live basis.
- No-rewrite / two-line forms -> unchanged.

New renderer kwargs wire_before/wire_after are honored only for basis="stored"
with both > 0. 11 new renderer tests + 2 updated gateway E2E tests (the LCM
case reproduces Ace's exact 303,201 scenario). 217 compaction/compress tests
green.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…self-identifying markers (#177)

* fix(compaction): gate in-turn stats on announce render-eligibility + self-identifying markers

The in-turn compaction-stats block (build_inturn_stats + gross-error guard +
COMPACTION_STATS_* WARNINGs) ran for EVERY LCM compress() call — including
no-ops — while the announce formatter default-denies non-rendering statuses
(noop/idle/running/bypassed; conditional statuses failing the post<pre render
check). Result: ~100% kept_tail APPROX_ATTRIBUTION and TAG_MISSING WARNINGs
fired for announces that never render — pure log noise that polluted the
daily compaction-stats watcher report (measured 2026-07-02: every real-session
degrade marker in apollo's agent.log rode an 'LCM compression no-op' line).

Fix: _inturn_stats_render_eligible() mirrors the formatter's own gating by
consuming the SAME allow-list objects (single source of truth, no copied
literals) + the conditional token check, applied on the exact variables the
announce call passes as pre_tokens/post_tokens. LCM-scoped: the built-in path
keeps its sid-rotation gating and may not expose _last_compression_status.

Also (D-4): every compaction-stats marker now carries session=<id>, plus
src=test under pytest (PYTEST_CURRENT_TEST), so the log watcher can attribute
markers by the line's own fields instead of fragile proximity joins — pytest
writing through the live logging config previously inflated production counts.

Tests: 13 new (gate predicate matrix vs the real formatter as oracle; drift
guards asserting the gate uses the formatter's allow-list objects and the
announce-call token variables; stats=None render tolerance; marker session/
src=test fields incl. from a threading.Thread worker — the measured polluter's
shape; throttle unaffected). Spec (4 Opus review passes):
plans/2026-07-02_inturn-noop-gate-multipass-provenance-SPEC.md v0.5, PR-A.

* fix(compaction): don't suppress non-LCM stats in the render gate (Greptile #177)

The render gate was written as (_engine_name == 'lcm') and render_eligible(...),
which short-circuits to False for the built-in (non-LCM) compressor and
permanently dropped its stats — silently degrading every built-in announce to
the two-line form. Intent was to gate LCM no-ops only. Restructure to an
explicit branch: LCM → render-eligibility; non-LCM → always eligible (unchanged
always-attempt behavior).

Also closes the Greptile test-coverage gap: _run_announce_block now mirrors the
real LCM-scoped gate and takes explicit pre/post tokens, and new tests exercise
(a) a conditional status ('sanitized') with a REAL compression post<pre through
build+validate, (b) a conditional no-op not built, and (c) non-LCM not
suppressed by the gate.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
… extension) (#178)

* feat(compaction): multi-pass provenance stamp, shadow-first (Option B extension)

The Option B _src_idx provenance stamp (#110) — the EXACT kept/folded
partition — was gated to single-pass compactions (leaf_passes == 1). Heavy
multi-pass sessions (dynamic leaf chunking, deferred maintenance) never got
the stamp, fell to whole-tail replay, and on sanitize-heavy tails degraded to
the A-floor's approximate split → the honest two-line announce instead of the
granular breakdown.

Widen the stamp gate to leaf_passes >= 1. Soundness (spec §0.6, five-step
line-cited proof): ingest is a 1:1 order-preserving rewrite; every in-loop
mutation removes rows ONLY from the front region (scaffold drop + compacted-
chunk removal both slice strictly before fresh_tail_start; the leaf chunk is
a contiguous FRONT prefix); summary rows go to the DAG, never into
working_messages (summaries are prepended inside _assemble_context, after the
stamp); frozen-K means 'the last K rows' names the same physical rows after
every pass; and the stub-inserting sanitize (_sanitize_tool_pairs — single
caller: _sanitize_active_context_messages) runs only inside/after
_assemble_context, strictly post-stamp. Hence the fresh tail is a contiguous
suffix of the ORIGINAL messages for any pass count, and the end-anchored
mapping is valid. Stamp guard hardened: role + tool_call_id + tool_calls
arity (structural fields that survive content rewrites).

SHADOW-FIRST (observe before trust): the consumer trusts single-pass stamps
exactly as before; for multi-pass it computes B, displays the current
replay/A-floor behavior byte-identical, and compares via an index-SET
predicate — diverge <=> sorted(B kept-origin indices) != sorted(current kept-
origin indices), never a token sum (an equal-token wrong-but-in-range swap
changes the set but not the sum). Logged as
COMPACTION_STATS_B_MULTIPASS_SHADOW agree (INFO) | diverge (WARNING) for the
soak; a follow-up trust-flip PR promotes multi-pass B to primary after a
clean soak. Engine exposes last_leaf_passes for the trust decision.

Tests: multi-pass stamp exactness across fold counts + boundary shifts +
scaffold drops; structural-guard rejection; shadow display byte-equality;
the diverge branch fires (incl. an EQUAL-TOKEN wrong-but-in-range fixture
that fails any summing comparator); corrupt stamps fall to A-floor; compare
exceptions never break the build; single-pass byte-identical.

Spec (4 Opus review passes, zero blockers at close):
plans/2026-07-02_inturn-noop-gate-multipass-provenance-SPEC.md v0.5, PR-B.

* fix(compaction): un-throttle the shadow diverge warning (Greptile #178)

The diverge arm of the multi-pass shadow compare used
_warn_compaction_stats_once, which fires at most once per (cause, session).
A heavy session that diverges multiple times would drop all but the first,
so the soak analysis couldn't measure within-session diverge frequency or
tell a one-off fluke from a systematic disagreement. Replace with a direct
logger.warning so every diverge event is independently observable; it carries
the same session=/src=test fields for watcher attribution.

Tests: assert two diverges in one session both emit (not throttled) + a
source-structure guard that the diverge arm uses logger.warning, not the
throttled helper.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Replace the /api/sessions/stats by_source hydrate with a SessionDB GROUP BY helper and run the blocking stats read in the default executor.\n\nVerified with scripts/run_tests.sh tests/test_web_server.py tests/test_hermes_state.py (308 passed). Post-spike read-only live-state measurement: stats max 1.314s, trivial endpoint max 0.597s, fd count returned to baseline.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Route /compress and session-reading slash commands through the live gateway session instead of the slash worker, add the manual compression no-op/concurrency guard, and block local file reveal/rename/delete for remote gateway files.\n\nVerified:\n- PYTHONPATH=$PWD ~/.hermes/runtime/hermes-agent/venv/bin/python -m pytest tests/test_tui_gateway_server.py -k 'compress_history_identity_noop or compress_history_same_length or slash_exec_compress_bypasses or slash_exec_compress_reports_existing or command_dispatch_compress_returns or live_read_commands' -q\n- PYTHONPATH=$PWD ~/.hermes/runtime/hermes-agent/venv/bin/python -m pytest tests/test_tui_gateway_server.py -k 'not browser_manage_connect_default_local_reports_launch_hint' -q\n- npm run test:ui -- src/app/session/hooks/use-prompt-actions/index.test.tsx src/store/file-actions.test.ts\n- npm run typecheck\n- npm run lint -- src/app/session/hooks/use-prompt-actions/index.test.tsx src/app/session/hooks/use-prompt-actions/slash.ts src/store/file-actions.ts src/store/file-actions.test.ts\n- git diff --check

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Route remote desktop file opens through an authenticated bytes-to-temp path instead of token/download URLs or client-local file:// fallbacks. Add private temp lifecycle helpers with start/periodic TTL sweeps and harden remote HTML previews with blob URLs, CSP, and JavaScript disabled.\n\nVerified:\n- node --test apps/desktop/electron/gateway-temp-files.test.cjs\n- npm --workspace apps/desktop exec vitest run --environment jsdom src/lib/media.remote.test.ts src/lib/html-preview.test.ts src/lib/local-preview.remote.test.ts\n- npm --workspace apps/desktop exec eslint src/lib/media.ts src/lib/media.remote.test.ts src/lib/html-preview.ts src/lib/html-preview.test.ts src/lib/local-preview.ts src/lib/local-preview.remote.test.ts src/app/chat/right-rail/preview-pane.tsx electron/gateway-temp-files.cjs electron/gateway-temp-files.test.cjs electron/main.cjs\n- npm --workspace apps/desktop run typecheck\n- npm --workspace apps/desktop run test:desktop:platforms

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Offloads the remaining dashboard SessionDB read sites from async handlers through sync helpers and adds Semaphore(2) bounding for heavyweight reads. Adds thread-assert coverage plus an AST-aware regression lint with bare/lambda/partial/inner-def/to_thread self-tests.

Verified: scripts/run_tests.sh tests/test_web_server_sessiondb_eventloop.py tests/test_web_server.py tests/test_hermes_state.py -q; scripts/run_tests.sh tests/hermes_cli/test_dashboard_auth_status_endpoint.py -q.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
The /reset and /new banner (_format_session_info) surfaces Model, Provider,
Context, and Endpoint so a user can immediately see the fresh session's
config. It omitted the reasoning effort, which is independently configurable
(agent.reasoning_effort) and just as material to cost/behavior.

Add a '◆ Reasoning: <level>' row between Provider and Context, resolved from
_load_reasoning_config():
  - None            -> 'medium (default)'
  - {enabled:False} -> 'none'
  - {enabled:True}  -> the effort level (minimal/low/medium/high/xhigh)

Global config (not the per-session override) is correct here: the banner
renders on /reset and /new, which clear any /reasoning session override, so
config.yaml is the effective level at reset time. Resolution is wrapped so a
failure omits the row rather than breaking the banner (mirrors the existing
Endpoint conditional-append).

Tests: 5 new cases in test_session_info.py covering each reasoning state,
row ordering, and the fail-open path.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…op PRD Phase 3) (#185)

* fix(desktop): gate remote media and file reveals

Route remote a/v markdown attachments through the gateway temp-file opener before local stream fallbacks, add the desktop remote file-boundary eslint rule/self-test, and show an honest remote reveal message.\n\nVerified:\n- npm run test:ui -- eslint.config.test.mjs src/components/assistant-ui/markdown-text.test.ts src/lib/desktop-fs.test.ts src/store/projects.test.ts src/lib/media.remote.test.ts\n- npx eslint src/components/assistant-ui/markdown-text.tsx src/components/assistant-ui/markdown-text.test.ts src/lib/desktop-fs.ts src/lib/desktop-fs.test.ts src/store/projects.ts src/store/projects.test.ts src/store/file-actions.ts src/components/assistant-ui/directive-text.tsx src/components/chat/generated-image-result.tsx src/app/chat/right-rail/preview-file.tsx src/i18n/en.ts src/i18n/zh.ts src/i18n/zh-hant.ts src/i18n/ja.ts src/i18n/types.ts\n- npm run typecheck\n- git diff --cached --check

* fix(desktop): reconcile B1/D1 remote file-action guard seam

D1's facade-throw guard (desktop-fs *File exports) is the enforced safety net;
restore B1's entry-point blockRemoteFileAction guards (union, defense-in-depth)
on all five entry points incl. executeFileRename/Delete for the data-loss bug.
Update file-actions.test.ts to the *File facade mock names.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Treat widthOverride as persistent per the pane-store feature commits and add a module-reload assertion that persisted width loads back into the pane store.\n\nVerified:\n- npm --prefix apps/desktop run test:ui -- src/store/panes.test.ts\n- npm --prefix apps/desktop run typecheck\n- npm exec prettier -- --check src/store/panes.test.ts (from apps/desktop)
Add a server-owned pinned flag to SessionDB, expose it through REST/session lists and the session.pin gateway RPC, and derive desktop pinned rows from server session data instead of localStorage.

Verified: scripts/run_tests.sh tests/test_hermes_state.py tests/test_tui_gateway_server.py tests/test_web_server_sessiondb_eventloop.py; npm --workspace apps/desktop run typecheck; npm --workspace apps/desktop run test:ui -- src/store/session.test.ts src/store/layout.test.ts src/store/updates.test.ts.

Note: npm --workspace apps/desktop run test:ui was also attempted after a temporary worktree node_modules symlink; it failed in desktop suites outside the touched pinned-session files (46 files / 21 tests).
The server-side pin PR (#186) made $pinnedSessionIds a read-only computed
atom, which dropped setPinnedSessionOrder + the sidebar's reorderPinned
handler and hardcoded sortable={false} on the pinned list — losing
drag-to-reorder. Restore it with a local $sidebarPinnedOrderIds order atom
(mirrors $sidebarSessionOrderIds) layered over the server-synced pin SET via
orderByIds: pin membership syncs across devices, visual order stays per-device.
Keyed by durable lineage-root pin ids so order survives compression.
* feat(dashboard): web chat client self-declares source='dashboard'

The in-browser dashboard chat (web/), the Electron desktop app, and the Ink TUI
all drive the same tui_gateway JSON-RPC server, which stamps platform='tui' by
default. PR #175 taught the Electron app + backend to thread a client 'source'
so turns are attributed per-client (desktop/dashboard/mobile) on the token
dashboard — but the WEB dashboard client (web/src/lib/gatewayClient.ts) was
never updated, so browser-dashboard turns still recorded as 'tui'.

Override GatewayClient.request to stamp source:'dashboard' on the
session-originating methods (session.create/session.resume) when the caller
hasn't set an explicit source (the sidebar sidecar session deliberately sends
source:'tool'). Mirrors the Electron app's requestGateway injection exactly.
Backend already sanitizes + threads source -> AIAgent(platform=...) (#175); no
backend change needed.

Test: web/src/lib/gatewayClient.test.ts (4) — stamps create/resume, skips other
methods, never overrides an explicit source.

* ci: re-trigger checks (apps/desktop/node_modules ENOTDIR runner flake, not a code issue)
…emote-mode correctness, pinned-session sync) (#191)

As-built specs with full review logs (eventloop: 4-pass super-pass; desktop remote:
9-pass super-pass; pin-sync: 1-pass light track), Phase-0 findings, live AC-2
results, and residual follow-ups (session.list projection cost, BUG-E facade,
transient-dirty build stamp).
@Kyzcreig
Kyzcreig requested a review from a team July 14, 2026 19:38
Kyzcreig and others added 13 commits July 14, 2026 19:46
By default the desktop app persists the last composer model pick to
localStorage and carries it across new chats (Cmd+N) and restarts, so
picking an expensive model makes it the sticky default for every new
session. This adds an opt-in config knob to instead reseed a fresh
draft's model from the profile default.

- HermesConfig type gains desktop.reset_model_on_new_session (optional).
- useHermesConfig mirrors it into a new $resetModelOnNewSession atom.
- refreshCurrentModel treats a fresh draft as a forced reseed when the
  flag is on, overwriting the sticky pick with the profile default.
  A live session still owns its model (never disturbed mid-conversation).
- The fresh-draft effect loads config before the model reseed so the flag
  is in the store first (avoids a cold-boot race).

Default false preserves the existing sticky-last-pick behavior. Backend
needs no change: /api/config already returns the whole config, so the
desktop section flows through untouched.

Tests: use-model-controls (flag on reseeds, flag off keeps pick, active
session untouched, explicit force still works) + use-hermes-config
(mirrors true, defaults false when absent). Full desktop UI suite green
(1283 passed), typecheck clean.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
… zombie writes (#341)

* fix(persist): append-time generation gate — suppress a /stop'd turn's zombie writes

A turn interrupted by /stop is a cooperative interrupt: the inner
retry/fallback loop can keep spinning across fallback providers, and its
incremental flush keeps appending rows to the session DB *after* the user
stopped the turn (the #666915-666942 incident: 14 assistant(tool_calls)
+ 14 tool results written post-/stop). Those zombie rows let a later
/undo land somewhere the drain then clobbers.

Add an append-time generation gate at the single persist chokepoint
(_flush_messages_to_session_db). When the gateway invalidates a turn's
run generation (/stop, /new), it sets agent._persist_superseded; the
flush then suppresses that turn's continued NEW content/tool rows.

PAIRING-SAFE (never orphans a tool call → NousResearch#48879 role-alternation
corruption):
- gate runs AFTER all already-durable skips, so a pre-/stop
  assistant(tool_calls) is never re-evaluated and its result passes (B1)
- the suppressed-tool-call-id set is AGENT-SCOPED, surviving the whole
  drain, because the assistant(tool_calls) row and its tool result flush
  in SEPARATE flushes within one iteration — a per-flush set would orphan
  the result (B1')
- the interrupt-close tail (finish_reason=interrupt_close) is ALWAYS
  persisted — it is the auto-continue / restart-loop signal, and the
  reader _is_interrupt_close_tail consumes only agent_history[-1] (I1)

Fail-open (I5): _persist_superseded resolved once via getattr-default-
False in try/except — a guard error persists normally (a stray late row
is cosmetic; #339 already stops /undo racing it; a dropped real row is
data loss). Best-effort: the flag-set (event loop) races the flush
(executor thread); reaper path scoped out (R6). logger.info on suppress.

Tests (all RED-proven): 8 in test_persist_superseded_gate.py
(same-flush pair, B1 owner-durable, B1' cross-flush, I1 tail, I1
synthetic-repair, I3 byte-identical, I5 fail-open, AC1 suppress) +
2 in test_undo_drain_guard.py (Phase-2 flag-set, R6 reaper-unset).

3 Opus review passes: BLOCK(B1) -> BLOCK(B1',B2') -> APPROVE-WITH-CHANGES.
Spec: ~/Projects/hermes-append-gen-gate/SPEC.md v1.3.

* fix(persist): gate fails open on unexpected non-assistant/tool rows (Greptile-P2)

The else-branch suppressed ALL non-tool roles. A zombie turn only writes
assistant+tool rows, but if scaffolding ever injects a NEW user/system
row mid-drain, dropping it would be data loss. Suppress only assistant
rows; fail OPEN (persist + debug-log) on any other role. Adds
test_unexpected_new_user_row_fails_open_during_supersede.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(blackbox): price MoA turns by physical model routes

* fix(blackbox): preserve turn accumulator append invariant

* fix(blackbox): keep MoA telemetry failures isolated

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(terminal): make snapshot temp files portable

* fix(tests): make the full suite hermetic on macOS

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(kanban): clear respawn guard for closed PRs

Resolve recent GitHub PR state with a per-tick capped memo, retain fail-safe guarding on errors, and surface guarded tasks in CLI, gateway logs, and stranded-ready diagnostics. Verified: scripts/run_tests.sh on 4 targeted files, 309 tests passed.

* fix(kanban): deduplicate respawn guard PR URLs

Avoid repeated resolver work when steering comments quote the same PR. Verified: targeted kanban suite, 310 tests passed.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…time (#348)

`cron/jobs.py` froze `CRON_DIR`/`JOBS_FILE`/`OUTPUT_DIR` (and the ticker
heartbeat paths) as module-level constants at *import* time, and
`_current_cron_store()` returned those frozen paths whenever no explicit
`use_cron_store()` override was active. So if the module was imported while
`HERMES_HOME` pointed at the real `~/.hermes` — e.g. a real agent booting from
a worktree/e2e harness that shares the live home — a *later* `HERMES_HOME`
redirect (the pytest hermetic fixture, or an e2e harness pointing at a tempdir)
was silently ignored, and `create_job` / the `cronjob` tool wrote straight into
the LIVE `cron/jobs.json`.

This actually happened: a real-agent blackbox e2e session booting from a
shared-home worktree wrote 11 unit-test fixture jobs (7x name=brief every-60m,
3x "claim job", 1x "paused job") into the production cron store, re-arming
cron-config-lint errors and paging cron-health.

Fix: `_current_cron_store()` resolves `get_hermes_home()` LIVE when there is no
explicit `use_cron_store()` override, instead of returning the import-frozen
constants. The same import-time freeze afflicted the ticker heartbeat/success
paths (sibling of the same bug class), so those now resolve live too. The
module-level constants remain as a back-compat surface for existing importers
(e.g. hermes_cli/backup.py's JOBS_FILE reference).

Migrated the 5 tests that coupled to the removed mechanism (patching
`cron.jobs.CRON_DIR` / `TICKER_HEARTBEAT_FILE` directly) to redirect
`HERMES_HOME` — the supported isolation path, and a behavior contract rather
than a change-detector on internals. This includes the concurrency-isolation
test `test_profile_call_cannot_retarget_ticker_store_mid_write`, whose invariant
(a per-thread profile override must not retarget a concurrent ticker's store)
is preserved: the profile call uses `use_cron_store()` (a per-thread ContextVar),
the ticker uses no override and now resolves the live default home.

Regression test `tests/cron/test_store_resolves_live_home.py` reproduces the
leak (a HERMES_HOME set after import must be honored; a job must not land in the
stale import-time home) and is RED against the frozen-constant implementation,
GREEN with the fix. Full cron suite: 752 passed.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
fix(desktop): distinguish xhigh and max reasoning
The desktop renders a sent turn optimistically (client-minted ids user-<ts>/assistant-<ts>) and stamps them to the committed DB ids carried on the message.complete frame. Two defects made stamping a no-op, so the live session-sync poll re-fetched the committed rows, failed to recognize them, and appended them as DUPLICATES (every message shown twice; DB stays clean):

1. Server (tui_gateway/server.py): message.complete never included the committed ids. Add _completed_turn_committed_ids() to extract [user_id, assistant_id] for the finished turn from session history (_db_persisted_row_id) and ship them as payload.message_ids.

2. Client (use-session-changes.ts): the stamp predicate only matched user-prefixed OR pending rows, but completeAssistantMessage() clears the streamed assistant row pending flag BEFORE markTurnComplete() stamps it, so the completed assistant row (id assistant-<ts>, pending:false) was skipped and duplicated. Recognize any optimistic id (user-/assistant- prefix) OR pending.

Gated by tests that FAIL on the naive code: tests/tui_gateway/test_completed_turn_committed_ids.py (6) + two new cases in use-session-changes.test.ts proving the already-completed assistant row stamps and does not duplicate end-to-end.
@Kyzcreig
Kyzcreig force-pushed the fix/desktop-optimistic-stamp-dup branch from 33429b5 to 0fc8088 Compare July 15, 2026 22:56
@alt-glitch alt-glitch added type/bug Something isn't working comp/desktop Electron desktop app (apps/desktop/*) comp/tui Terminal UI (ui-tui/ + tui_gateway/) P3 Low — cosmetic, nice to have needs-decision Awaiting maintainer decision before any implementation sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Jul 15, 2026
@teknium1

Copy link
Copy Markdown
Collaborator

Thanks for tracing the optimistic-row failure mode. This automated hermes-sweeper review found that current main has already moved to a different reconciliation path, so this PR's message_ids payload is no longer consumed.

  • apps/desktop/src/app/session/hooks/use-message-stream/gateway-event.ts:383-384 completes the assistant message from final text only.
  • apps/desktop/src/app/session/hooks/use-message-stream/index.ts:405-458 finalizes the existing stream bubble and uses stored-history hydration only as a fallback.
  • apps/desktop/src/app/contrib/wiring.tsx:253-259 replaces the current transcript with authoritative stored messages rather than appending a session.changes result.
  • The current-main desktop session-hooks refactor is 2afbe777636078622ae39c60ddb85300d0690296; the old stamping symbols are absent from current main.

Closing as implemented on main: the duplicate-producing append/poll seam described here no longer exists.

@teknium1 teknium1 closed this Jul 16, 2026
@teknium1 teknium1 added sweeper:implemented-on-main Sweeper: behavior already present on current main sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation sweeper:risk-caching Sweeper risk: may break/degrade prompt caching or cache-key stability (invariant) sweeper:blast-broad Sweeper blast radius: broad — a core path most sessions hit labels Jul 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/desktop Electron desktop app (apps/desktop/*) comp/tui Terminal UI (ui-tui/ + tui_gateway/) needs-decision Awaiting maintainer decision before any implementation P3 Low — cosmetic, nice to have sweeper:blast-broad Sweeper blast radius: broad — a core path most sessions hit sweeper:implemented-on-main Sweeper: behavior already present on current main sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation sweeper:risk-caching Sweeper risk: may break/degrade prompt caching or cache-key stability (invariant) sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants