Repository navigation
Conversation
…nfig override) (#158) * feat(desktop): parseDesktopBackendRoot pinned-grammar config reader (P1.1, AC-4) No-YAML-dep targeted scalar read of desktop.backend_root with indentation-anchored grammar: nested-under-desktop only, comment strip, quote unwrap, ~/\$HOME expand, first-wins; rejects tab-indent/flow-style/multi-doc -> null (fail-safe). 18/18. * feat(desktop): backend runs from runtime tree by default (P1.1-1.5, AC-1/3/4/5) resolveBackendRoot 3-tier precedence (config override > runtime > dev) wired into createActiveBackend; parseDesktopBackendRoot no-YAML-dep config read; ready-file emits project_root (AC-3 effect gate, fail-closed); shared fallback surface (AC-5). No live behavior change on installs without a runtime tree. 313/314 desktop mechanics green. * fix(desktop): reason-accurate backend-fallback log (override-unusable may still land on runtime tree) * fix(desktop): honor leading --- YAML document-start marker in backend_root config reader --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…n guard) (#159) * test(telegram): lock restart/reconnect message-loss parity (regression guard) Discord had a drain-window message-loss bug (PR #157) because it is a *push* transport. Telegram is a *pull* transport (getUpdates long-poll) with a server-side queue, so downtime is backfilled for free — but only because of three independent properties a refactor could each silently break. Lock them as behavior contracts so the loss bug can never appear on Telegram. INV-1 reconnect(is_reconnect=True) preserves the queue (drop_pending_updates= False); cold boot drops it. (positive + negative control) INV-2 every start_polling recovery ladder preserves the queue unconditionally (AST source-contract; bootstrap-vs-recovery discriminated by enclosing function, not line number; MISSING kwarg = explicit FAIL). INV-3a disconnect() drains in order updater.stop -> app.stop -> app.shutdown. INV-3b PTB-provenance guard: Application.stop() in the installed PTB actually calls update_queue.join(); fails loudly (naming version) on a PTB bump. Test-only + doc note; no production code change. All 6 tests mutation-proven to gate (see PR body). Ground-truthed against pinned PTB 22.7. SPEC: ~/.hermes/plans/2026-07-01_telegram-parity-regression-SPEC.md Spike: ~/.hermes/plans/2026-07-01_telegram-parity-SPIKE.md * test(telegram): harden mock guard + PTB provenance walk (Greptile follow-ups) Two Greptile correctness nits (Confidence 4/5, 'safe to merge'): - _ensure_telegram_mock: hasattr(__file__) is True for a MagicMock too, so a prior test's mock in sys.modules was misread as 'real telegram present'. Now distinguishes a genuine module (real str __file__, not a MagicMock). - PTB provenance test: outer walk didn't stop at the first Application class and could match a nested/inner class. Iterate module-level statements only and break at the first match. Both mutation-re-verified to still gate; 6 passed; ruff clean. --------- Co-authored-by: Apollo <apollo@ang.ventures>
…pe B) (#160) Telegram is loss-safe (#159) but had a KNOWN-OPEN residual: a hard kill (SIGKILL/OOM/power) after the agent answered an update but before PTB confirmed its offset makes Telegram re-deliver the answered update on next boot -> the agent answers it twice. This guard suppresses a proven-duplicate re-delivery, exactly once, and fails OPEN in every uncertain case (never drops a genuine message -- suppressing on uncertainty would convert a rare dup into a loss). Two Telegram id spaces, NEVER interchanged (SPEC section 0): - update_id (platform_update_id): envelope seq, used ONLY to SCOPE the guard to genuine candidate re-deliveries (update_id <= persisted HWM). - message_id (stamped platform_message_id): used ONLY for the answerability lookup (was this message already answered?). Design: - gateway/telegram_redelivery.py: TelegramHwmTracker (in-memory HWM, coalesced <=30s + shutdown flush, atomic 0600, fail-open read), in_redelivery_scope (HWM-primary, 120s time-cap fallback only when no HWM), decide_redelivery (the fail-open decision table), RedeliverySuppressionCounter (content-free PHASE=tg_redelivery_suppressed observability). - gateway/run.py: _is_telegram_boot_redelivered_duplicate wired into the general message path after session resolution, before the agent run; HWM flush on graceful stop; INV-6 companion-row persist for aggregated turns. - telegram/adapter.py: _enqueue_text_event accumulates ALL constituent message_ids of a buffered multi-update turn. - INV-6: aggregated turns stamp NON-first constituents as observed=1 companion rows (reuses shipped observed column + exact-match primitive, no schema change, cache-safe) so a re-delivery of any constituent is answerable. Config: telegram.redelivery_guard (default on). Deploys on next gateway restart. Tests (26, mutation-proven, two tiers): - Tier-1 unit: decision table + HWM tracker + fail-open reads. - Tier-2 LIVE (real SessionStore): the B-1 update_id/message_id conflation mutation (feed update_id -> miss -> the silent no-op), AC-8 aggregate companion rows make the middle constituent answerable, AC-11 side-effects. Mutations proven RED: fail-closed-on-unanswerable, drop out-of-scope short-circuit, B-1 conflation, stamp-only-first-constituent. SPEC: ~/.hermes/plans/2026-07-01_telegram-redelivery-guard-SPEC.md (v1.3, 3 Opus review passes: BLOCK->AWC->AWC, all folded). Co-authored-by: Apollo <apollo@ang.ventures>
…ully (>100% basis artifact) (#161) The in-turn A-floor observability marker printed 'kept_tail 75378 / pre 74447 = 101.3%' — a confusing >100% that reads as 'kept more than existed'. raw_tail_tokens is a documented UPPER BOUND estimated on the RAW (pre-sanitize) suffix, while pre_tokens is counted on the sanitized/in-context basis; the two are DIFFERENT bases, so the raw bound can legitimately exceed pre. Compaction itself is unaffected — only the displayed fraction was misleading. New _fmt_gross_frac() caps the displayed ratio at 100% and marks it as a bound with a basis note ('≥100% (raw-tail bound N ≥ pre M; raw vs sanitized basis)') when raw exceeds pre; normal within-bound cases still render as a plain percent. The threshold COMPARISON is unchanged (a bound over the 10% gross-error line still correctly triggers the two-line degrade) — this is display-only. 3 unit tests incl. the real captured 75378/74447 case; 116 compaction tests green, ruff clean. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…file HERMES_HOME (#162) The author-to-shared create path (#122) resolved the shared tree as HERMES_HOME/'skills-shared'. A specialist-profile gateway runs with HERMES_HOME=<root>/profiles/<name>, so it resolved profiles/<name>/skills-shared — which doesn't exist (the shared tree is top-level only). Result: _valid_shared_groups() returned [] and EVERY create from a non-default profile hard-errored ('category X is not a shared group … (none found)'). Only the default profile worked. Fix: _shared_skills_root() derives the top-level root with the same profile rule as hermes_constants.get_default_hermes_root (parent dir named 'profiles' -> grandparent), applied to the module's HERMES_HOME so the established test seam (patching the module global) keeps working. Covers standard (~/.hermes/profiles/<p>) and Docker (/opt/data/profiles/<p>) layouts; a non-profile home resolves exactly as before (vanilla upstream unchanged). Fail-open on any resolution error to the old HERMES_HOME-relative path. Tests: new TestProfileHomeSharedRoot suite — profile-home resolution, groups found, create lands in the top-level tree (and NOT under the profile), Docker layout, a RED-proof documenting the old expression's miss, vanilla/fleet layouts unchanged (INV-1/2), and fail-open. 35/35 module, 222/222 blast radius.
…ead code) + hoist compaction_stats imports (#163) The #137 reaper comment justified ignoring the idle heuristic by claiming _touch_activity is dead code with a fake monotonically-growing idle signal. Verified against live code: _touch_activity (run_agent.py) IS wired at ~20 call sites (every API call, tool call, stream delta) and two watchdogs depend on it (cron inactivity kill, gateway poll-loop eviction). The reaper's task-liveness design is still correct — idleness is the wrong question for leaked slots (agent object may be GC'd; a quiet live turn can look idle) — but the stated rationale would mislead the next person touching restart lifecycle. Comment-only change in gateway/run.py. Also hoist function-local 'import hashlib' / 'from collections import Counter' (x4) in agent/compaction_stats.py to module level — _row_signature runs per-row in the A-floor fallback path. Tests: 108 compaction-stats tests green under worktree PYTHONPATH. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(pricing): add claude-fable-5 snapshot rate ($10/$50, premium tier) Real claude-app/claude-fable-5 turns landed unpriced (cost_usd NULL) — flagged by the tokens-reprice-sweep on 2026-07-01. Add the official-docs snapshot entry at its premium tier ($10/M in, $50/M out, cache read $1.00, cache write $12.50; matches the OpenRouter catalog rate) so the engine prices it and the nightly sweep back-heals the NULL rows. Test: behavior contract through estimate_usage_cost across all notional providers + bare anthropic. * test: also assert empty/None provider (vendor-inference fallback path)
Replace duplicate private partition helpers with one consume-once _signature_partition helper and reuse it for hygiene and in-turn compaction stats. Verified: python -m pytest tests/agent/test_inturn_afloor_partition.py tests/agent/test_compaction_stats_reconcile.py -q -o addopts="" (97 passed); python -m pytest tests/agent/test_compaction_*.py tests/agent/test_inturn_compaction_stats_replay.py tests/agent/test_inturn_afloor_partition.py tests/run_agent/test_infinite_compaction_loop.py tests/run_agent/test_in_place_compaction.py tests/tui_gateway/test_compaction_status.py tests/context_engine/test_compaction_render_e2e.py -q -o addopts="" (259 passed); python -m ruff check agent/compaction_stats.py tests/agent/test_inturn_afloor_partition.py tests/agent/test_compaction_stats_reconcile.py (passed); git diff --check (passed). Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…del_ids merge (#166) Plugin provider catalogs namespace curated entries with the provider's own slug ("my-proxy/claude-x") while the provider's live /models endpoint returns bare ids ("claude-x"). The curated+live merge compared spellings with a naive lowercased exact match, so every model appeared TWICE per provider in every /model picker (desktop Models dialog, TUI, gateway) — adapters display the bare tail, so both rows render identically. Fix: canonicalize both sides on the provider's OWN "<slug>/" prefix only, matched on the "/" boundary. Deliberately NOT a blanket namespace strip — vendor-namespaced aggregator entries (openai/gpt-5 vs azure/gpt-5) that merely share a trailing name must not collapse. Same identifier-representation bug class as #135, different site. RED->GREEN: 4 behavior tests fail on the old merge loop, pass on the fix; plus over-collapse, live-only-discovery, live-first-order, and boundary guards. 200 adjacent picker/provider tests green.
A pasted/clipboard/dropped/picked composer image is written to the CLIENT's local Electron userData dir (saveImageBuffer / saveClipboardImage) even when the app runs against a REMOTE gateway. The preview readback, however, went through readDesktopFileDataUrl(), which in remote mode routes to the backend's GET /api/fs/read-data-url. The backend is a different machine and has no such file, so it returns 404 "File not found" — surfaced to the user as the "Image preview failed" toast on every paste. The attachment itself was unaffected (submit uploads the bytes via image.attach_bytes, which already reads the local path through the Electron bridge) — only the thumbnail preview was broken. Fix: add readComposerImagePreview(), which reads the preview from the LOCAL Electron bridge first and only falls back to the remote-aware read when the local bridge is unavailable or can't read the path (e.g. a path dragged from the remote file tree). This mirrors what the submit-time upload already does and keeps the local-mode behavior identical. Tests: 3 new cases in desktop-fs.test.ts assert remote-mode preview reads the local bridge and never touches the backend FS REST, plus the remote-file-tree fallback path. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…tail (#168) * feat(lcm): honor compression.target_ratio via a token-budgeted fresh tail LCM previously ignored compression.target_ratio entirely and kept a fixed 32-message fresh tail (~6.5K tokens on live traffic) regardless of context size. This maps the fleet-standard knob to a token budget: budget = target_ratio × threshold_tokens (= target_ratio × threshold × context_length) capped at lcm.fresh_tail_max_tokens (default 60K, guards 1M-window models) and 0.9 × threshold_tokens (convergence clamp — compaction can never no-op loop even at ratio=1.0). At compress() entry the dynamic count K is computed ONCE (frozen-K): walk backward accumulating count_message_tokens until the budget is spent, floored at the legacy fresh_tail_count (the tail can only widen, never shrink below today's behavior). All in-pass cut sites use len(list) - K — sound because the leaf-loop span is removal-only (the only list-growing transform, stub insertion via _sanitize_active_context_messages, runs strictly after the last cut; an AST source-contract test with planted- violation negative controls guards this invariant against future edits). Out-of-pass readers (preflight, _raw_backlog_messages) compute fresh. Also: - protect_last_n now tracks the dynamic K after each compaction so the compaction-stats consumer (find_inturn_kept_cut slack window) stays centered; pre-first-compaction value unchanged. - rotate_active_session preserves max(static, last-dynamic-K) rows so the lifecycle frontier can never advance past the wider in-memory tail. - fail-open fallback compressor now reads the configured target_ratio (was a hardcoded 0.20). - lcm_status exposes the resolved budget + last dynamic count for live verification. - disable switch: lcm.fresh_tail_token_budget_enabled: false (or env LCM_FRESH_TAIL_TOKEN_BUDGET_ENABLED=0), fail-safe ON for garbage values. Degenerate inputs (no context window, flag off, invalid ratio, estimator exception) reproduce the legacy fixed-count cut exactly. Ingest/replay/dedup paths untouched (PR #107 territory): zero edits to _reconcile_ingest_cursor_from_store, the scaffold-overlap guard, or the min-evidence bar at the replay-skip site. Verification: 45 new tests (config plumbing, budget resolution, chokepoint helper, frozen-K boundary consistency, AST guard + planted violations, rotate floor incl. window-grow, fail-open ratio); full context_engine + compaction-stats suite 348 passed; offline A/B shakedown on a mixed-density corpus at 200K: ON-arm tail 98 msgs / 36,946 tokens vs budget 37,500, legacy arm 32 msgs, 3 fires per arm, cadence_ratio 1.01 (gate ≤ 2.0). Spec: ~/.hermes/plans/2026-07-01_lcm-target-ratio-token-budget-tail-SPEC.md (v0.4, three Opus review passes: BLOCK → AWC → AWC, all blockers folded) * fix(lcm): Greptile review — env-path (0,1] clamp for LCM_TARGET_RATIO + AugAssign list-growth coverage in the frozen-K AST guard - LCM_TARGET_RATIO env override now applies the same (0,1] range guard as the config-file path (out-of-range → default 0.20), with tests. - The frozen-K AST guard also flags augmented-assignment list growth (working_messages += [...]) inside the leaf loop — scoped to message-list vars / list-literal RHS so integer counters don't false-positive — with a planted += negative control. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Tests-workflow wall time is 95% raw pytest (install is ~10s w/ warm uv cache). Slices are independent parallel jobs and the per-file subprocess model is import/IO-bound, not CPU-bound, so both knobs have headroom: - ci.yml: slice_count 8 -> 10 - tests.yml: HERMES_TEST_WORKERS=12 (was cpu_count*2 = 8 on 4-vCPU) - run_tests.sh: pass HERMES_TEST_WORKERS through the hermetic env -i (without this the env var silently never reached the runner) Expected: ~5min -> ~3min wall. This PR's own CI run is the measurement. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
#172) * fix(gateway): reconcile /compress feedback across both transcript axes The gateway /compress path compresses a chat-only projection of the stored transcript and, when the rewrite happens, drops every non-chat row (tool results / system / contentless turns) — usually the bulk of a tool-heavy session. But every user-facing number was computed on whichever axis was convenient, so a real 453K→59K compaction reported itself as: 🗜️ No changes from compression: 179 messages Chat size: ~31,406 tokens (unchanged; ...) Full request size: 453,542 → ~32,036 tokens ... — a self-contradicting message (observed live 2026-07-02). Fixes, per the honesty spec: - F1: summarize_manual_compression grows an opt-in enhanced mode (non_chat_count/tokens, transcript_rewritten, full_before_count) that reconciles BOTH axes: CASE A 'Compacted stored transcript: N → M' (chat compact, tool rows dropped), CASE B both axes shrank, CASE C 'No changes: transcript preserved (N messages: X chat + Y tool/system)'. Classic positional callers (CLI, TUI) are byte-identical. - F2: the after-size basis is now the rows the next request actually carries — the compressed list only when the store was rewritten, the ORIGINAL history otherwise (no more fabricated shrink on a no-op), and the no-rewrite full-request line says 'unchanged' instead of printing estimator noise as a delta. - F4: last_prompt_tokens is only zeroed when the transcript was actually rewritten; a true no-op keeps the provider-measured figure. - F3: the fixed overhead (resident agent's real system prompt + tool schemas) is resolved ONCE up front and used for BOTH the before and after full-request estimates — the before was previously measured with the memory-only temp agent's empty overhead while claiming to include system + tools. - New locale key gateway.compress.full_request_unchanged (all 16 catalogs). Tests: 7 new unit tests for the helper's cases, 2 new gateway E2E-style regressions (tool-heavy CASE A + CASE C incl. token-reset gating), all existing compress/i18n/CLI tests green (70 passed). * fix(review): fold Greptile P2s — all-or-nothing enhanced gate, CASE C token-line invariant, skip unused after-estimate - enhanced gate now requires non_chat_tokens too (no silent '~0 tokens reclaimed') - CASE C forces the unchanged token wording in the helper itself instead of trusting callers to pass equal before/after - gateway no-rewrite path skips the after-estimate entirely (the reply says 'unchanged' and never reads it; estimating over a tool-heavy history is not free) - 2 regression tests --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
#173) * feat(gateway): granular CompactionStats breakdown for manual /compress Follow-up to #172. The honest two-line form fixed the contradiction, but the rich per-bucket breakdown the auto-compaction announce renders (Messages / Context / Removed-from-live-context with the tool-result sub-split / Model / recovery pointer) was never wired into manual /compress — despite compaction_stats.py's docstring claiming it as a consumer. Ace asked for the detailed version. - After a real rewrite that changed the transcript, /compress now builds build_hygiene_stats over (history, chat-projection, compressed) and renders _format_granular_announce — the exact same reconciling renderer and validate() contract the hygiene announce uses. Example: 🗜️ Compressed: 7 → 3 stored messages Messages: 7 → 3 (kept 3 recent chat) Context: ~564 → ~49 tokens (freed ~515, 91% smaller) Removed from live context (5 messages): • 2 tool-result messages → ~502 reclaimed (raw tool output) • 1 other messages → ~19 reclaimed (system + tool-call turns, cleared) • 2 folded messages → ~22 reclaimed (older chat condensed into 1 summary) Model: claude-app/claude-opus-4-8 ↩ previous transcript preserved: sess-1 (searchable via session_search) Full request size: 453,542 → ~49 tokens (...) - Store-correct recovery pointer: lcm.db (lcm_grep/lcm_expand) for the LCM engine, session_search + old session id for a rotated built-in store. - Degrades to the #172 two-line enhanced form on ANY stats build/validate failure (a reconcile bug can never break /compress or ship wrong math), and on no-op passes (len unchanged) where there is nothing to break down. - 2 new gateway regressions: full granular render (buckets, sub-split, model line, recovery pointer) + forced-failure degradation. * fix(review): classify built-in SUMMARY_PREFIX rows as summary; warn on granular build failure Greptile P2s on #173: - _is_summary_message only matched the LCM summary regex, so a built-in compressor summary row ([CONTEXT COMPACTION — REFERENCE ONLY] ...) classified as 'kept chat' and the Messages line miscounted (kept N instead of kept N-1 + 1 summary). Now the built-in marker is detected too; regression assert added (kept 2 recent chat + 1 summary). - Unexpected exceptions in the granular stats build now log at WARNING (parity with the validate() failure path); DEBUG hid real regressions. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…nscript basis (#174) * fix(gateway): label manual /compress granular block by its stored-transcript basis The granular CompactionStats renderer (#173) was written for the auto- compaction announce, whose input IS the live message list — so 'Context:' / 'Removed from live context' / 'kept in context' accurately describe wire savings there. Manual /compress reuses the same renderer but feeds it the STORED transcript (session_store.load_transcript), which under LCM has already had its bulk compacted OFF the wire in earlier passes. The reduction is storage reclaimed, not request-size — so the block overstated wire savings (Ace's live case: '~689K → ~37K, freed 651K, 95% smaller' against a real 303K request). Make _format_granular_announce basis-aware via a keyword-only arg (default 'live' → byte-identical auto-announce output). Manual /compress passes basis='stored', rendering 'Stored transcript:' / 'Removed from stored transcript' / 'reclaimed' / 'kept in transcript'. The separate provider-measured 'Full request size:' line stays the wire truth beside it, so the user sees both storage-reclaimed AND real request-size cut. Numbers unchanged (build_hygiene_stats / CompactionStats / validate untouched) — only labels. Adds TestGranularBasisLabel (default==live, stored relabels without changing any number, renderer never emits the wire line) + an LCM-engine E2E through the real _handle_compress_command reproducing Ace's exact scenario. * test: assert kept-in-transcript unconditionally when Replacement cost present (Greptile P2 #174) --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
… aren't all "tui" (#175) * feat(tui_gateway): client-identity source so desktop/dashboard/mobile aren't all "tui" The desktop app, the in-browser dashboard chat, and the Ink stdio TUI all drive the same tui_gateway JSON-RPC server, which hardcoded platform="tui" on every agent it built. So every desktop-app turn recorded platform="tui" in blackbox (turns.platform) — the source dimension on the token/cost dashboard couldn't tell the desktop client from the terminal, and a future mobile app would fold in too. Let a client self-declare a `source` on session.create/resume; thread it through _make_agent -> AIAgent(platform=...) (which already feeds the blackbox on_session_end hook), and store it on the session dict so the deferred/lazy build, /new reset, resume, and branch paths all preserve it. `source` is untrusted client input feeding a persisted DB dimension, so it is sanitized to a conservative slug (`^[a-z][a-z0-9_-]{0,23}$`) with a `tui` fallback — the label stays open-ended (new clients name themselves; no central enum to keep in sync, per the existing "fixed set of platform names goes stale" note). Desktop app: stamp source:'desktop' once, centrally, in the requestGateway wrapper for the session-originating methods (create/resume) — never overriding an explicit caller source (e.g. the dashboard sidecar's 'tool'). No schema/blackbox change (the platform pipe already existed); stdio Ink and internal rebuilds keep the historical "tui" default so nothing regresses. Historical desktop turns are indistinguishable from real TUI turns, so the clean split starts at deploy (no backfill possible). Tests: tui_gateway/test_client_source_platform.py (sanitizer table + _make_agent platform derivation); desktop use-gateway-request.test.tsx (injection contract: stamps create/resume, skips other methods, never overrides explicit source). * test: accept _make_agent(source=…) in remaining narrow test fakes The new source kwarg on _make_agent tripped three narrow monkeypatch fakes in tests/test_tui_gateway_server.py (session_id/session_db-only signatures). Widen them to **_kw so they tolerate the new kwarg, matching the sibling fakes that already splat. No behavior change. * fix(tui_gateway): reject non-str source (Greptile P2) _sanitize_client_source used str(raw or ''), so a JSON-RPC boolean true would coerce to the slug 'true' instead of falling back to 'tui'. Guard with isinstance(raw, str) up front; add True/0/[] cases to the sanitizer table. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…ead, one number story (#176) The manual /compress granular block led with the stored-transcript estimate (Stored transcript: ~328K -> ~43K, chars/4 over the archive) while the footer showed the REAL provider-measured wire count (e.g. 236K). Two numbers, two populations, two tokenizers -> they disagreed by 100K+ every time and confused the operator (and the agent explaining it). The real provider-measured before-count already exists at /compress time (session_entry.last_prompt_tokens). This makes the granular block lead with it. Change (stored basis only, when a real measured count is available): - Context line becomes the WIRE story: measured before (exact, no ~) -> next-request estimate (~), freed/pct computed on wire numbers, tagged "before measured, after next-request estimate". - Stored-transcript totals are demoted into the "Removed from stored transcript" header as an explicit "token-est reclaimed from archive" parenthetical, so they can never be read as request-size savings. - The trailing "Full request size:" line is skipped when the block already carried the wire story (no double-reporting). Fallbacks unchanged and covered: - No real measured count (post-restart) -> prior stored-basis rendering. - Auto-compaction announce (basis="live") -> byte-identical; wire kwargs are ignored on the live basis. - No-rewrite / two-line forms -> unchanged. New renderer kwargs wire_before/wire_after are honored only for basis="stored" with both > 0. 11 new renderer tests + 2 updated gateway E2E tests (the LCM case reproduces Ace's exact 303,201 scenario). 217 compaction/compress tests green. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…self-identifying markers (#177) * fix(compaction): gate in-turn stats on announce render-eligibility + self-identifying markers The in-turn compaction-stats block (build_inturn_stats + gross-error guard + COMPACTION_STATS_* WARNINGs) ran for EVERY LCM compress() call — including no-ops — while the announce formatter default-denies non-rendering statuses (noop/idle/running/bypassed; conditional statuses failing the post<pre render check). Result: ~100% kept_tail APPROX_ATTRIBUTION and TAG_MISSING WARNINGs fired for announces that never render — pure log noise that polluted the daily compaction-stats watcher report (measured 2026-07-02: every real-session degrade marker in apollo's agent.log rode an 'LCM compression no-op' line). Fix: _inturn_stats_render_eligible() mirrors the formatter's own gating by consuming the SAME allow-list objects (single source of truth, no copied literals) + the conditional token check, applied on the exact variables the announce call passes as pre_tokens/post_tokens. LCM-scoped: the built-in path keeps its sid-rotation gating and may not expose _last_compression_status. Also (D-4): every compaction-stats marker now carries session=<id>, plus src=test under pytest (PYTEST_CURRENT_TEST), so the log watcher can attribute markers by the line's own fields instead of fragile proximity joins — pytest writing through the live logging config previously inflated production counts. Tests: 13 new (gate predicate matrix vs the real formatter as oracle; drift guards asserting the gate uses the formatter's allow-list objects and the announce-call token variables; stats=None render tolerance; marker session/ src=test fields incl. from a threading.Thread worker — the measured polluter's shape; throttle unaffected). Spec (4 Opus review passes): plans/2026-07-02_inturn-noop-gate-multipass-provenance-SPEC.md v0.5, PR-A. * fix(compaction): don't suppress non-LCM stats in the render gate (Greptile #177) The render gate was written as (_engine_name == 'lcm') and render_eligible(...), which short-circuits to False for the built-in (non-LCM) compressor and permanently dropped its stats — silently degrading every built-in announce to the two-line form. Intent was to gate LCM no-ops only. Restructure to an explicit branch: LCM → render-eligibility; non-LCM → always eligible (unchanged always-attempt behavior). Also closes the Greptile test-coverage gap: _run_announce_block now mirrors the real LCM-scoped gate and takes explicit pre/post tokens, and new tests exercise (a) a conditional status ('sanitized') with a REAL compression post<pre through build+validate, (b) a conditional no-op not built, and (c) non-LCM not suppressed by the gate. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
… extension) (#178) * feat(compaction): multi-pass provenance stamp, shadow-first (Option B extension) The Option B _src_idx provenance stamp (#110) — the EXACT kept/folded partition — was gated to single-pass compactions (leaf_passes == 1). Heavy multi-pass sessions (dynamic leaf chunking, deferred maintenance) never got the stamp, fell to whole-tail replay, and on sanitize-heavy tails degraded to the A-floor's approximate split → the honest two-line announce instead of the granular breakdown. Widen the stamp gate to leaf_passes >= 1. Soundness (spec §0.6, five-step line-cited proof): ingest is a 1:1 order-preserving rewrite; every in-loop mutation removes rows ONLY from the front region (scaffold drop + compacted- chunk removal both slice strictly before fresh_tail_start; the leaf chunk is a contiguous FRONT prefix); summary rows go to the DAG, never into working_messages (summaries are prepended inside _assemble_context, after the stamp); frozen-K means 'the last K rows' names the same physical rows after every pass; and the stub-inserting sanitize (_sanitize_tool_pairs — single caller: _sanitize_active_context_messages) runs only inside/after _assemble_context, strictly post-stamp. Hence the fresh tail is a contiguous suffix of the ORIGINAL messages for any pass count, and the end-anchored mapping is valid. Stamp guard hardened: role + tool_call_id + tool_calls arity (structural fields that survive content rewrites). SHADOW-FIRST (observe before trust): the consumer trusts single-pass stamps exactly as before; for multi-pass it computes B, displays the current replay/A-floor behavior byte-identical, and compares via an index-SET predicate — diverge <=> sorted(B kept-origin indices) != sorted(current kept- origin indices), never a token sum (an equal-token wrong-but-in-range swap changes the set but not the sum). Logged as COMPACTION_STATS_B_MULTIPASS_SHADOW agree (INFO) | diverge (WARNING) for the soak; a follow-up trust-flip PR promotes multi-pass B to primary after a clean soak. Engine exposes last_leaf_passes for the trust decision. Tests: multi-pass stamp exactness across fold counts + boundary shifts + scaffold drops; structural-guard rejection; shadow display byte-equality; the diverge branch fires (incl. an EQUAL-TOKEN wrong-but-in-range fixture that fails any summing comparator); corrupt stamps fall to A-floor; compare exceptions never break the build; single-pass byte-identical. Spec (4 Opus review passes, zero blockers at close): plans/2026-07-02_inturn-noop-gate-multipass-provenance-SPEC.md v0.5, PR-B. * fix(compaction): un-throttle the shadow diverge warning (Greptile #178) The diverge arm of the multi-pass shadow compare used _warn_compaction_stats_once, which fires at most once per (cause, session). A heavy session that diverges multiple times would drop all but the first, so the soak analysis couldn't measure within-session diverge frequency or tell a one-off fluke from a systematic disagreement. Replace with a direct logger.warning so every diverge event is independently observable; it carries the same session=/src=test fields for watcher attribution. Tests: assert two diverges in one session both emit (not throttled) + a source-structure guard that the diverge arm uses logger.warning, not the throttled helper. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Replace the /api/sessions/stats by_source hydrate with a SessionDB GROUP BY helper and run the blocking stats read in the default executor.\n\nVerified with scripts/run_tests.sh tests/test_web_server.py tests/test_hermes_state.py (308 passed). Post-spike read-only live-state measurement: stats max 1.314s, trivial endpoint max 0.597s, fd count returned to baseline. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Route /compress and session-reading slash commands through the live gateway session instead of the slash worker, add the manual compression no-op/concurrency guard, and block local file reveal/rename/delete for remote gateway files.\n\nVerified:\n- PYTHONPATH=$PWD ~/.hermes/runtime/hermes-agent/venv/bin/python -m pytest tests/test_tui_gateway_server.py -k 'compress_history_identity_noop or compress_history_same_length or slash_exec_compress_bypasses or slash_exec_compress_reports_existing or command_dispatch_compress_returns or live_read_commands' -q\n- PYTHONPATH=$PWD ~/.hermes/runtime/hermes-agent/venv/bin/python -m pytest tests/test_tui_gateway_server.py -k 'not browser_manage_connect_default_local_reports_launch_hint' -q\n- npm run test:ui -- src/app/session/hooks/use-prompt-actions/index.test.tsx src/store/file-actions.test.ts\n- npm run typecheck\n- npm run lint -- src/app/session/hooks/use-prompt-actions/index.test.tsx src/app/session/hooks/use-prompt-actions/slash.ts src/store/file-actions.ts src/store/file-actions.test.ts\n- git diff --check Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Route remote desktop file opens through an authenticated bytes-to-temp path instead of token/download URLs or client-local file:// fallbacks. Add private temp lifecycle helpers with start/periodic TTL sweeps and harden remote HTML previews with blob URLs, CSP, and JavaScript disabled.\n\nVerified:\n- node --test apps/desktop/electron/gateway-temp-files.test.cjs\n- npm --workspace apps/desktop exec vitest run --environment jsdom src/lib/media.remote.test.ts src/lib/html-preview.test.ts src/lib/local-preview.remote.test.ts\n- npm --workspace apps/desktop exec eslint src/lib/media.ts src/lib/media.remote.test.ts src/lib/html-preview.ts src/lib/html-preview.test.ts src/lib/local-preview.ts src/lib/local-preview.remote.test.ts src/app/chat/right-rail/preview-pane.tsx electron/gateway-temp-files.cjs electron/gateway-temp-files.test.cjs electron/main.cjs\n- npm --workspace apps/desktop run typecheck\n- npm --workspace apps/desktop run test:desktop:platforms Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Offloads the remaining dashboard SessionDB read sites from async handlers through sync helpers and adds Semaphore(2) bounding for heavyweight reads. Adds thread-assert coverage plus an AST-aware regression lint with bare/lambda/partial/inner-def/to_thread self-tests. Verified: scripts/run_tests.sh tests/test_web_server_sessiondb_eventloop.py tests/test_web_server.py tests/test_hermes_state.py -q; scripts/run_tests.sh tests/hermes_cli/test_dashboard_auth_status_endpoint.py -q. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
The /reset and /new banner (_format_session_info) surfaces Model, Provider,
Context, and Endpoint so a user can immediately see the fresh session's
config. It omitted the reasoning effort, which is independently configurable
(agent.reasoning_effort) and just as material to cost/behavior.
Add a '◆ Reasoning: <level>' row between Provider and Context, resolved from
_load_reasoning_config():
- None -> 'medium (default)'
- {enabled:False} -> 'none'
- {enabled:True} -> the effort level (minimal/low/medium/high/xhigh)
Global config (not the per-session override) is correct here: the banner
renders on /reset and /new, which clear any /reasoning session override, so
config.yaml is the effective level at reset time. Resolution is wrapped so a
failure omits the row rather than breaking the banner (mirrors the existing
Endpoint conditional-append).
Tests: 5 new cases in test_session_info.py covering each reasoning state,
row ordering, and the fail-open path.
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…op PRD Phase 3) (#185) * fix(desktop): gate remote media and file reveals Route remote a/v markdown attachments through the gateway temp-file opener before local stream fallbacks, add the desktop remote file-boundary eslint rule/self-test, and show an honest remote reveal message.\n\nVerified:\n- npm run test:ui -- eslint.config.test.mjs src/components/assistant-ui/markdown-text.test.ts src/lib/desktop-fs.test.ts src/store/projects.test.ts src/lib/media.remote.test.ts\n- npx eslint src/components/assistant-ui/markdown-text.tsx src/components/assistant-ui/markdown-text.test.ts src/lib/desktop-fs.ts src/lib/desktop-fs.test.ts src/store/projects.ts src/store/projects.test.ts src/store/file-actions.ts src/components/assistant-ui/directive-text.tsx src/components/chat/generated-image-result.tsx src/app/chat/right-rail/preview-file.tsx src/i18n/en.ts src/i18n/zh.ts src/i18n/zh-hant.ts src/i18n/ja.ts src/i18n/types.ts\n- npm run typecheck\n- git diff --cached --check * fix(desktop): reconcile B1/D1 remote file-action guard seam D1's facade-throw guard (desktop-fs *File exports) is the enforced safety net; restore B1's entry-point blockRemoteFileAction guards (union, defense-in-depth) on all five entry points incl. executeFileRename/Delete for the data-loss bug. Update file-actions.test.ts to the *File facade mock names. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Treat widthOverride as persistent per the pane-store feature commits and add a module-reload assertion that persisted width loads back into the pane store.\n\nVerified:\n- npm --prefix apps/desktop run test:ui -- src/store/panes.test.ts\n- npm --prefix apps/desktop run typecheck\n- npm exec prettier -- --check src/store/panes.test.ts (from apps/desktop)
Add a server-owned pinned flag to SessionDB, expose it through REST/session lists and the session.pin gateway RPC, and derive desktop pinned rows from server session data instead of localStorage. Verified: scripts/run_tests.sh tests/test_hermes_state.py tests/test_tui_gateway_server.py tests/test_web_server_sessiondb_eventloop.py; npm --workspace apps/desktop run typecheck; npm --workspace apps/desktop run test:ui -- src/store/session.test.ts src/store/layout.test.ts src/store/updates.test.ts. Note: npm --workspace apps/desktop run test:ui was also attempted after a temporary worktree node_modules symlink; it failed in desktop suites outside the touched pinned-session files (46 files / 21 tests).
The server-side pin PR (#186) made $pinnedSessionIds a read-only computed atom, which dropped setPinnedSessionOrder + the sidebar's reorderPinned handler and hardcoded sortable={false} on the pinned list — losing drag-to-reorder. Restore it with a local $sidebarPinnedOrderIds order atom (mirrors $sidebarSessionOrderIds) layered over the server-synced pin SET via orderByIds: pin membership syncs across devices, visual order stays per-device. Keyed by durable lineage-root pin ids so order survives compression.
* feat(dashboard): web chat client self-declares source='dashboard' The in-browser dashboard chat (web/), the Electron desktop app, and the Ink TUI all drive the same tui_gateway JSON-RPC server, which stamps platform='tui' by default. PR #175 taught the Electron app + backend to thread a client 'source' so turns are attributed per-client (desktop/dashboard/mobile) on the token dashboard — but the WEB dashboard client (web/src/lib/gatewayClient.ts) was never updated, so browser-dashboard turns still recorded as 'tui'. Override GatewayClient.request to stamp source:'dashboard' on the session-originating methods (session.create/session.resume) when the caller hasn't set an explicit source (the sidebar sidecar session deliberately sends source:'tool'). Mirrors the Electron app's requestGateway injection exactly. Backend already sanitizes + threads source -> AIAgent(platform=...) (#175); no backend change needed. Test: web/src/lib/gatewayClient.test.ts (4) — stamps create/resume, skips other methods, never overrides an explicit source. * ci: re-trigger checks (apps/desktop/node_modules ENOTDIR runner flake, not a code issue)
…emote-mode correctness, pinned-session sync) (#191) As-built specs with full review logs (eventloop: 4-pass super-pass; desktop remote: 9-pass super-pass; pin-sync: 1-pass light track), Phase-0 findings, live AC-2 results, and residual follow-ups (session.list projection cost, BUG-E facade, transient-dirty build stamp).
By default the desktop app persists the last composer model pick to localStorage and carries it across new chats (Cmd+N) and restarts, so picking an expensive model makes it the sticky default for every new session. This adds an opt-in config knob to instead reseed a fresh draft's model from the profile default. - HermesConfig type gains desktop.reset_model_on_new_session (optional). - useHermesConfig mirrors it into a new $resetModelOnNewSession atom. - refreshCurrentModel treats a fresh draft as a forced reseed when the flag is on, overwriting the sticky pick with the profile default. A live session still owns its model (never disturbed mid-conversation). - The fresh-draft effect loads config before the model reseed so the flag is in the store first (avoids a cold-boot race). Default false preserves the existing sticky-last-pick behavior. Backend needs no change: /api/config already returns the whole config, so the desktop section flows through untouched. Tests: use-model-controls (flag on reseeds, flag off keeps pick, active session untouched, explicit force still works) + use-hermes-config (mirrors true, defaults false when absent). Full desktop UI suite green (1283 passed), typecheck clean. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
… zombie writes (#341) * fix(persist): append-time generation gate — suppress a /stop'd turn's zombie writes A turn interrupted by /stop is a cooperative interrupt: the inner retry/fallback loop can keep spinning across fallback providers, and its incremental flush keeps appending rows to the session DB *after* the user stopped the turn (the #666915-666942 incident: 14 assistant(tool_calls) + 14 tool results written post-/stop). Those zombie rows let a later /undo land somewhere the drain then clobbers. Add an append-time generation gate at the single persist chokepoint (_flush_messages_to_session_db). When the gateway invalidates a turn's run generation (/stop, /new), it sets agent._persist_superseded; the flush then suppresses that turn's continued NEW content/tool rows. PAIRING-SAFE (never orphans a tool call → NousResearch#48879 role-alternation corruption): - gate runs AFTER all already-durable skips, so a pre-/stop assistant(tool_calls) is never re-evaluated and its result passes (B1) - the suppressed-tool-call-id set is AGENT-SCOPED, surviving the whole drain, because the assistant(tool_calls) row and its tool result flush in SEPARATE flushes within one iteration — a per-flush set would orphan the result (B1') - the interrupt-close tail (finish_reason=interrupt_close) is ALWAYS persisted — it is the auto-continue / restart-loop signal, and the reader _is_interrupt_close_tail consumes only agent_history[-1] (I1) Fail-open (I5): _persist_superseded resolved once via getattr-default- False in try/except — a guard error persists normally (a stray late row is cosmetic; #339 already stops /undo racing it; a dropped real row is data loss). Best-effort: the flag-set (event loop) races the flush (executor thread); reaper path scoped out (R6). logger.info on suppress. Tests (all RED-proven): 8 in test_persist_superseded_gate.py (same-flush pair, B1 owner-durable, B1' cross-flush, I1 tail, I1 synthetic-repair, I3 byte-identical, I5 fail-open, AC1 suppress) + 2 in test_undo_drain_guard.py (Phase-2 flag-set, R6 reaper-unset). 3 Opus review passes: BLOCK(B1) -> BLOCK(B1',B2') -> APPROVE-WITH-CHANGES. Spec: ~/Projects/hermes-append-gen-gate/SPEC.md v1.3. * fix(persist): gate fails open on unexpected non-assistant/tool rows (Greptile-P2) The else-branch suppressed ALL non-tool roles. A zombie turn only writes assistant+tool rows, but if scaffolding ever injects a NEW user/system row mid-drain, dropping it would be data loss. Suppress only assistant rows; fail OPEN (persist + debug-log) on any other role. Adds test_unexpected_new_user_row_fails_open_during_supersede. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(blackbox): price MoA turns by physical model routes * fix(blackbox): preserve turn accumulator append invariant * fix(blackbox): keep MoA telemetry failures isolated --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(terminal): make snapshot temp files portable * fix(tests): make the full suite hermetic on macOS --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
* fix(kanban): clear respawn guard for closed PRs Resolve recent GitHub PR state with a per-tick capped memo, retain fail-safe guarding on errors, and surface guarded tasks in CLI, gateway logs, and stranded-ready diagnostics. Verified: scripts/run_tests.sh on 4 targeted files, 309 tests passed. * fix(kanban): deduplicate respawn guard PR URLs Avoid repeated resolver work when steering comments quote the same PR. Verified: targeted kanban suite, 310 tests passed. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…time (#348) `cron/jobs.py` froze `CRON_DIR`/`JOBS_FILE`/`OUTPUT_DIR` (and the ticker heartbeat paths) as module-level constants at *import* time, and `_current_cron_store()` returned those frozen paths whenever no explicit `use_cron_store()` override was active. So if the module was imported while `HERMES_HOME` pointed at the real `~/.hermes` — e.g. a real agent booting from a worktree/e2e harness that shares the live home — a *later* `HERMES_HOME` redirect (the pytest hermetic fixture, or an e2e harness pointing at a tempdir) was silently ignored, and `create_job` / the `cronjob` tool wrote straight into the LIVE `cron/jobs.json`. This actually happened: a real-agent blackbox e2e session booting from a shared-home worktree wrote 11 unit-test fixture jobs (7x name=brief every-60m, 3x "claim job", 1x "paused job") into the production cron store, re-arming cron-config-lint errors and paging cron-health. Fix: `_current_cron_store()` resolves `get_hermes_home()` LIVE when there is no explicit `use_cron_store()` override, instead of returning the import-frozen constants. The same import-time freeze afflicted the ticker heartbeat/success paths (sibling of the same bug class), so those now resolve live too. The module-level constants remain as a back-compat surface for existing importers (e.g. hermes_cli/backup.py's JOBS_FILE reference). Migrated the 5 tests that coupled to the removed mechanism (patching `cron.jobs.CRON_DIR` / `TICKER_HEARTBEAT_FILE` directly) to redirect `HERMES_HOME` — the supported isolation path, and a behavior contract rather than a change-detector on internals. This includes the concurrency-isolation test `test_profile_call_cannot_retarget_ticker_store_mid_write`, whose invariant (a per-thread profile override must not retarget a concurrent ticker's store) is preserved: the profile call uses `use_cron_store()` (a per-thread ContextVar), the ticker uses no override and now resolves the live default home. Regression test `tests/cron/test_store_resolves_live_home.py` reproduces the leak (a HERMES_HOME set after import must be honored; a job must not land in the stale import-time home) and is RED against the frozen-constant implementation, GREEN with the fix. Full cron suite: 752 passed. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
fix(desktop): distinguish xhigh and max reasoning
The desktop renders a sent turn optimistically (client-minted ids user-<ts>/assistant-<ts>) and stamps them to the committed DB ids carried on the message.complete frame. Two defects made stamping a no-op, so the live session-sync poll re-fetched the committed rows, failed to recognize them, and appended them as DUPLICATES (every message shown twice; DB stays clean): 1. Server (tui_gateway/server.py): message.complete never included the committed ids. Add _completed_turn_committed_ids() to extract [user_id, assistant_id] for the finished turn from session history (_db_persisted_row_id) and ship them as payload.message_ids. 2. Client (use-session-changes.ts): the stamp predicate only matched user-prefixed OR pending rows, but completeAssistantMessage() clears the streamed assistant row pending flag BEFORE markTurnComplete() stamps it, so the completed assistant row (id assistant-<ts>, pending:false) was skipped and duplicated. Recognize any optimistic id (user-/assistant- prefix) OR pending. Gated by tests that FAIL on the naive code: tests/tui_gateway/test_completed_turn_committed_ids.py (6) + two new cases in use-session-changes.test.ts proving the already-completed assistant row stamps and does not duplicate end-to-end.
Kyzcreig
force-pushed
the
fix/desktop-optimistic-stamp-dup
branch
from
July 15, 2026 22:56
33429b5 to
0fc8088
Compare
Collaborator
|
Thanks for tracing the optimistic-row failure mode. This automated hermes-sweeper review found that current
Closing as implemented on main: the duplicate-producing append/poll seam described here no longer exists. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Symptom
In the desktop app (remote mode), every message in a session renders twice — reproducible by opening a new session and sending a single message. The duplication is display-only: the session DB stores each message exactly once (verified: N rows, N distinct contents). A reload shows the transcript correctly.
Root cause (two defects on the optimistic-stamp seam)
The desktop renders a sent turn optimistically with client-minted ids (
user-<ts>/assistant-<ts>) and relies on themessage.completeframe carrying the committed DB row ids to stamp those optimistic rows to their persisted integer ids. The live session-sync poll (session.changes) then recognizes the committed rows by id and skips them. Two defects broke that contract, so the poll re-appended every committed row as a duplicate:Server (
tui_gateway/server.py) — themessage.completepayload never included the committed ids ({text, usage, status, +reasoning/warning/rendered}only). SoextractCommittedMessageIds()returned[],markTurnComplete()early-returned, and no stamping happened.Client (
use-session-changes.ts) — even with ids present, the stamp predicate only matcheduser-prefixed orpendingrows. ButcompleteAssistantMessage()clears the streamed assistant row'spendingflag beforemarkTurnComplete()runs, so the completed assistant row (idassistant-<ts>,pending:false) was skipped → kept its optimistic id → duplicated by the poll. The existing unit test masked this by passingpending:true, a state that doesn't occur at runtime.Fix
_completed_turn_committed_ids(history)— extracts[user_id, assistant_id]for the finished turn fromsession["history"]via the_db_persisted_row_ideach message carries (stamped inrun_agent._flush_messages_to_session_db) — and ship it aspayload["message_ids"]. Tool rows are excluded (they're not optimistic client rows; the poll handles them by their own ids), matching the client's two-optimistic-row contract.user-/assistant-prefix) or a still-pending row in bothstampOptimisticTranscriptRowsandoptimisticTranscriptIds.Tests (fail on the naive code, pass on the fix)
tests/tui_gateway/test_completed_turn_committed_ids.py(6): order, tool-loop turn (returns[user, final-assistant-text]), empty-tool-calls-row exclusion, missing/non-int id skipping, invalid input.use-session-changes.test.ts(2 new): an already-completed (pending:false) assistant row still stamps; end-to-end no-duplicate when the poll returns the own committed rows.Verified RED→GREEN on both halves; existing suites green (session_changes 7, finalize_persist 11, desktop chat-hooks + message-stream 38). Server helper additionally verified against real session data (returns
[user_id, assistant_id]in order from an actual completed desktop turn).