Skip to content

feat: /compress with focus topic in Web UI (replace /compact, add transcript-aware UX) - #619

Closed
franksong2702 wants to merge 14 commits into
nesquena:masterfrom
franksong2702:codex/issue-469-compress
Closed

franksong2702 wants to merge 14 commits into
nesquena:masterfrom
franksong2702:codex/issue-469-compress

Conversation

@franksong2702

Copy link
Copy Markdown
Contributor

Thinking Path

  • Hermes WebUI should support manual conversation compression in a way that feels explicit and stable, closer to the Hermes CLI mental model.
  • Issue feat: /compress with focus topic in Web UI (replace /compact) #469 asked for /compress [focus topic] as the canonical manual compression entry point, with /compact kept only as a compatibility alias.
  • The backend path already existed, but the WebUI still lacked a transcript-aware interaction for command, progress, completion, and persisted context handoff.
  • The main UX challenge was not just triggering compression, but making the command understandable in the transcript, preserving a visible context-compaction handoff, and keeping that handoff stable across continued chatting and reloads.
  • This PR adds that manual compression interaction and aligns the command semantics, visual states, and persistence behavior with the current transcript flow.

What Changed

Command behavior

  • Promoted /compress [focus topic] as the primary manual compression command.
  • Kept /compact as a legacy alias for compatibility.
  • Updated command copy and i18n strings so /compress matches CLI wording and /compact is clearly labeled as the old alias.

Manual compression UX

  • Added an explicit transcript-visible manual compression flow with:
    • Command
    • Compressing
    • Compression complete
    • Context compaction / Reference only
  • Running state appears immediately when /compress is sent.
  • During compression, new user input is queued instead of interrupting the current compression request.
  • The context-compaction handoff is rendered as its own card instead of appearing like a normal assistant message.
  • The reference card is collapsed by default, expandable, and copyable.

Transcript anchoring / persistence

  • While compression is active and immediately after completion, the compression cards behave like part of the current transcript and continue to move upward as new turns are added.
  • After refresh/reopen, transient cards disappear while the Reference only card persists.
  • Added session-level compression anchor metadata so the persisted reference card remains near the point where compression occurred instead of drifting to an incorrect earlier location.
  • The completed block now anchors to the end of the compressed transcript so follow-up turns naturally appear beneath it.

Backend integration

  • Manual compression continues to use the direct POST /api/session/compress path.
  • focus_topic is passed through to the context compressor.
  • Session state now persists compression anchor metadata used by reload-time placement.

Why It Matters

This makes manual compression feel like a first-class interaction rather than a hidden backend action.

Users now get:

  • a clear command entry point
  • immediate progress feedback
  • a stable context-compaction handoff they can inspect later
  • focus-topic-aware compression
  • safer follow-up behavior when they continue chatting during or after compression

That brings the WebUI much closer to the intended CLI-style manual compression workflow.

Verification

Automated

  • node --check static/ui.js
  • node --check static/commands.js
  • node --check static/messages.js
  • uv run --with pytest --with requests --with pyyaml pytest tests/test_sprint46.py -q

Manual

Verified in the browser that:

  • /compress shows immediate running feedback
  • /compress [focus topic] preserves focus-topic context
  • /compact still works as a compatibility alias
  • input during compression is queued instead of interrupting the compression request
  • the Reference only card is collapsed by default, expandable, and copyable
  • after refresh/reopen, transient cards disappear and the reference card remains
  • after compression completes, follow-up messages continue below the compression block and push it upward naturally

Risks / Follow-ups

  • This is user-visible transcript behavior, so additional polish may still be warranted as transcript UX evolves.
  • Compression summary quality still depends on the underlying Hermes context compressor and summary model behavior.
  • The PR intentionally focuses on the manual compression interaction only; it does not attempt to redesign unrelated transcript cards.

Model Used

  • Provider: OpenAI Codex
  • Model: GPT-5.4
  • Notable usage: local coding agent workflow with terminal-based inspection, patching, transcript UX iteration, and local verification

Images

This is a net-new interaction rather than a replacement of an existing UI flow, so the attached images show the main states of the feature rather than a traditional before/after pair:

  • running state
  • completed state
  • queued follow-up / continued conversation state
  • reload / persisted reference state

(Images to be attached in the GitHub UI.)

@franksong2702

Copy link
Copy Markdown
Contributor Author
3 - queued follow-up : continued conversation state 4 - reload:persisted reference state 2- complete state 1 - running state

@franksong2702
franksong2702 force-pushed the codex/issue-469-compress branch from 27be242 to a7169f7 Compare April 17, 2026 05:53
@franksong2702 franksong2702 changed the title feat(compress): align manual compression cards with transcript flow feat: /compress with focus topic in Web UI (replace /compact, add transcript-aware UX) Apr 17, 2026
@franksong2702

Copy link
Copy Markdown
Contributor Author

Looping in @aronprins here as well since a big part of this PR is the transcript-side UX for manual /compress. Would love your take on the interaction/placement when you have a chance.

@franksong2702
franksong2702 force-pushed the codex/issue-469-compress branch from a7169f7 to 474a321 Compare April 17, 2026 06:56
@aronprins

Copy link
Copy Markdown
Contributor

@franksong2702 looking good!

For the compression complete message ,can you do it similar to a "thinking" card but colored green (not sure if there is a success state yet, if not add it based on existing color schemes) so that it too collapses once completed?

@franksong2702

Copy link
Copy Markdown
Contributor Author

@franksong2702 looking good!

For the compression complete message ,can you do it similar to a "thinking" card but colored green (not sure if there is a success state yet, if not add it based on existing color schemes) so that it too collapses once completed?

yeah I see what u mean……working on it

@franksong2702

Copy link
Copy Markdown
Contributor Author

@aronprins implemented this direction on the compress flow:\n\n- the completion state now uses the transcript-style card treatment instead of a one-off toast\n- it has the green success styling and collapsible behavior like the thinking card pattern\n- the running state was also aligned to the transcript dot treatment so the whole manual flow reads consistently in the timeline\n\nThe updated UX is now part of the draft PR and matches the transcript redesign much more closely.

@franksong2702

Copy link
Copy Markdown
Contributor Author

@aronprins implemented this direction on the compress flow:

  • the completion state now uses the transcript-style card treatment instead of a one-off toast
  • it has the green success styling and collapsible behavior like the thinking card pattern
  • the running state was also aligned to the transcript dot treatment so the whole manual /compress flow reads consistently in the timeline

The updated UX is now part of the draft PR and matches the transcript redesign much more closely.

@franksong2702

Copy link
Copy Markdown
Contributor Author
截屏2026-04-17 17 52 43 截屏2026-04-17 17 54 14 these are screenshots after the fix

@aronprins

Copy link
Copy Markdown
Contributor

截屏2026-04-17 17 52 43 截屏2026-04-17 17 54 14 these are screenshots after the fix

Much better, however I see the command tool call (or whatever that is) has a differnt bottom margin - is that on your branch or global?

@franksong2702

Copy link
Copy Markdown
Contributor Author
截屏2026-04-17 18 54 18 compressing 截屏2026-04-17 18 55 13 finished compression 截屏2026-04-17 18 55 32 reloading session 截屏2026-04-17 18 58 05 and the last one: continuing conversation

@franksong2702
franksong2702 marked this pull request as ready for review April 17, 2026 13:21
@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Review: /compress with Focus Topic — Transcript-Aware UX ✅

This is a well-executed feature PR. The transcript-card approach for the compression flow feels right — having command / running / complete / reference as visible turn-anchored cards is much better than a toast or a hidden background action.

What I looked at

  • api/routes.py — _handle_session_compress and the resilience fallbacks
  • static/commands.js — command dispatch and state machine
  • static/ui.js — _compressionCardsHtml and card rendering
  • static/style.css — new .live-compression-cards, .compression-*, .tool-card-compress-* classes
  • tests/test_sprint46.py — new test suite (157 lines)

Backend (api/routes.py) ✅

The resilience fix in the last commit is correct and necessary: wrapping estimate_messages_tokens_rough and summarize_manual_compression in local try/except with fallbacks means the endpoint stays functional even in stripped-down test environments or when agent helpers are unavailable. The fallback summary is minimal but honest.

One small observation: _fallback_estimate_messages_tokens_rough uses word count as a proxy for tokens. That's fine as a rough estimate, but worth noting in a comment that it's intentionally ~4× off from BPE token count (and that's acceptable for a fallback).


CSS ✅ (with one flag)

Color approach is consistent with the rest of the codebase — var(--blue), var(--gold), var(--green), var(--muted) for text/foreground; rgba() for translucent borders and backgrounds. This matches the pattern used elsewhere.

Flag for #627 merge sequence: PR #627 introduces a full light/dark + accent skin system. Some of the new rgba() values here (e.g. rgba(201,168,76,.04) for gold-tinted command card background) are hardcoded to dark-theme color values. Once #627 lands, these should be expressed via the new accent skin variables if the theme system exposes semantic background tokens. Worth noting as a follow-up but not blocking — the values are structurally fine today.


Tests ✅

test_sprint46.py is solid:

  • _FakeCompressor and _FakeAgent properly stub the agent layer
  • Tests cover the compress API endpoint, focus topic passthrough, session persistence after compression, and the CI-safe isolation approach (HERMES_BASE_HOME)
  • The _make_session helper follows the same pattern as other sprint test files

UX States ✅

The four-state progression (command → running → complete → reference) matches the mental model and the screenshots in the thread look consistent. Aronprins's earlier feedback about the completion card being a collapsed green card (like the thinking card pattern) has been addressed.

Re: the margin question from @aronprins — the updated screenshots at 10:58 look good. If the bottom margin difference is still present, it can be tracked under #629 (already filed as a polish follow-up).


Summary

  • ✅ Backend: resilient, correct fallback path
  • ✅ Tests: good coverage including CI isolation
  • ✅ CSS: structurally consistent, theme-compat follow-up tracked via Compress UI polish: align command/compressing card density #629
  • ✅ UX: four-state flow, transcript-anchored, persistence handled across reload
  • 🔁 Minor: consider a comment on _fallback_estimate_messages_tokens_rough noting intentional word-count approximation

Ready to merge once the margin polish question is resolved (or confirmed as tracked in #629).

@franksong2702

Copy link
Copy Markdown
Contributor Author

I agree to merge this PR as-is from an overall UX/behavior perspective. I’m taking the follow-up correction items into issue #629 (including any remaining styling/polish items like spacing/consistency and theme/color variable follow-up) rather than expanding scope here. So this PR can be merged.

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

@franksong2702 — confirmed. The condition from my earlier review ("ready to merge once the margin polish question is resolved or confirmed as tracked in #629") is now met. #629 is filed and you've acknowledged it.

PR #619 is merge-ready. ✅

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

New commit (chore(compress): align command card density with running card) — this looks like it directly addresses the card density difference @aronprins flagged. If the command card spacing now visually matches the running card, the margin concern from the review thread is resolved in-PR rather than deferred to #629.

PR #619 remains merge-ready. ✅ The density alignment commit is a welcome polish addition.

@franksong2702

Copy link
Copy Markdown
Contributor Author

Added the requested follow-up note in documenting that the fallback uses word-count approximation (not true BPE token count) and is intentional for resilience. commit: b338189 on codex/issue-469-compress

@franksong2702

Copy link
Copy Markdown
Contributor Author

Done on the same 619 branch (commit b338189): added an explicit note on _fallback_estimate_messages_tokens_rough that the fallback uses word-count as a rough heuristic and is intentionally approximate (non-BPE / not exact token counting), for resilience when metadata helpers are unavailable.

@franksong2702

franksong2702 commented Apr 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Update: PR #619 now also includes the reference-card fallback recovery for compression completion.

Merged on this branch:

So the polish and the small robustness fix stay together on the same PR chain, and #629 is treated as a closed follow-up once this lands.

@franksong2702

franksong2702 commented Apr 18, 2026 •

Copy link
Copy Markdown
Contributor Author

before
截屏2026-04-18 11 13 27

after
截屏2026-04-18 11 22 45

Parameter diff:

  • compression-card-row + .compression-card-row: margin-top: 2px -> margin-top: 4px
  • .tool-card-compress-command: background: rgba(201,168,76,.04) -> rgba(201,168,76,.02)
  • .tool-card-compress-command: border-color: rgba(201,168,76,.22) -> rgba(201,168,76,.16)
  • .tool-card-compress-command .tool-card-name: kept at font-weight: 600 so the lighter command card still reads clearly

So the screenshot diff is intentionally subtle: slightly looser spacing and a lighter command card weight to better match the running/completion/reference cards.

@franksong2702

Copy link
Copy Markdown
Contributor Author

@nesquena @aronprins PR #619 now includes the #629 polish fixes as well, including the reference-card fallback recovery. If everything looks good on your side, please review this one PR and we can merge it in one shot.

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Re-reviewing after the latest commits (b338189, d986f6f, 015a66a).

What's been added since my last pass

  1. b338189 — fallback token estimate documentation: Comment added to _fallback_estimate_messages_tokens_rough clarifying it's intentionally word-count-based (not BPE) for resilience. This addresses the reviewer note cleanly — future readers won't wonder why it's approximate.

  2. 015a66a8 — command card density alignment: Subtle CSS delta (margin-top: 2px → 4px, lighter amber tints rgba(...,.04)→.02 / rgba(...,.22)→.16) to visually match the running/completion/reference card weight. The before/after screenshots confirm the intent — the command card now reads lighter and slightly more spacious. This resolves the density concern @aronprins raised.

  3. d986f6f — reference-card fallback recovery: Restores the reference card from the summary content on compression completion. This is the right robustness fix — the third card no longer goes blank if state was dropped between the compressing and completed phases.

Overall

All three open points from the review thread are now addressed in-PR:

  • ✅ Fallback token approximation is documented
  • ✅ Card density aligned across compress lifecycle states
  • ✅ Reference-card persistence after completion

PR #619 is ready to merge. ✅ The decision to land #629 fixes directly here rather than as a follow-up is the right call — keeps the feature cohesive.

@nesquena

Copy link
Copy Markdown
Owner

Independent End-to-End Review — PR #619

Independent fourth-pass review after the three prior approvals. Feature PR, 1249 additions across 10 files. Parallelized security audit and test run; found one minor hardening that I pushed.

TL;DR

Merge-ready after rebase onto current master. Feature is well-scoped, tests green, security clean. Pushed one small defensive cap on focus_topic input length (commit fc0ec04). Branch is 7 commits behind master — author or merger will need to rebase, which will also pull in the v0.50.71–v0.50.76 CHANGELOG/version history the branch is missing.

Test results ✅

  • 1322 passed, 42 skipped, 0 failed (full suite in isolated worktree)
  • All 3 tests in test_sprint46.py pass:
    • test_session_compress_requires_session_id ✅
    • test_session_compress_roundtrip ✅
    • test_static_commands_js_registers_compress_alias ✅
  • CI green on Python 3.11/3.12/3.13 ✅

Security audit ✅

Traced every user-input and rendering path. Clean.

Area Finding
/api/session/compress auth/CSRF ✅ Inherits both via handle_post() + _check_csrf() + check_auth()
session_id input ✅ Stringified + stripped; get_session() handles sanitization
focus_topic input ⚠️ No length cap (fixed below)
Concurrency ✅ Properly guarded by _get_session_agent_lock(sid) + active_stream_id check (409 on race)
Error paths ✅ Uses _sanitize_error() to strip filesystem paths
Response serialization ✅ Wrapped in redact_session_data(...)
Session metadata persistence ✅ compression_anchor_* only written server-side; client reads bound-checked with typeof guards
XSS in card rendering ✅ Every user/agent string passes through esc() — cmdText, statusLabel, previewText, focusText, summary.*, errorText, reference card contents, data-raw-text attrs
CSS attacks ✅ No attr() selectors pulling user data into content, no external url(), no @import
Fallback handlers ✅ Both _fallback_estimate_messages_tokens_rough and _fallback_summarize_manual_compression fail safely — pure string/int arithmetic, no external deps

Follow-up pushed (fc0ec04)

Cap focus_topic to 500 chars:

# Before:
focus_topic = str(body.get("focus_topic") or body.get("topic") or "").strip() or None
# After:
focus_topic = str(body.get("focus_topic") or body.get("topic") or "").strip()[:500] or None

Not a security vuln (user prompting themselves — no privilege boundary crossed), but matches the defensive input-size pattern used elsewhere (session title :80], first-exchange snippets :500]). Cheap bound-checking against a multi-MB payload wasting tokens/memory before the LLM call errors. Suggested by the security audit.

Code review observations

Backend (api/routes.py) — Well-structured. The resilience layer (_fallback_* helpers) is a nice touch for CI environments and stripped-down deployments. The prior reviewer's suggestion of a code comment on _fallback_estimate_messages_tokens_rough is addressed in commit b338189.

Card state machine (static/ui.js) — Four-state flow (command / running / complete / reference) correctly gates on:

  • _compressionStates[sid] map for in-flight state
  • session.compression_anchor_* metadata for persisted reference placement
  • visWithIdx[insertionAnchor] with out-of-range fallback to inner.appendChild()

All bounds-checked.

Command dispatch (static/commands.js) — /compress primary + /compact alias is clean. _compressionAnchorMessageKey() truncates to 160 chars (internal dedup key, not prompt input). encodeURIComponent(sid) correctly used on GET preflight.

Persistence — Session metadata fields (compression_anchor_visible_idx, compression_anchor_message_key) added to models.py::compact() and defensively read with type guards on the client. No assumption about field presence.

Outstanding: rebase needed before merge

The branch is 7 commits behind master (v0.50.71 through v0.50.76 content not present):

  • f3f23ab fix(csp): allow external https images
  • d6267f4 chore: CHANGELOG v0.50.75 + version badge
  • e7b8ab4 fix: harden test server isolation
  • 79428f9 fix: catch OSError from SETTINGS_FILE.exists()
  • a2ea15b fix: add favicon + static MIME types
  • 692ba68 fix(title): strip markdown labels
  • 2484409 fix: HERMES_WEBUI_DEFAULT_WORKSPACE

A rebase before merge will pull these in, bringing the branch's CHANGELOG up to current. Then a v0.50.81 entry + version badge bump can be added on top (the other 4 in-flight PRs — #647, #648, #649, #640 — are bumping to v0.50.77, v0.50.78, v0.50.79, v0.50.80 respectively if they all land first).

I didn't preemptively add a v0.50.81 CHANGELOG entry because it would conflict with master's v0.50.71–v0.50.76 entries during the rebase.

Summary

Aspect Status
Tests ✅ 1322 passed, 0 failed (3 new sprint46 tests)
Security ✅ Clean (focus_topic cap added as minor hardening)
Auth/CSRF ✅ Inherited correctly
XSS in new card rendering ✅ All user input escaped
Session metadata persistence ✅ Bounds-checked on read
Prior review items (fallback docs, density, reference fallback) ✅ All addressed in commits b338189, 015a66a, d986f6f
Rebase on master ⏳ Needed before merge
CHANGELOG + version bump ⏳ Add after rebase (v0.50.81 suggested)

Merge-ready pending rebase. The feature is cohesive, well-tested, and the prior reviewers' feedback is all addressed. Thanks @franksong2702 for the thorough UX iteration and @aronprins for the card-design feedback. Great collaboration pattern in this PR.

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Review — merge-ready after one CSS fix

Full end-to-end review complete. The feature is well-implemented. One bug required a fix before merge:

Bug: var(--green) undefined on all themes

--green was used in two places (.tool-card-compress-complete .tool-card-name and .compression-complete-header) but is not defined in any :root, :root.dark, or skin block. The completion card header and name would render with no color (browser falls back to inherited text color, usually muted gray).

Fixed by replacing var(--green) with #4ec984 (matches rgba(78,201,132,...) already used for the card's background/border — consistent visual treatment).

Also added: CHANGELOG entry for v0.50.82 describing the /compress feature and attribution to @franksong2702.

What looks good:

  • api/routes.py _handle_session_compress: input validation, 409 streaming guard, agent lock, token estimation fallback, error sanitization — all correct. SAFE.
  • /compress and /compact (alias) registration in commands.js — lock lifecycle correct on all exit paths
  • Transcript cards (command/running/complete/reference) cleanly rendered via msgInner injection
  • i18n: all 5 locales complete
  • api/models.py anchor fields backward-compatible (None defaults, **kwargs)
  • Tests: 3/3 compress-specific tests pass

Minor items tracked in #629 (not blocking this PR):

  • #liveCompressionCards div in index.html is always cleared (unused DOM leftover)
  • Historical compaction messages don't show timestamp labels (tsTitle silently dropped)
  • "No reduction" case shows N→N message count without special handling

Test results: 4 failed (pre-existing test_sprint34.py), 1372 passed. Clean.

The fix commit is on the integration branch. This PR is ready for independent review and merge.

nesquena-hermes pushed a commit that referenced this pull request Apr 18, 2026
…469 (PR #619)

POST /api/session/compress runs real compression via the agent's context_compressor.
Accepts optional focus_topic (capped at 500 chars). Replaces the old /compact
agent-message hack with a proper transcript-inline UX: command card (gold),
running card (blue, animated), collapsible complete card (green, shows delta),
reference card (full compaction summary). /compact is kept as an alias.
Fallback token estimation uses word-count (intentional, for resilience).

Fix (review): var(--green) was undefined on all themes — replaced with #4ec984.
Fix (review): focus_topic capped at 500 chars (fc0ec04 by @nesquena).

Co-Authored-By: franksong2702 <138988108+franksong2702@users.noreply.github.com>
Co-Authored-By: Nathan Esquenazi <nesquena@gmail.com>
nesquena-hermes added a commit that referenced this pull request Apr 18, 2026
…469 (PR #619 by @franksong2702)

POST /api/session/compress with optional focus_topic. Transcript-inline cards: command, running, complete (collapsible green), reference. /compact alias kept. Fixes: var(--green) undefined color, focus_topic 500-char cap. Independent review by @nesquena (4 passes).
@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Merged via integration branch #663 as commit b49de92 (v0.50.82). Full attribution to @franksong2702 for this substantial feature PR — the /compress flow, transcript cards, and /compact alias are all working. Thanks also to @aronprins for the UX feedback on the completion card styling.

@franksong2702
franksong2702 deleted the codex/issue-469-compress branch April 25, 2026 00:35
JKJameson pushed a commit to JKJameson/hermes-webui that referenced this pull request Apr 25, 2026
…esquena#469 (PR nesquena#619 by @franksong2702)

POST /api/session/compress with optional focus_topic. Transcript-inline cards: command, running, complete (collapsible green), reference. /compact alias kept. Fixes: var(--green) undefined color, focus_topic 500-char cap. Independent review by @nesquena (4 passes).
SysAdminDoc pushed a commit to SysAdminDoc/hermes-webui that referenced this pull request Jun 26, 2026
…esquena#469 (PR nesquena#619 by @franksong2702)

POST /api/session/compress with optional focus_topic. Transcript-inline cards: command, running, complete (collapsible green), reference. /compact alias kept. Fixes: var(--green) undefined color, focus_topic 500-char cap. Independent review by @nesquena (4 passes).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants