Skip to content

fix: stop long sessions ratcheting into full-transcript re-renders - #5

Draft
alanjds wants to merge 12 commits into
masterfrom
claude/webui-transcript-reload-ratchet
Draft

alanjds wants to merge 12 commits into
masterfrom
claude/webui-transcript-reload-ratchet

Conversation

@alanjds

@alanjds alanjds commented Aug 24, 2026 •

Copy link
Copy Markdown
Owner

Split 1 of 4 from a combined branch (alanjds/hermes-webui#2). Affects the default path (virtualize_transcript off) — this is the one that removes the cliff.

Problem

renderMessages() rebuilds the whole transcript and runs on every send, every SSE batch and every refresh. On a synthetic 2000-message session (desktop Chromium; a phone is several times slower), with the shipped default virtualize_transcript=false:

loaded rows render DOM nodes
37 46 ms 960
537 546 ms 13,460
2000 2,883 ms 50,007

Cost tracks S.messages.length exactly. The server is not involved — Session.load() is 4.6 ms and GET /api/session ~330 ms flat at both 200 and 2000 messages. So the bug is whatever loads the whole transcript into the browser.

Two things did, and together they formed a one-way ratchet nothing undid short of a page reload.

1. The bare full-transcript refetch

_messageReloadLimitForSession() returned null — "send no msg_limit" — when the previous load wasn't truncated, or when the reload hint exceeded the server ceiling. The hint grows to the loaded row count, so paging back crossed the 500-row ceiling and from then on every same-session refresh (tab focus, SSE catch-up, visibility change — which fire aggressively on mobile) refetched and re-rendered all 2000 rows.

That null was deliberate: a clamped window would silently drop already-loaded older rows (nesquena#6152/nesquena#6154). The window is now always bounded, and the invariant is upheld differently — rows older than the returned window that we already hold are retained and re-prepended.

The re-join is skipped, letting the server window win, when the prior rows can't be trusted to line up: the session shrank server-side (fork / undo / truncate), or we don't hold enough rows to cover the gap. Losing older rows to a re-fetch is recoverable via "load earlier"; splicing a stale prefix onto a rewritten transcript is not.

loadSession() clears S.messages and _oldestIdx before the reload fetch, so the pre-clear offset is stashed alongside the existing carry-forward snapshot and consumed with it.

2. The render-window ratchet

nesquena#6999 capped the auto-expanding render window at 4× the default, but the two stream-completion sites in messages.js still expanded it to every loaded row — one completed turn undid the cap. All three sites now share one bounded helper.

This restores the nesquena#6999 invariant; it does not by itself bound render cost (with virtualization off the DOM is built for every loaded row regardless). Part 1 is what removes the cliff.

Deliberately out of scope

The render window is also expanded to the full transcript by jump-to-session-start, jump-to-message when the target is outside the window, and the outline jump; and outline.js still issues a bare no-msg_limit fetch. Those are explicit user requests to reveal the whole transcript, addressed by absolute index — windowing them needs a jump-to-index-with-context fetch.

Verification

  • Reload requests always carry msg_limit, verified by intercepting the page's real /api/session traffic.
  • Render window stays at the 200 cap across repeated turn completions.
  • Targeted suite green on current master. Note the full tests/ run is too flaky to use as a signal — ~45–50 failures on both this branch and pristine master, differing in both directions, all in test_workspace_git.py / test_webui_state_db_reconciliation.py.

test_webui_external_refresh_frontend.py pinned the exact source strings this changes; those assertions are updated to the new mechanism, plus two new tests (the retained-prefix re-join and its bail-outs; no expansion site still using the raw uncapped count).

claude and others added 12 commits August 24, 2026 20:22
…ders

A conversation of a few thousand messages becomes progressively unusable:
every renderMessages() rebuilds the whole transcript, and renderMessages()
runs on every send, every SSE batch and every refresh. Measured on a
synthetic 2000-message session (desktop Chromium; a phone is several times
slower), with the shipped default of virtualize_transcript=false:

    loaded rows   render   DOM nodes
             37     46 ms         960
            537    546 ms      13,460
           2000  2,883 ms      50,007

Render cost tracks S.messages.length exactly, so the bug is whatever loads
the whole transcript into the browser. Two things did, and together they
formed a one-way ratchet that no later action undid.

1. _messageReloadLimitForSession() returned null — "no msg_limit", i.e. the
   bare full-transcript request — in two cases: when the previous load was
   not truncated, and when the reload hint exceeded the server's msg_limit
   ceiling. The hint grows to the loaded row count, so paging back through a
   long conversation crossed the 500-row ceiling and every subsequent
   same-session refresh (tab focus, SSE catch-up, visibility change) refetched
   and re-rendered all 2000 rows. The null was deliberate: a clamped window
   would silently drop already-loaded older rows (nesquena#6152/nesquena#6154).

   The window is now always bounded, and that invariant is upheld a different
   way — rows older than the returned window that we already hold are retained
   and re-prepended, so a bounded refresh still cannot lose history. The
   re-join is skipped, letting the server window win, when the prior rows
   can't be trusted to line up (the session shrank server-side via fork/undo/
   truncate, or we don't hold enough rows to cover the gap): losing older rows
   to a re-fetch is recoverable via "load earlier", splicing a stale prefix
   onto a rewritten transcript is not. loadSession() clears S.messages and
   _oldestIdx before the reload fetch, so the pre-clear offset is stashed
   alongside the existing carry-forward snapshot and consumed with it.

2. nesquena#6999 capped the auto-expanding render window at 4x the default so it could
   not grow to the whole loaded transcript, but the two stream-completion sites
   in messages.js still expanded it to every loaded row. A single completed
   turn silently undid the cap. All three sites now share one bounded helper in
   ui.js so the bound cannot drift apart again. This restores the nesquena#6999
   invariant (it does not by itself bound render cost — with virtualization off
   the DOM is built for every loaded row regardless of the window).

Deliberately out of scope, per the "fix the class" rule: the render window is
also expanded to the full transcript by jump-to-session-start (ui.js),
jump-to-message when the target is outside the window (ui.js), and the outline
jump (outline.js), and outline.js still issues a bare no-msg_limit fetch.
Those are explicit user requests to reveal the whole transcript, addressed by
absolute index; windowing them needs a jump-to-index-with-context fetch, which
is a separate change.

Verified: reload requests always carry msg_limit; the render window stays at
the 200 cap across repeated turn completions instead of ratcheting to the
loaded row count; 2455 tests across the 155 files that read the changed JS
pass (the 13 failures in test_update_banner_fixes.py are pre-existing
test-ordering pollution — identical on a clean tree, and the file passes
standalone both before and after).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018y9pr7ppZgDuCZCT6x96kY
…x is provably identical; compression-anchor coverage (nesquena#6826 r3)
Removes the local re-import of api.streaming._session_payload_with_full_messages
inside test_in_tail_duplicate_guard_refuses_bounded_and_full_revision_accepted;
it is already imported at module level (line 17) and unused in the function body,
tripping F401 on CI's hosted lint. Mechanical maintainer fix per auto-fix policy.

Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
Bounded sidecar-anchored tail read for regenerate (closes nesquena#6826). Thanks @webtecnica — six rounds of adversarial re-gate to convergence.
…ail read (nesquena#7204, @webtecnica)

Adds the CHANGELOG entry for nesquena#7204 (closes nesquena#6826). Release-metadata only;
the code change ships via the merge of nesquena#7204 (contributor webtecnica).
Release exp-v0.52.264: fast regenerate via bounded sidecar-anchored tail read (nesquena#7204, @webtecnica)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants